What GPU infrastructure does Nebius currently offer?
Nebius lists NVIDIA HGX H100, H200, B200 and B300 as self-service GPU instances, alongside GB200 NVL72 and GB300 NVL72 systems that require direct contact. It also lists RTX PRO 6000 and L40S for workloads that do not require the highest-end training platforms. The self-service page states that these GPUs are available in data centers in Europe and the United States.
The portfolio spans two NVIDIA generations. H100 and H200 are Hopper systems. B200 and B300 are Blackwell-generation platforms, with B300 described by Nebius as Blackwell Ultra. The technical difference matters because accelerator memory, memory bandwidth, interconnect generation and model support can change the economics of a workload even when the hourly price is higher.
| GPU | Memory | Current on-demand GPU-hour | Effective 1 Oct 2026 | Network on self-service page |
|---|---|---|---|---|
| HGX H100 | 80 GB | $3.85 | $4.50 | Quantum-2 InfiniBand |
| HGX H200 | 141 GB | $4.50 | $5.40 | Quantum-2 InfiniBand |
| HGX B200 | 180 GB | $7.15 | $8.50 | Quantum-2 InfiniBand |
| HGX B300 | 270 GB | $7.85 | $9.50 | Quantum-X800 InfiniBand |
These figures are time-sensitive. Nebius already publishes the rates scheduled to take effect on 1 October 2026, so a buyer evaluating the platform after that date should use the new column rather than the September rates. Taxes are excluded from the published figures.
How should Nebius on-demand, preemptible and committed pricing be compared?
Nebius separates on-demand capacity, preemptible capacity and commitment discounts. The current price page states that large-scale clusters reserved for multiple months can receive commitment discounts of up to 35%. Preemptible GPU prices are dynamic and can change with capacity, utilization, demand and customer pricing preferences.
The current preemptible table starts H100, H200, B200 and B300 at $0.79, $0.79, $0.99 and $0.99 per GPU-hour respectively. Those figures should not be compared directly with on-demand rates without accounting for interruption risk. Nebius states that spot prices may change as frequently as every 15 minutes and that the active price is displayed in the console.
Decision implication
For checkpoint-friendly training, batch inference or experimentation, preemptible capacity can materially change compute economics. For long distributed training runs with expensive restart costs, on-demand or reserved capacity may deliver a lower effective cost even when the nominal GPU-hour rate is higher.
A complete TCO comparison should also include storage and data movement. Nebius currently lists shared filesystem storage, WEKA, object storage, block volumes and local SSD pricing separately. Its price page lists networking ingress/egress and public IP addresses as free, but buyers should still validate the exact service scope and workload architecture before treating transfer economics as zero.
What orchestration options distinguish Nebius from simple GPU rental?
Nebius provides managed Kubernetes as a default cloud-native orchestration path and Managed Soperator for Slurm-on-Kubernetes. Soperator is Nebius's Kubernetes operator for Slurm. The managed service provisions nodes, drivers and dependencies, and Nebius describes topology-aware placement, monitoring and fault-tolerance features for large training workloads.
This matters because GPU procurement alone does not solve cluster operations. Distributed training requires scheduling, health monitoring, checkpointing, storage throughput, topology awareness and failure recovery. A lower GPU-hour rate can lose its advantage if engineering teams spend significant time building and operating those layers themselves.
Nebius also documents support for SkyPilot, Ray and Anyscale workflows. That gives teams several paths: remain Kubernetes-native, use managed Slurm, or adopt higher-level workload schedulers. The relevant comparison with AWS, Google Cloud, Azure, CoreWeave or RunPod should therefore include the orchestration model that the engineering team already uses.
How do networking and storage affect Nebius cluster economics?
The self-service catalog identifies Quantum-2 InfiniBand for H100, H200 and B200 and Quantum-X800 InfiniBand for B300. High-bandwidth, low-latency scale-out networking is important when a training job spans multiple GPU nodes because communication overhead can reduce useful accelerator utilization.
Storage is equally important. Nebius offers shared filesystems, object storage and block storage, while its Managed Soperator materials describe a shared root filesystem for Slurm clusters. For checkpoint-heavy training, the ability to feed and persist data at sufficient throughput can be as important as raw GPU specifications. Buyers should therefore compare measured workload throughput and checkpoint recovery rather than using GPU model alone as the selection criterion.
For very large clusters, availability becomes a procurement variable. Nebius offers immediate self-service for selected GPU instances, while large-scale reserved capacity is a separate sales path. A sourcing plan should distinguish what can be launched from the console today from what requires a capacity agreement.
Which workloads fit Nebius, and when should alternatives be compared?
H100 remains relevant for established training and inference stacks where software compatibility and broad availability matter. H200 adds substantially more memory per GPU and can be attractive for memory-bound inference, long-context workloads and models that otherwise require more accelerator sharding. B200 and B300 target teams moving to newer Blackwell capabilities and larger memory footprints.
Nebius can be particularly relevant when a team wants specialist AI infrastructure with self-service access plus managed Kubernetes or Slurm. Hyperscalers such as AWS, Google Cloud and Microsoft Azure may be preferable when GPU compute must sit inside a broader enterprise cloud estate, identity model or managed-service stack. Specialist providers such as CoreWeave, Lambda, RunPod and Crusoe should also be compared where procurement flexibility, specific GPU inventory or cluster design is decisive.
For accelerator-specific sourcing, use the H100, H200 and B200 guides. The GPU Cloud Pricing Comparator provides a separate utility layer for normalized pricing decisions.
Methodology, evidence and freshness
This page uses Nebius's official pricing, self-service, H200 and orchestration pages as primary evidence. Pricing and current product availability were checked on 29 September 2026. Nebius pricing is classified as fast-changing information: the provider already publishes different rates effective 1 October 2026, and preemptible rates can change dynamically.
Primary evidence: Nebius AI Cloud pricing, Nebius self-service GPU catalog, Nebius H200 infrastructure, and Nebius orchestration documentation.
The comparison framework treats posted prices as inputs rather than conclusions. A normalized decision should record currency, billing unit, GPU model, region, commitment assumption, interruption tolerance, storage, networking and timestamp. It should then relate those inputs to expected utilization and workload completion time.
Nebius GPU procurement can also intersect with technology transactions, AI infrastructure commercialization, patent strategy and cross-border deployment. Relevant specialist resources include Patent Business Lawyer, GIP Research, GIPResearch.org, US Tech Law Attorney, Patent Business Attorney, International Patents, TechCorpLegal and Advocate Rahul Dev.
Frequently asked questions
Does Nebius offer NVIDIA H100 and H200?
Yes. Nebius currently lists both HGX H100 and HGX H200 as self-service GPU instances, with H200 providing 141 GB of GPU memory and Quantum-2 InfiniBand.
Does Nebius offer NVIDIA B200 and B300?
Yes. HGX B200 and HGX B300 are currently listed in self-service. Nebius also lists GB200 NVL72 and GB300 NVL72 through a contact-led procurement path.
Does Nebius publish GPU prices?
Yes. Nebius publishes on-demand GPU-hour prices and separate preemptible rates. It also publishes new on-demand rates that become effective on 1 October 2026, so the timestamp is material.
Can Nebius run Slurm workloads?
Yes. Managed Soperator provides a managed Slurm-on-Kubernetes environment, while Kubernetes, SkyPilot, Ray and Anyscale workflows are also supported.

