What Azure currently offers for H100, H200 and GB200 workloads
Microsoft Azure separates large, tightly coupled AI clusters from smaller applied-AI VM shapes. The ND H100 v5 family uses eight NVIDIA H100 Tensor Core GPUs with 80 GB per accelerator. Microsoft documents 96 vCPUs, 1,900 GiB of system memory and dedicated 400 Gbps NVIDIA Quantum-2 CX7 InfiniBand connectivity for each GPU. The design is intended for deep-learning training, tightly coupled generative AI and HPC workloads that need scale-out communication across many GPUs.
The ND H200 v5 family keeps the eight-GPU architecture but moves to NVIDIA H200 accelerators with 141 GB of memory each, for 1,128 GB of aggregate accelerator memory per VM. Microsoft documents 900 GB/s NVLink within the VM and 3.2 Tb/s of scale-out interconnect bandwidth per VM, again with dedicated 400 Gbps InfiniBand connectivity per GPU. That makes H200 relevant when model size, KV-cache requirements or scientific datasets create a meaningful memory-capacity bottleneck relative to H100.
The ND GB200 v6 family changes the system architecture more substantially. Azure documents two NVIDIA Grace CPUs and four Blackwell GPUs per VM, with 192 GiB of accelerator memory per GPU. Eighteen of these VMs can participate in an NVIDIA GB200 NVL72 rack-scale system containing up to 72 GPUs. This is a different procurement and deployment class from a conventional single VM with attached GPUs: the relevant unit becomes the rack-scale AI system, its NVLink domain and the surrounding network and storage architecture.
Azure also exposes H100 in smaller shapes. The NCads H100 v5 family provides one or two H100 NVL GPUs with 94 GB each, while NCCads H100 v5 provides a one-GPU H100 configuration with a trusted execution environment spanning the CPU and attached GPU. These options suit applied AI, batch inference, development and security-sensitive workloads that do not require an eight-GPU ND node.
Azure’s VM family is part of the accelerator decision
| Azure family | GPU layout | Memory context | Primary decision signal |
|---|---|---|---|
| ND H100 v5 | 8× H100 | 80 GB per GPU | Tightly coupled H100 training and HPC |
| ND H200 v5 | 8× H200 | 141 GB per GPU | Memory-intensive training and inference |
| ND GB200 v6 | 4× Blackwell GPU + 2× Grace CPU per VM | 192 GiB per GPU | NVL72-oriented frontier AI infrastructure |
| NCads H100 v5 | 1–2× H100 NVL | 94 GB per GPU | Smaller applied-AI and inference workloads |
| NCCads H100 v5 | 1× H100 | 94 GB | Confidential GPU workload requirements |
“Azure H100” is therefore not a single product. ND H100 and NCads H100 differ in GPU count, memory, networking and scaling model, so comparisons should name the exact Azure series and size.
Networking is central to Azure’s distributed-training proposition
For ND H100 v5, Azure documents eight H100 GPUs linked within the VM and a dedicated 400 Gbps Quantum-2 CX7 InfiniBand connection for each GPU. The H200 family follows the same scale-out logic, supporting GPUDirect RDMA and distributed workloads built around NCCL. This matters because distributed training can become communication-bound; adding GPUs does not guarantee useful scaling when synchronization and checkpoint traffic overwhelm the fabric.
GB200 shifts the scale-up boundary further. Azure documents fifth-generation NVLink connectivity inside the GB200 VM and an NVL72 architecture in which multiple VMs form a rack-scale system. For buyers, the key question becomes how much work remains inside the high-bandwidth scale-up fabric before traffic reaches the scale-out network.
Smaller NC H100 families are designed for different patterns and should not be assumed to expose the same InfiniBand topology as ND clusters. That makes Azure’s broad GPU catalog also creates a comparison hazard: two H100 offers may target different workload architectures.
Azure GPU pricing requires a full VM and purchasing-model comparison
Azure does not expose one permanent global “H100 price” or “H200 price” that can responsibly represent the whole service. Microsoft directs users to Azure pricing pages and the pricing calculator, where cost depends on the exact VM series and size, region, operating system, reservation or savings commitment, and other attached resources. Availability and quota can also constrain the practical buying path for high-demand accelerator families.
For AIDataCenterHQ comparisons, the minimum pricing record should therefore preserve the Azure VM SKU, accelerator count, accelerator memory, region, currency, billing unit, commitment assumption, timestamp, and whether storage or networking costs are included. An eight-GPU ND node should be normalized to both VM-hour and, where useful, a derived GPU-hour figure. That derived number still does not make the service economically equivalent to a marketplace GPU rental because the Azure VM also contains substantial CPU, system memory, local storage and network capacity.
Decision rule
Use published GPU-hour figures only after normalizing the entire workload unit. For distributed training, compare time-to-train, cluster utilization, storage throughput, checkpoint behavior and inter-node efficiency. A lower nominal GPU-hour can produce a higher training-run cost if the infrastructure scales poorly or remains underutilized.
Security requirements can change the Azure GPU choice
Azure’s NCCads H100 v5 series introduces a useful differentiator for organizations that need a trusted execution environment extending across both CPU and GPU. Microsoft describes the series as supporting confidential GPU offload of data, models and computation. It creates a distinct infrastructure option where confidential computing is part of the technical requirement.
Provider comparisons should therefore consider hardware-backed isolation, key management, identity, logging, data residency, model-weight handling and attestation, not certifications alone.
Which Azure GPU family fits which workload?
ND H100 v5 fits mature distributed training stacks that need eight H100s per node and high-bandwidth scale-out networking. It can also support large inference deployments where H100 remains sufficient.
ND H200 v5 becomes more compelling when memory capacity and bandwidth are primary constraints. Larger models, longer contexts and KV-cache-heavy inference can benefit from H200’s 141 GB memory footprint, subject to model architecture and serving strategy.
ND GB200 v6 is aimed at a higher scale of AI system design. Its Grace Blackwell architecture and NVL72 orientation make it relevant to frontier training and high-throughput inference where rack-scale compute, memory bandwidth and fast accelerator-to-accelerator communication are fundamental requirements.
NCads H100 v5 can be a better fit for one- or two-GPU development, fine-tuning, inference and applied-AI jobs that do not need an eight-GPU ND node. NCCads H100 v5 deserves separate evaluation when confidential GPU execution is required. These smaller families broaden Azure beyond eight-GPU nodes.
How Azure compares conceptually with AWS, Google Cloud and neoclouds
Azure belongs to the hyperscaler category, so the GPU decision is usually tied to a broader cloud operating environment: identity, virtual networks, managed storage, monitoring, governance, enterprise agreements and adjacent data or AI services. This can reduce integration friction for organizations already standardized on Microsoft infrastructure.
Specialist GPU clouds may emphasize faster self-service access, simpler accelerator-specific pricing or a narrower AI-first operational surface. AWS and Google Cloud bring their own hyperscaler architectures and purchasing models. The useful comparison is workload-level: accelerator availability, network topology, security, region, commitment structure, surrounding services and total job economics.
Continue with the AWS GPU Cloud and Google Cloud GPU profiles, or use the GPU Cloud Providers directory to compare specialist providers including CoreWeave, Lambda, RunPod and Crusoe.
Constraints, caveats and comparison risks
- Regional availability changes. A VM family documented by Azure may not be provisionable in every region or subscription at the moment a workload needs capacity.
- Quota and capacity matter. High-end accelerators can require quota planning, capacity reservations or sales engagement rather than immediate self-service provisioning.
- VM families are not interchangeable. NC H100, ND H100 and confidential H100 configurations target different workload patterns.
- Hourly VM price is incomplete. Storage, data transfer, managed services, support and idle time can materially change total cost.
- Specifications are upper bounds, not workload guarantees. Real scaling depends on software, communication patterns, data pipelines and tuning.
Methodology, evidence and freshness
This page uses Microsoft Learn documentation for ND H100 v5, ND H200 v5, ND GB200 v6, NCads H100 v5 and NCCads H100 v5 as primary technical evidence. Specifications and current family mappings were checked on 28 September 2026. Azure pricing is treated as a fast-changing field and is intentionally described through a normalization methodology rather than a static universal rate where a single official figure cannot represent all regions and purchasing models.
Primary evidence: Azure ND H100 v5 documentation, Azure ND H200 v5 documentation, Azure ND GB200 v6 documentation, Azure NCads H100 v5 documentation, and Azure confidential H100 documentation.
Azure GPU procurement can also intersect with AI infrastructure contracting, technology commercialization, patent strategy and cross-border deployment. Relevant specialist resources include Patent Business Lawyer, GIP Research, GIPResearch.org, US Tech Law Attorney, Patent Business Attorney, International Patents, TechCorpLegal and Advocate Rahul Dev.
Frequently asked questions
Does Microsoft Azure offer NVIDIA H100 GPUs?
Yes. Azure documents H100 in ND H100 v5 for eight-GPU scale-out systems and in smaller NCads/NCCads H100 families for applied-AI or confidential-computing workloads.
Does Azure offer NVIDIA H200?
Yes. Azure’s ND H200 v5 series uses eight H200 GPUs with 141 GB of accelerator memory each and high-bandwidth InfiniBand for distributed AI and HPC.
Does Azure offer NVIDIA GB200?
Yes. Azure documents ND GB200 v6 with two Grace CPUs and four Blackwell GPUs per VM, designed to participate in GB200 NVL72 rack-scale systems.
Is Azure GPU pricing directly comparable with RunPod or CoreWeave?
Only after normalizing the exact VM series, GPU count, region, purchasing model, network, storage and expected utilization. Hyperscaler VM pricing and specialist GPU rental are not automatically equivalent units.

