What AWS currently offers for H100, H200 and B200 workloads

AWS separates its highest-end NVIDIA accelerators into different EC2 families. The P5 family includes p5.4xlarge with one NVIDIA H100 GPU and p5.48xlarge with eight H100 GPUs. The eight-GPU system exposes 640 GiB of total accelerator memory, while the single-GPU size provides 80 GiB. That single-GPU option is important because it gives AWS a smaller H100 entry point for workloads that do not need a full HGX-class eight-GPU node.

P5e and P5en move to NVIDIA H200. AWS documents both as eight-GPU configurations with 1,128 GiB of aggregate accelerator memory, which corresponds to 141 GiB per GPU. P5en also changes the host platform and is designed around high-bandwidth networking for distributed AI and HPC workloads.

P6-B200 is the Blackwell step. AWS documents p6-b200.48xlarge with eight NVIDIA B200 GPUs. The current EC2 instance specification lists 1,432 GiB of aggregate accelerator memory. AWS also has newer P6-B300 and P6e-GB200 families, but this page keeps B200 as the primary Blackwell comparison because it connects directly to AIDataCenterHQ's live B200 cloud sourcing guide.

Primary source: AWS EC2 accelerated computing instance specifications, checked 25 September 2026.

AWS Capacity Blocks show why billing context matters

AWS publishes Capacity Blocks for ML rates for selected accelerator families. These are reserved-capacity offers, not a universal price for every launch method. AWS states that a Capacity Block price depends on supply and demand when the block is purchased, is charged up front, and does not change after reservation. The pricing page also provides an effective hourly rate for comparison.

InstanceAcceleratorExample published Capacity Blocks rateDerived accelerator rateContext
p5.48xlarge8 × H100$41.528 per instance-hour$5.191 per H100-hourExample rate shown for several supported regions
p5e.48xlarge8 × H200$47.76 per instance-hour$5.97 per H200-hourCapacity Blocks example rate
p5en.48xlarge8 × H200$49.928 to $54.920 per instance-hour in listed regions$6.241 to $6.865 per H200-hourPublished regional variation
p6-b200.48xlarge8 × B200$98.84 per instance-hour in listed commercial regions$12.355 per B200-hourCapacity Blocks example rate

These figures are useful as transparent reference observations, but they should not be collapsed into a permanent ranking. Capacity Blocks require a reservation, availability is time-sensitive, regions differ, and operating-system or ancillary charges may apply. The correct comparison record therefore retains purchase model, region, instance, GPU count and timestamp.

Primary source: Amazon EC2 Capacity Blocks for ML Pricing, checked 25 September 2026. Billing mechanics: AWS Capacity Blocks pricing and billing documentation.

AWS GPU procurement is broader than one hourly price

For GPU buyers, AWS can expose several commercial paths depending on instance type and region. The H100 p5.4xlarge announcement states that the single-GPU size can be purchased through Capacity Blocks in listed regions and is also available through On-Demand, Spot or Savings Plans in selected locations. That means the same accelerator label can sit behind different availability guarantees, interruption risk and commitment assumptions.

Capacity Blocks are useful when a team needs known capacity for a defined future period. Spot can reduce cost for interruption-tolerant workloads but adds execution risk. On-Demand can improve flexibility where available, while commitment structures may make sense for predictable longer-lived usage. None of those should be mixed in a single price column without labeling the purchase model.

This procurement diversity is one reason AWS should be compared with AI-focused specialists such as CoreWeave, Lambda, RunPod and Crusoe on more than price. Specialist providers can expose different reservation mechanics, cluster packaging, support and deployment models.

Source for single-GPU P5 purchasing modes and regions: AWS, P5 single-GPU general availability.

Networking and cluster architecture can dominate large-training economics

AWS documents 3,200 Gigabit networking for p5.48xlarge, P5e, P5en and P6-B200 configurations in its current accelerated-computing specifications. These families support Elastic Fabric Adapter, which is relevant for distributed training and HPC because inter-node communication can determine scaling efficiency once a workload expands beyond one server.

For P6-B200, AWS states that the platform uses up to 3.2 terabits per second of EFA networking and can scale within EC2 UltraClusters to very large GPU counts. The value of that architecture depends on the workload. A small inference service may gain little from an expensive multi-node fabric, while large foundation-model training can be sensitive to network topology and communication overhead.

For procurement, this means a lower GPU-hour from another provider is not automatically lower workload cost. Buyers should compare measured throughput, scaling efficiency, storage path, checkpoint behavior and time-to-result. The AI Infrastructure ROI Calculator can then model scenarios rather than treating list prices as complete economics.

Sources: AWS accelerated instance specifications and Amazon EC2 P6/P6e overview.

Which AWS GPU family fits which workload?

H100 P5: H100 remains a practical baseline for training, fine-tuning and inference where the software stack is mature and 80 GiB per GPU is sufficient. The single-GPU p5.4xlarge is particularly relevant for teams that need H100 capability without paying for eight accelerators at once.

H200 P5e/P5en: H200 is more attractive where larger memory per GPU can reduce sharding pressure or increase inference capacity. The H200 cloud guide explains why memory-sensitive workloads can justify a higher accelerator rate when they reduce runtime or simplify model placement.

B200 P6: B200 is aimed at newer Blackwell-class training and inference. AWS states that P6-B200 can deliver up to twice the performance of P5en for AI training and inference, but that provider claim should be tested against the actual model, precision, framework and utilization. Higher Capacity Blocks pricing means a buyer should quantify time-to-result rather than assume the newest accelerator automatically has the lowest total cost.

Teams should also consider the wider AWS environment. Data already stored in S3, workflows integrated with AWS identity and networking, or applications using AWS managed services can reduce migration friction. Conversely, a specialist GPU provider may offer a simpler or cheaper compute path for a workload that does not need the broader AWS platform.

AWS geographic availability is an independent decision field

GPU availability differs materially by instance family and region. AWS documentation lists instance availability region by region, and newer accelerators tend to reach fewer locations first. For example, AWS announced P6-B200 availability in Asia Pacific (Hyderabad) on 3 September 2026 and stated that the family was then available in US West (Oregon), US East (N. Virginia and Ohio), GovCloud regions, Hyderabad and Mumbai.

This has direct implications for data residency, latency, capacity planning and migration. A global organization should avoid assuming that an instance visible in the AWS catalog is launchable in every region. Region and actual capacity state belong in the provider record alongside price.

That geographic constraint will become increasingly important as AIDataCenterHQ expands into country-specific GPU-cloud intelligence. Until then, the Countries hub provides the parent geography layer.

Current B200 regional expansion source: AWS, 3 September 2026.

Constraints and questions to resolve before choosing AWS GPU infrastructure

First, confirm quota and actual capacity. A supported instance family is not the same as immediately available capacity at the required scale. Second, model the full bill. GPU compute can sit alongside storage, data transfer, snapshots, managed services, support and operational tooling. Third, identify the commercial state of every price. Capacity Blocks, On-Demand, Savings Plans and Spot carry different commitments and risks.

Fourth, test workload portability. Deep integration with the AWS platform can create real operational value, but it can also increase migration effort if economics or capacity assumptions change. Fifth, benchmark scaling behavior rather than relying only on accelerator specifications. Network and storage can become bottlenecks, especially for multi-node training.

Finally, compare cloud with owned or colocated infrastructure where usage is persistent. AIDataCenterHQ's Data Centers, Power and Cooling sections provide the physical-infrastructure context for that decision.

Methodology, freshness and limitations

This page uses current AWS documentation, AWS product announcements and AWS Capacity Blocks pricing as the primary evidence for instance configurations, published reference rates and regional availability. Values were checked on 25 September 2026 and should be treated as fast-changing F2 data.

Capacity Blocks prices shown here preserve the AWS effective hourly rate and the published per-accelerator equivalent. They do not represent private discounts, every On-Demand rate, future capacity or universal pricing across regions. Before procurement, recheck AWS pricing, confirm quota and capacity in the target region, and model ancillary costs.

Related GPU cloud pages

Adjacent technology, IP and commercialization resources

AWS GPU procurement can intersect with AI infrastructure contracting, technology commercialization, patent strategy and cross-border deployment. Relevant specialist resources include Patent Business Lawyer, GIP Research, GIPResearch.org, Tech Law Attorney, US Tech Law Attorney, Patent Business Attorney, International Patents, TechCorpLegal and Advocate Rahul Dev.

Frequently asked questions

Does AWS offer single-GPU H100 instances?

Yes. AWS documents p5.4xlarge with one H100 GPU as well as p5.48xlarge with eight H100 GPUs.

Which AWS instances use H200?

P5e and P5en use eight NVIDIA H200 GPUs with 141 GiB of accelerator memory per GPU.

Does AWS offer B200?

Yes. AWS documents P6-B200 with eight NVIDIA B200 GPUs, with regional availability that should be checked before planning deployment.

Are AWS Capacity Blocks prices the same as On-Demand prices?

No. Capacity Blocks are reserved capacity purchased for a defined period and their pricing should be treated separately from On-Demand or Spot.

Should AWS be compared only on GPU-hour price?

No. Region, networking, storage, reservation method, wider AWS integration, workload runtime and scaling efficiency can materially change total economics.

Nebius GPU Cloud