What Google Cloud currently offers for H100, H200, B200 and GB200 workloads

Google Cloud separates its accelerator-optimized infrastructure into machine families rather than presenting every GPU as an interchangeable add-on. For current frontier NVIDIA accelerators, the core progression is A3 for H100, A3 Ultra for H200, A4 for B200 and A4X for GB200. Google also documents A4X Max for GB300 Ultra Superchips, extending the Blackwell generation beyond the B200 and GB200 pages already relevant to most current procurement comparisons.

The A3 family includes High, Mega and Edge variants using NVIDIA H100 SXM GPUs with 80 GB of accelerator memory per GPU. Google positions A3 for compute- and memory-intensive, network-bound ML training and HPC. A3 Ultra moves to H200 and uses eight accelerators with 141 GB HBM3e each, giving 1,128 GB of aggregate GPU memory in the documented a3-ultragpu-8g configuration.

A4 is the B200 machine family. Google documents a4-highgpu-8g with eight B200 GPUs, 224 vCPUs, 3,968 GB of host memory, 12,000 GiB of Local SSD and 1,440 GB of total GPU memory. A4X is architecturally different: the documented a4x-highgpu-4g machine uses four GB200 Grace Blackwell Superchips, while the platform is based on an NVL72-oriented design and is intended for large foundation-model training and serving.

Primary sources: Google Cloud GPU machine types and accelerator-optimized machine documentation, checked 25 September 2026.

Current Google Cloud pricing shows why procurement context matters

Google Cloud publishes accelerator-optimized VM pricing by machine family and region. The current pricing sheet includes multiple consumption models rather than one permanent provider-wide GPU rate. The Iowa price sheet, for example, lists standard hourly, Dynamic Workload Scheduler Flex-start, Calendar Mode, Spot and committed-use values where supported.

MachineAcceleratorPublished pricing observationDerived per-GPU figureImportant context
a3-highgpu-8g8 × H100$88.4900/hour standard reference$11.061/H100-hourIowa price-sheet observation; other models/regions differ
a3-ultragpu-8g8 × H200$84.8069/hour standard reference$10.601/H200-hourIowa price-sheet observation
a4-highgpu-8g8 × B200$64.4400/hour DWS Flex-start$8.055/B200-hourFlex-start, not a standard on-demand quote

The table is not a ranking. The H100 and H200 observations use the standard price column shown on the current Iowa sheet, while the B200 observation is explicitly a Flex-start price because the same current table shows no ordinary standard price in that field. Mixing those values without the consumption-model label would be misleading.

Google also states that Spot prices are variable and can change, while Flex-start and Calendar Mode are Dynamic Workload Scheduler mechanisms. Buyers should therefore model effective cost using the exact region, consumption model, runtime, storage, network traffic and interruption tolerance.

Primary source: Google Cloud accelerator-optimized VM pricing, checked 25 September 2026.

Consumption model can change both price and capacity certainty

Google Cloud supports several paths for accelerator-optimized capacity. Depending on the machine family, a workload may use ordinary provisioning, Spot VMs, Flex-start, reservations, future reservations or committed-use structures. Google specifically notes that some future reservation requests for supported GPU machine types require discussion with an account or sales team before submission.

This distinction matters because AI teams usually optimize for more than price. A short inference experiment may value immediate self-service access. A multi-week training run may value capacity certainty. A flexible batch workload may accept delayed start or interruption in exchange for lower compute cost. Procurement architecture should therefore be treated as part of the workload design, not as an afterthought.

For comparison with specialist providers such as CoreWeave, Lambda, RunPod and Crusoe, preserve the same fields: GPU generation, count, host configuration, network, storage, region, reservation term and timestamp.

Networking is a major differentiator for distributed training

Google Cloud documents substantial network bandwidth for the current frontier accelerator families. A4 B200 and A3 Ultra H200 machines are listed with maximum network bandwidth of 3,600 Gbps, while A4X GB200 is listed at 2,000 Gbps. Google documents GPUDirect RDMA for these platforms, which is relevant to multi-node communication-intensive workloads.

The provider also positions selected GPU machine types inside AI Hypercomputer, a broader system for performance-optimized AI infrastructure with integrations for Google Kubernetes Engine and Slurm schedulers. This can matter when the operational objective is a repeatable training platform rather than an isolated VM.

However, headline bandwidth is not the same as delivered application throughput. Model parallelism, collective communication patterns, topology, software configuration and storage can all become bottlenecks. Buyers comparing the H100, H200 and B200 should therefore evaluate end-to-end cluster architecture rather than GPU specifications alone.

Primary source: Google Cloud networking and GPU machines, checked 25 September 2026.

Google Cloud region availability is an independent procurement variable

The Compute Engine pricing and GPU documentation span a broad regional footprint, but a region being present in Google Cloud does not mean every accelerator family is available there at the same time or under the same capacity conditions. Current price-sheet navigation includes locations across North America, Europe, Asia, Australia, the Middle East, Africa and South America, while machine-family documentation directs users to check GPU regions and zones separately.

This creates a practical decision field for latency, data residency, existing data location and capacity planning. A team with data already concentrated in Google Cloud may accept a higher apparent accelerator rate if it avoids data transfer, migration and operational rework. A team starting from a portable containerized workload may compare specialist providers more aggressively on compute economics and availability.

Future AIDataCenterHQ geographic pages should therefore treat Google Cloud availability as time-sensitive evidence rather than copy one global provider statement into every country page.

Which Google Cloud GPU family fits which workload?

H100 A3: A3 remains a practical choice for mature H100 training, fine-tuning and inference workflows. Multiple A3 variants allow buyers to distinguish scale and networking needs rather than treat H100 as a single machine shape.

H200 A3 Ultra: H200 is more attractive where 141 GB per GPU reduces memory pressure or improves inference placement for larger models. The higher-memory design can be valuable even when the hourly rate is not the lowest headline number.

B200 A4: A4 targets Blackwell-class training and serving. The 8-GPU configuration and high network bandwidth make it relevant for large distributed jobs, but availability and consumption model should be confirmed before committing a workload plan.

GB200 A4X: A4X is a system-level Grace Blackwell proposition rather than simply another discrete-GPU VM. Teams evaluating it should examine software compatibility, cluster topology, scheduler design and whether the workload can benefit from the NVL72-oriented architecture.

Google Cloud is especially relevant where the broader application already depends on GCP services, data, identity, networking or Kubernetes. The correct comparison question is therefore not “Which GPU is cheapest?” but “Which provider configuration produces the best workload outcome after infrastructure dependencies are included?”

Constraints, caveats and comparison risks

  • Pricing is not one number: standard, Spot, Flex-start, Calendar Mode and commitments represent different purchase conditions.
  • Availability is regional: a machine family can exist globally while capacity remains constrained or absent in a specific zone.
  • Newer machines may require reservation workflows: A4/A4X-class infrastructure can involve capacity planning rather than instant self-service provisioning.
  • Per-GPU arithmetic can hide host value: CPU, memory, Local SSD and network differ materially between machine families.
  • Migration has a cost: a lower external GPU rate may not offset data movement, platform engineering and ecosystem integration.

For these reasons, AIDataCenterHQ does not label a single provider universally “best.” The useful decision is workload-specific and should preserve commercial and technical assumptions explicitly.

Methodology, evidence and freshness

This page uses Google Cloud Compute Engine documentation and official Google Cloud pricing pages as primary evidence. Machine specifications, consumption-model descriptions and networking information were checked on 25 September 2026. Derived per-GPU figures are simple division of a published eight-GPU machine price by eight and are labelled as derived calculations rather than provider-quoted per-GPU prices.

GPU pricing, regional availability, reservation terms and machine-family availability are fast-changing facts. Re-check the provider before procurement. AIDataCenterHQ classifies these fields as high-freshness infrastructure data and avoids presenting a timestamped observation as a permanent market ranking.

Google Cloud GPU procurement can also intersect with AI infrastructure contracting, technology commercialization, patent strategy and cross-border deployment. Relevant specialist resources include Patent Business Lawyer, GIP Research, GIPResearch.org, US Tech Law Attorney, Patent Business Attorney, International Patents, TechCorpLegal and Advocate Rahul Dev.

Frequently asked questions

Does Google Cloud offer NVIDIA H100 GPUs?

Yes. Google Cloud documents H100 GPUs in its A3 machine family, including A3 High, Mega and Edge variants.

Does Google Cloud offer H200 and B200?

Yes. H200 is documented in A3 Ultra and B200 in A4. Google also documents GB200 through A4X and GB300 through A4X Max.

Is Google Cloud GPU pricing directly comparable with RunPod or CoreWeave?

Only after normalizing GPU count, machine shape, region, network, storage and purchasing model. Google Cloud exposes several consumption options, so a single hourly figure can be misleading without context.

What is AI Hypercomputer?

Google describes AI Hypercomputer as a performance-optimized AI infrastructure system that integrates supported accelerator infrastructure with services including GKE and Slurm-oriented deployment workflows.

Nebius GPU Cloud