NVIDIA H200

NVIDIA H200

Compare NVIDIA H200 specifications, variants and workload fit before choosing a cloud GPU provider.

NVIDIA H200 cloud GPU

NVIDIA Hopper data-center GPU

Larger, faster memory for generative AI and HPC

NVIDIA H200 is a Hopper-based data-center GPU designed for memory-intensive generative AI, large language model inference, model training and high-performance computing. Its defining advantage is 141 GB of HBM3e memory with 4.8 TB/s of memory bandwidth.

  • HopperGPU architecture
  • 141 GBHBM3e GPU memory
  • 4.8 TB/sMemory bandwidth
  • 900 GB/sNVLink bandwidth

Who should choose an H200 cloud GPU?

Choose H200 when GPU memory is the bottleneck. It is especially relevant for large-model inference, long contexts, larger batches and scientific workloads that benefit from 141 GB of memory and 4.8 TB/s bandwidth. H100 may remain sufficient when 80 GB is enough.

NVIDIA H200 specifications

NVIDIA offers H200 in SXM and NVL configurations. Both provide 141 GB of HBM3e memory and 4.8 TB/s bandwidth, but their power, form factor, MIG allocation and server options differ.

SpecificationH200 SXMH200 NVL
GPU memory141 GB HBM3e141 GB HBM3e
Memory bandwidth4.8 TB/s4.8 TB/s
FP6434 TFLOPS30 TFLOPS
FP64 Tensor Core67 TFLOPS60 TFLOPS
TF32 Tensor Core*989 TFLOPS835 TFLOPS
BF16 / FP16 Tensor Core*1,979 TFLOPS1,671 TFLOPS
FP8 / INT8 Tensor Core*3,958 TFLOPS3,341 TFLOPS
Maximum TDPUp to 700 W, configurableUp to 600 W, configurable
MIGUp to 7 instances at 18 GB eachUp to 7 instances at 16.5 GB each
Form factorSXMDual-slot PCIe, air cooled
InterconnectNVLink 900 GB/s; PCIe Gen5 128 GB/s2- or 4-way NVLink bridge, 900 GB/s per GPU; PCIe Gen5 128 GB/s
Media engines7 NVDEC and 7 JPEG decoders7 NVDEC and 7 JPEG decoders
NVIDIA AI EnterpriseAdd-onIncluded

* Tensor Core figures are published with sparsity enabled. Cloud configurations vary by provider, server platform, region and allocated GPU count.

H200 advantage

Choose H200 for memory-intensive workloads

H200 increases GPU memory from 80 GB on H100 SXM to 141 GB and increases memory bandwidth from 3.35 TB/s to 4.8 TB/s. That can reduce model sharding and support larger batches or contexts.

Important distinction

Raw compute is not the main upgrade

H200 and comparable H100 variants publish the same Tensor Core performance figures. The H200 purchase decision should therefore be driven mainly by memory capacity, bandwidth and workload fit.

What makes H200 different?

Open each capability for a practical explanation of how it affects cloud AI and HPC workloads.

  • 141 GB of HBM3e memoryMore room for large models and datasets

    H200 provides 141 GB of HBM3e memory per GPU. The larger memory capacity can reduce model partitioning and support larger batches, longer contexts and more demanding scientific datasets.

  • 4.8 TB/s memory bandwidthBuilt for memory-bound AI and HPC workloads

    H200 delivers 4.8 TB/s of GPU memory bandwidth. This helps workloads that frequently move large amounts of model weights or scientific data between memory and the GPU.

  • Transformer Engine with FP8Accelerated transformer training and inference

    The Hopper Transformer Engine uses FP8 and FP16 precision to accelerate transformer workloads while helping preserve the accuracy needed for large AI models.

  • Fourth-generation Tensor CoresMixed-precision AI and scientific computing

    H200 supports FP64, TF32, FP32, BF16, FP16, FP8 and INT8 formats for large-model AI, inference and high-performance computing.

  • NVLink and NVSwitch scalingHigh-speed communication between GPUs

    H200 supports up to 900 GB/s of NVLink bandwidth per GPU. HGX and DGX systems use NVSwitch to connect multiple GPUs for distributed AI and HPC workloads.

  • Multi-Instance GPUPartition one GPU into isolated instances

    H200 supports up to seven Multi-Instance GPU partitions. MIG can improve utilization when several isolated workloads share the same physical GPU.

Best workloads for an H200 cloud GPU

  • Large-model inference

    Well suited to high-throughput inference for large language models where model weights, KV cache and long contexts place pressure on GPU memory.

  • Long-context AI workloads

    The 141 GB memory capacity can support larger context windows, bigger batches and fewer compromises when serving memory-intensive generative AI applications.

  • Large-model training and fine-tuning

    Useful for transformer training, fine-tuning and distributed workloads that benefit from high memory capacity, FP8 support and fast GPU interconnects.

  • HPC and scientific computing

    A strong fit for simulations, genomics, computational chemistry and other memory-intensive applications that need high FP64 performance and memory bandwidth.

H200 is a strong fit when…

  • Your model, KV cache or scientific dataset exceeds the practical memory limits of an 80 GB GPU.
  • Your workload is constrained by GPU memory capacity or memory bandwidth.
  • You need larger inference batches or longer context windows.
  • You want to reduce model partitioning across multiple GPUs.
  • You need NVLink or NVSwitch for distributed AI or HPC.
  • You expect sustained utilization that can justify a premium data-center GPU.

Consider another GPU when…

H200 may not deliver enough additional value when your models fit comfortably within 80 GB, your utilization is low or your workload is primarily cost-sensitive inference.

  • Compare H100 when 80 GB is sufficient for the workload.
  • Compare L40S for smaller inference, rendering and visual AI.
  • Compare B200 when maximum Blackwell-generation capability matters.

Provider comparison

Cloud providers offering NVIDIA H200

Compare the H200 variant, GPU count, attached compute, deployment region, starting price and purchasing options. Check whether the listing uses H200 SXM, H200 NVL or a complete multi-GPU HGX system.

View all providers
ProviderH200 offeringGPU memoryExample configurationRegionsStarting priceBilling optionsDetails
AceCloud1× NVIDIA H200 NVL141 GB HBM3e16 vCPU · 128 GB RAMIndia and United States₹381.46/hour · ₹2,22,775/monthHourly · monthly · 6-month · 12-monthView provider
E2E Networks1× NVIDIA H200141 GB HBM3e30 vCPU · 375 GB RAMIndia₹436/hour · ₹2,72,401/monthOn-demand · monthly · 6 months · annual · committed pricingView provider
Utho1× NVIDIA H200141 GB HBM3e16 vCPU · 96 GB RAMIndiaFrom ₹1,79,995/monthMonthly · custom quote · reserved capacityView provider
Cyfuture AINVIDIA H200141 GB HBM3eCustom configurationIndia, United States and EuropeCustom quoteEnterprise quote · reserved capacityView provider
Microsoft AzureND H200 v58× 141 GB · 1,128 GB total96 vCPU · 1,850 GiB RAMSelected Azure regionsUse Azure Pricing CalculatorPay as you go · reservations · savings plan · SpotView provider
DigitalOcean1× or 8× NVIDIA HGX H200141 GB HBM3e per GPU24 vCPU · 240 GiB RAM per GPUNew York and Atlanta$4.47/GPU-hour · $3.40 reservedPer-second on-demand · 12-month reservedView provider
AWSEC2 P5e / P5en8× 141 GB · 1,128 GB total8-GPU multi-node training instanceSelected AWS regions, including MumbaiFrom $5.97/GPU-hour via Capacity BlocksOn-Demand · Savings Plans · Spot · Capacity BlocksView provider

Compare the complete instance rather than only the displayed GPU rate. Providers may offer different H200 variants, GPU counts, vCPU, RAM, storage, network fabric, regions, commitment periods and tax treatment. Multi-GPU instance pricing must not be compared directly with a single-GPU rate, and Spot capacity may be interrupted.

NVIDIA H200 FAQs

How much GPU memory does NVIDIA H200 have?

Both H200 SXM and H200 NVL provide 141 GB of HBM3e memory per GPU.

What is the memory bandwidth of NVIDIA H200?

NVIDIA lists 4.8 TB/s of GPU memory bandwidth for both H200 SXM and H200 NVL.

Is NVIDIA H200 better than H100 for LLM inference?

H200 is usually the stronger option when inference is limited by memory capacity or bandwidth. Its main advantage over H100 is 141 GB HBM3e memory and 4.8 TB/s bandwidth rather than higher raw Tensor Core specifications.

Is NVIDIA H200 suitable for LLM training?

Yes. H200 supports Transformer Engine, FP8 and high-speed NVLink, making it suitable for large-model training and fine-tuning, especially when memory is a constraint.

Does NVIDIA H200 support MIG?

Yes. H200 supports up to seven MIG instances. NVIDIA lists up to seven 18 GB instances for H200 SXM and up to seven 16.5 GB instances for H200 NVL.

What is the difference between H200 SXM and H200 NVL?

Both provide 141 GB HBM3e and 4.8 TB/s bandwidth. H200 SXM is designed for HGX systems and supports up to 700 W, while H200 NVL uses a dual-slot PCIe form factor and supports up to 600 W.

Compare H200 configurations before choosing a provider

Check whether the listing uses H200 SXM or H200 NVL, then compare GPU count, region, interconnect, attached compute, storage, networking and the complete hourly or monthly price.