NVIDIA H200
Compare NVIDIA H200 specifications, variants and workload fit before choosing a cloud GPU provider.

NVIDIA Hopper data-center GPU
Larger, faster memory for generative AI and HPC
NVIDIA H200 is a Hopper-based data-center GPU designed for memory-intensive generative AI, large language model inference, model training and high-performance computing. Its defining advantage is 141 GB of HBM3e memory with 4.8 TB/s of memory bandwidth.
- HopperGPU architecture
- 141 GBHBM3e GPU memory
- 4.8 TB/sMemory bandwidth
- 900 GB/sNVLink bandwidth
Who should choose an H200 cloud GPU?
Choose H200 when GPU memory is the bottleneck. It is especially relevant for large-model inference, long contexts, larger batches and scientific workloads that benefit from 141 GB of memory and 4.8 TB/s bandwidth. H100 may remain sufficient when 80 GB is enough.
NVIDIA H200 specifications
NVIDIA offers H200 in SXM and NVL configurations. Both provide 141 GB of HBM3e memory and 4.8 TB/s bandwidth, but their power, form factor, MIG allocation and server options differ.
| Specification | H200 SXM | H200 NVL |
|---|---|---|
| GPU memory | 141 GB HBM3e | 141 GB HBM3e |
| Memory bandwidth | 4.8 TB/s | 4.8 TB/s |
| FP64 | 34 TFLOPS | 30 TFLOPS |
| FP64 Tensor Core | 67 TFLOPS | 60 TFLOPS |
| TF32 Tensor Core* | 989 TFLOPS | 835 TFLOPS |
| BF16 / FP16 Tensor Core* | 1,979 TFLOPS | 1,671 TFLOPS |
| FP8 / INT8 Tensor Core* | 3,958 TFLOPS | 3,341 TFLOPS |
| Maximum TDP | Up to 700 W, configurable | Up to 600 W, configurable |
| MIG | Up to 7 instances at 18 GB each | Up to 7 instances at 16.5 GB each |
| Form factor | SXM | Dual-slot PCIe, air cooled |
| Interconnect | NVLink 900 GB/s; PCIe Gen5 128 GB/s | 2- or 4-way NVLink bridge, 900 GB/s per GPU; PCIe Gen5 128 GB/s |
| Media engines | 7 NVDEC and 7 JPEG decoders | 7 NVDEC and 7 JPEG decoders |
| NVIDIA AI Enterprise | Add-on | Included |
* Tensor Core figures are published with sparsity enabled. Cloud configurations vary by provider, server platform, region and allocated GPU count.
H200 advantage
Choose H200 for memory-intensive workloads
H200 increases GPU memory from 80 GB on H100 SXM to 141 GB and increases memory bandwidth from 3.35 TB/s to 4.8 TB/s. That can reduce model sharding and support larger batches or contexts.
Important distinction
Raw compute is not the main upgrade
H200 and comparable H100 variants publish the same Tensor Core performance figures. The H200 purchase decision should therefore be driven mainly by memory capacity, bandwidth and workload fit.
What makes H200 different?
Open each capability for a practical explanation of how it affects cloud AI and HPC workloads.
141 GB of HBM3e memoryMore room for large models and datasets
H200 provides 141 GB of HBM3e memory per GPU. The larger memory capacity can reduce model partitioning and support larger batches, longer contexts and more demanding scientific datasets.
4.8 TB/s memory bandwidthBuilt for memory-bound AI and HPC workloads
H200 delivers 4.8 TB/s of GPU memory bandwidth. This helps workloads that frequently move large amounts of model weights or scientific data between memory and the GPU.
Transformer Engine with FP8Accelerated transformer training and inference
The Hopper Transformer Engine uses FP8 and FP16 precision to accelerate transformer workloads while helping preserve the accuracy needed for large AI models.
Fourth-generation Tensor CoresMixed-precision AI and scientific computing
H200 supports FP64, TF32, FP32, BF16, FP16, FP8 and INT8 formats for large-model AI, inference and high-performance computing.
NVLink and NVSwitch scalingHigh-speed communication between GPUs
H200 supports up to 900 GB/s of NVLink bandwidth per GPU. HGX and DGX systems use NVSwitch to connect multiple GPUs for distributed AI and HPC workloads.
Multi-Instance GPUPartition one GPU into isolated instances
H200 supports up to seven Multi-Instance GPU partitions. MIG can improve utilization when several isolated workloads share the same physical GPU.
Best workloads for an H200 cloud GPU
Large-model inference
Well suited to high-throughput inference for large language models where model weights, KV cache and long contexts place pressure on GPU memory.
Long-context AI workloads
The 141 GB memory capacity can support larger context windows, bigger batches and fewer compromises when serving memory-intensive generative AI applications.
Large-model training and fine-tuning
Useful for transformer training, fine-tuning and distributed workloads that benefit from high memory capacity, FP8 support and fast GPU interconnects.
HPC and scientific computing
A strong fit for simulations, genomics, computational chemistry and other memory-intensive applications that need high FP64 performance and memory bandwidth.
H200 is a strong fit when…
- Your model, KV cache or scientific dataset exceeds the practical memory limits of an 80 GB GPU.
- Your workload is constrained by GPU memory capacity or memory bandwidth.
- You need larger inference batches or longer context windows.
- You want to reduce model partitioning across multiple GPUs.
- You need NVLink or NVSwitch for distributed AI or HPC.
- You expect sustained utilization that can justify a premium data-center GPU.
Consider another GPU when…
H200 may not deliver enough additional value when your models fit comfortably within 80 GB, your utilization is low or your workload is primarily cost-sensitive inference.
- Compare H100 when 80 GB is sufficient for the workload.
- Compare L40S for smaller inference, rendering and visual AI.
- Compare B200 when maximum Blackwell-generation capability matters.
Provider comparison
Cloud providers offering NVIDIA H200
Compare the H200 variant, GPU count, attached compute, deployment region, starting price and purchasing options. Check whether the listing uses H200 SXM, H200 NVL or a complete multi-GPU HGX system.
| Provider | H200 offering | GPU memory | Example configuration | Regions | Starting price | Billing options | Details |
|---|---|---|---|---|---|---|---|
| AceCloud | 1× NVIDIA H200 NVL | 141 GB HBM3e | 16 vCPU · 128 GB RAM | India and United States | ₹381.46/hour · ₹2,22,775/month | Hourly · monthly · 6-month · 12-month | View provider |
| E2E Networks | 1× NVIDIA H200 | 141 GB HBM3e | 30 vCPU · 375 GB RAM | India | ₹436/hour · ₹2,72,401/month | On-demand · monthly · 6 months · annual · committed pricing | View provider |
| Utho | 1× NVIDIA H200 | 141 GB HBM3e | 16 vCPU · 96 GB RAM | India | From ₹1,79,995/month | Monthly · custom quote · reserved capacity | View provider |
| Cyfuture AI | NVIDIA H200 | 141 GB HBM3e | Custom configuration | India, United States and Europe | Custom quote | Enterprise quote · reserved capacity | View provider |
| Microsoft Azure | ND H200 v5 | 8× 141 GB · 1,128 GB total | 96 vCPU · 1,850 GiB RAM | Selected Azure regions | Use Azure Pricing Calculator | Pay as you go · reservations · savings plan · Spot | View provider |
| DigitalOcean | 1× or 8× NVIDIA HGX H200 | 141 GB HBM3e per GPU | 24 vCPU · 240 GiB RAM per GPU | New York and Atlanta | $4.47/GPU-hour · $3.40 reserved | Per-second on-demand · 12-month reserved | View provider |
| AWS | EC2 P5e / P5en | 8× 141 GB · 1,128 GB total | 8-GPU multi-node training instance | Selected AWS regions, including Mumbai | From $5.97/GPU-hour via Capacity Blocks | On-Demand · Savings Plans · Spot · Capacity Blocks | View provider |
Compare the complete instance rather than only the displayed GPU rate. Providers may offer different H200 variants, GPU counts, vCPU, RAM, storage, network fabric, regions, commitment periods and tax treatment. Multi-GPU instance pricing must not be compared directly with a single-GPU rate, and Spot capacity may be interrupted.
NVIDIA H200 FAQs
How much GPU memory does NVIDIA H200 have?
Both H200 SXM and H200 NVL provide 141 GB of HBM3e memory per GPU.
What is the memory bandwidth of NVIDIA H200?
NVIDIA lists 4.8 TB/s of GPU memory bandwidth for both H200 SXM and H200 NVL.
Is NVIDIA H200 better than H100 for LLM inference?
H200 is usually the stronger option when inference is limited by memory capacity or bandwidth. Its main advantage over H100 is 141 GB HBM3e memory and 4.8 TB/s bandwidth rather than higher raw Tensor Core specifications.
Is NVIDIA H200 suitable for LLM training?
Yes. H200 supports Transformer Engine, FP8 and high-speed NVLink, making it suitable for large-model training and fine-tuning, especially when memory is a constraint.
Does NVIDIA H200 support MIG?
Yes. H200 supports up to seven MIG instances. NVIDIA lists up to seven 18 GB instances for H200 SXM and up to seven 16.5 GB instances for H200 NVL.
What is the difference between H200 SXM and H200 NVL?
Both provide 141 GB HBM3e and 4.8 TB/s bandwidth. H200 SXM is designed for HGX systems and supports up to 700 W, while H200 NVL uses a dual-slot PCIe form factor and supports up to 600 W.
Compare H200 configurations before choosing a provider
Check whether the listing uses H200 SXM or H200 NVL, then compare GPU count, region, interconnect, attached compute, storage, networking and the complete hourly or monthly price.