NVIDIA H100

NVIDIA H100

Compare NVIDIA H100 specifications, variants and workload fit before choosing a cloud GPU provider.

NVIDIA H100 cloud GPU

NVIDIA Hopper data-center GPU

Built for large AI and HPC workloads

NVIDIA H100 is a high-end data-center GPU for large language model training, real-time inference, accelerated data analytics and high-performance computing. It combines the Hopper architecture, fourth-generation Tensor Cores, Transformer Engine, high-bandwidth memory and fast multi-GPU interconnects.

  • HopperGPU architecture
  • 80 GBHBM3 on H100 SXM
  • 3.35 TB/sSXM memory bandwidth
  • 900 GB/sSXM NVLink bandwidth

Who should choose an H100 cloud GPU?

Choose H100 when your workload needs high transformer throughput, strong mixed-precision performance, fast memory access or multi-GPU scaling. It is usually more capability than small models, light inference and low-utilization development environments require.

NVIDIA H100 specifications

H100 is available in different server configurations. The table below compares the current H100 SXM and H100 NVL specifications published by NVIDIA.

SpecificationH100 SXMH100 NVL
GPU memory80 GB HBM394 GB HBM3 per GPU
Memory bandwidth3.35 TB/s3.9 TB/s
FP6434 TFLOPS30 TFLOPS
FP64 Tensor Core67 TFLOPS60 TFLOPS
TF32 Tensor Core*989 TFLOPS835 TFLOPS
BF16 / FP16 Tensor Core*1,979 TFLOPS1,671 TFLOPS
FP8 Tensor Core*3,958 TFLOPS3,341 TFLOPS
Maximum TDPUp to 700 W, configurable350–400 W, configurable
MIGUp to 7 instances at 10 GB eachUp to 7 instances at 12 GB each
Form factorSXMDual-slot PCIe, air cooled
InterconnectNVLink 900 GB/s; PCIe Gen5 128 GB/sNVLink 600 GB/s; PCIe Gen5 128 GB/s

* Tensor Core figures are published with sparsity enabled. Cloud configurations vary by provider, server platform and allocated GPU count.

What makes H100 different?

Open each capability for a practical explanation of how it affects cloud workloads.

  • Transformer Engine with FP8Built for transformer training and inference

    The H100 Transformer Engine dynamically uses FP8 and FP16 precision to accelerate transformer workloads while maintaining the accuracy required by large AI models.

  • Fourth-generation Tensor CoresHigh-throughput mixed-precision compute

    H100 supports FP64, TF32, FP32, BF16, FP16, FP8 and INT8 workloads, giving teams one accelerator for AI training, inference and scientific computing.

  • NVLink and NVSwitch scalingDesigned for multi-GPU systems

    H100 SXM provides up to 900 GB/s of NVLink GPU-to-GPU bandwidth. HGX and DGX platforms use NVSwitch to connect multiple H100 GPUs for large distributed workloads.

  • Multi-Instance GPUPartition one GPU into isolated instances

    MIG can divide one H100 into as many as seven hardware-isolated GPU instances. This helps improve utilization when several smaller workloads share the same accelerator.

  • Confidential computingHardware-backed protection for data in use

    Hopper introduced confidential computing capabilities for accelerated workloads, helping protect sensitive applications and data while they are being processed.

  • DPX instructionsAcceleration for dynamic programming

    Dedicated DPX instructions accelerate algorithms such as Smith-Waterman sequence alignment, supporting demanding genomics and scientific-computing workloads.

Best workloads for an H100 cloud GPU

  • Large-model training

    Suitable for transformer pre-training, supervised fine-tuning and distributed training where high Tensor Core throughput and fast GPU interconnects matter.

  • LLM inference

    A strong fit for high-throughput and low-latency inference, particularly for larger models that benefit from FP8 support and high memory bandwidth.

  • HPC and scientific computing

    Designed for simulation, genomics, computational science and other workloads that depend on FP64 performance, memory bandwidth and multi-GPU scaling.

  • Accelerated data analytics

    Useful for large data-processing pipelines built with GPU-accelerated frameworks such as RAPIDS and Spark.

H100 is a strong fit when…

  • Your model or batch sizes need high memory bandwidth.
  • You are training or serving large transformer models.
  • Your workload can use FP8, BF16, FP16 or TensorFloat-32.
  • You need NVLink or NVSwitch for multi-GPU scaling.
  • You expect consistently high GPU utilization.
  • You need MIG for isolated shared-GPU environments.

Consider another GPU when…

H100 may be unnecessary for small models, development environments or cost-sensitive inference that cannot keep the GPU busy.

  • Compare L4 or L40S for lighter inference and visual AI workloads.
  • Compare A100 when Hopper-specific features are not required.
  • Compare H200 when memory capacity and bandwidth are the main constraint.

Provider comparison

Cloud providers offering NVIDIA H100

Compare the listed H100 variant, GPU memory, region, starting price and purchasing options. Prices use each provider's published currency and are not directly comparable without checking the complete instance configuration.

View all providers
ProviderH100 offeringGPU memoryRegionsStarting priceBilling optionsDetails
AceCloud1× NVIDIA H100 HGX80 GB HBM3India and United StatesFrom ₹1,80,000/monthMonthly · 6/12-month plans · hourly availableView provider
E2E Networks1× NVIDIA H100 SXM80 GB HBM3IndiaFrom ₹362/hourOn-demand · monthly · annual · spot subject to capacityView provider
UthoNVIDIA H10080 GB HBM3IndiaCustom quoteOn-demand · reserved capacity · spot to confirmView provider
Cyfuture AI1× NVIDIA H100 SXM80 GB HBM3India, United States and Europe₹329/hour on-demand · ₹219/hour for 12 monthsOn-demand · 1/6/12-month reservedView provider
Microsoft AzureNCads H100 v594 GB H100 NVLSelected global regionsUse Azure Pricing CalculatorPay as you go · reservations · savings plan · SpotView provider
DigitalOcean1× or 8× NVIDIA H10080 GB HBM3 per GPUSelected US, Canada and Europe regions$3.39/GPU-hour on-demand · $3.26 reservedPer-second on-demand · 12-month reservedView provider
AWSEC2 P5 H10080 GB HBM3 per GPUSelected global regions, including MumbaiFrom $4.326/GPU-hour with Capacity BlocksOn-Demand · Savings Plans · Spot · Capacity BlocksView provider

Compare the full instance, not only the displayed GPU rate. Provider prices can use different regions, H100 variants, GPU counts, vCPU, RAM, storage, network capacity, commitment periods and tax treatment. Spot capacity may be interrupted and is not available for every configuration.

NVIDIA H100 FAQs

How much GPU memory does NVIDIA H100 have?

H100 SXM has 80 GB of HBM3 memory. H100 NVL has 94 GB per GPU, or 188 GB across a bridged two-GPU configuration.

Is NVIDIA H100 suitable for LLM training?

Yes. H100 is designed for large AI training workloads and combines fourth-generation Tensor Cores, Transformer Engine, FP8 support and high-speed NVLink.

Can NVIDIA H100 run LLM inference?

Yes. H100 is well suited to large-model inference where throughput, latency and memory bandwidth justify a high-end data-center GPU.

What is the difference between H100 SXM and H100 NVL?

H100 SXM has 80 GB memory, 3.35 TB/s memory bandwidth and up to 900 GB/s NVLink. H100 NVL has 94 GB memory, 3.9 TB/s bandwidth and supports a bridged two-GPU configuration.

Does NVIDIA H100 support MIG?

Yes. H100 supports up to seven Multi-Instance GPU partitions, allowing several isolated workloads to share one physical GPU.

How much power does NVIDIA H100 use?

H100 SXM has a configurable TDP of up to 700 W. H100 NVL is configurable from 350 W to 400 W.

Compare H100 before choosing a provider

Check whether the listing uses H100 SXM, H100 PCIe or H100 NVL, then compare memory, interconnect, GPU count, region, attached compute and the complete hourly or monthly price.