NVIDIA B200 Cloud GPU

NVIDIA B200 Cloud GPU

Compare NVIDIA B200 specifications, Blackwell capabilities and workload fit before choosing a cloud GPU provider.

NVIDIA B200 cloud GPU

NVIDIA Blackwell data-center GPU

Built for large-scale generative AI

NVIDIA B200 is a Blackwell-generation data-center GPU designed for large-model training, real-time inference and AI factory workloads. In DGX B200, each GPU provides 180 GB of HBM3e memory, 8 TB/s of memory bandwidth and fifth-generation NVLink.

  • BlackwellGPU architecture
  • 180 GBHBM3e memory per GPU
  • 8 TB/sMemory bandwidth per GPU
  • 1.8 TB/sNVLink bandwidth per GPU

Who should choose a B200 cloud GPU?

Choose B200 when your training or inference stack can benefit from Blackwell FP4, very high memory bandwidth and faster multi-GPU communication. H100 or H200 may offer better value when your model already fits comfortably and your software cannot use FP4.

NVIDIA B200 specifications

NVIDIA's supplied datasheet describes the eight-GPU DGX B200 system. The per-GPU values below are calculated from the published system totals and are shown beside the complete DGX configuration.

SpecificationB200 per GPUDGX B200 system
ArchitectureNVIDIA Blackwell8 NVIDIA Blackwell GPUs
GPU memory180 GB HBM3e1,440 GB HBM3e total
Memory bandwidth8 TB/s64 TB/s aggregate
FP4 Tensor Core18 PFLOPS sparse / 9 PFLOPS dense144 PFLOPS sparse / 72 PFLOPS dense
FP8 Tensor Core9 PFLOPS sparse / 4.5 PFLOPS dense72 PFLOPS sparse / 36 PFLOPS dense
NVLink bandwidth1.8 TB/s14.4 TB/s aggregate
NVSwitchConnected through fifth-generation NVLink2 NVIDIA NVSwitch units
Maximum powerUp to 1,000 W per GPUApproximately 14.3 kW system maximum
CPUDepends on cloud configuration2 Intel Xeon Platinum 8570 processors, 112 cores total
System memoryDepends on cloud configuration2 TB, configurable to 4 TB
NetworkingDepends on cloud providerConnectX-7 and BlueField-3, up to 400 Gb/s
Form factorProvider and platform dependent10 RU DGX system

Per-GPU memory, bandwidth, Tensor Core and NVLink values are derived by dividing NVIDIA's published eight-GPU DGX totals by eight. Dense FP8 is one-half of the published sparse value, following NVIDIA's datasheet note. Cloud instances may expose different CPU, RAM, storage, networking and GPU topology.

Blackwell advantage

Choose B200 for FP4 and higher throughput

B200 adds Blackwell Tensor Cores, second-generation Transformer Engine, FP4 support, 8 TB/s memory bandwidth and 1.8 TB/s NVLink. These features matter most for optimized large-model training and high-volume inference.

Check software readiness

Blackwell value depends on your stack

B200's strongest advantages require software that can use FP4, Blackwell-optimized kernels and the available GPU topology. H100 or H200 can remain more economical for workloads that are not compute constrained or cannot use these features.

What makes B200 different?

Open each capability for a practical explanation of its effect on cloud AI infrastructure.

  • Second-generation Transformer EngineFP4 acceleration for training and inference

    Blackwell's second-generation Transformer Engine combines Blackwell Tensor Cores with fine-grained scaling techniques to support FP4 AI while preserving model accuracy.

  • 180 GB of HBM3e memoryMore room for large models and KV cache

    Each B200 GPU in DGX B200 provides 180 GB of HBM3e memory. This can reduce model partitioning and provide more headroom for larger batches, longer contexts and large scientific datasets.

  • 8 TB/s memory bandwidthDesigned for data-intensive AI workloads

    The DGX B200 specification publishes 64 TB/s of aggregate HBM3e bandwidth across eight GPUs, equivalent to 8 TB/s per B200 GPU.

  • Fifth-generation NVLinkFaster GPU-to-GPU communication

    Fifth-generation NVLink provides 1.8 TB/s of interconnect bandwidth per GPU in the DGX B200 configuration, helping large distributed training and inference workloads communicate efficiently.

  • Confidential computingHardware-backed protection for AI workloads

    Blackwell includes confidential-computing capabilities designed to protect sensitive data and AI models during training, inference and federated-learning workloads.

  • Reliability and serviceability engineDesigned for resilient AI infrastructure

    Blackwell adds a dedicated Reliability, Availability and Serviceability engine that helps identify potential faults, improve diagnostics and reduce infrastructure downtime.

Best workloads for a B200 cloud GPU

  • Large-scale LLM inference

    A strong fit for high-throughput inference, real-time generative AI and large mixture-of-experts models that benefit from FP4 compute and high memory bandwidth.

  • Foundation-model training

    Designed for distributed pre-training, fine-tuning and reinforcement-learning workloads that need large memory, fast interconnects and high low-precision throughput.

  • Long-context and high-concurrency AI

    The 180 GB memory capacity provides additional room for model weights, KV cache, larger batches and concurrent inference requests.

  • AI-enabled scientific computing

    Useful for large simulations, data-intensive research and mixed AI-HPC pipelines that benefit from high memory capacity and multi-GPU scale.

B200 is a strong fit when…

  • Your workload can use FP4 or FP8 to improve training or inference throughput.
  • You need more GPU memory than an H100 or H200 configuration provides.
  • Your workload benefits from 8 TB/s of HBM3e bandwidth.
  • You need fifth-generation NVLink for tightly coupled multi-GPU workloads.
  • You are training or serving very large transformer or mixture-of-experts models.
  • You expect sustained utilization that can justify a premium Blackwell GPU.

Consider another GPU when…

B200 may be unnecessary when your models fit comfortably on Hopper, your stack cannot use FP4 or your workload does not maintain enough utilization to justify Blackwell pricing.

  • Compare H200 when high memory capacity matters but FP4 does not.
  • Compare H100 for established enterprise AI training and inference.
  • Compare L40S for smaller inference, rendering and visual AI.

Provider comparison

Cloud providers offering NVIDIA B200 or B200-based systems

Compare direct B200 rentals, HGX nodes and rack-scale Blackwell systems by GPU count, memory, attached compute, region, starting price and purchasing model. Azure GB200 is a Grace Blackwell system rather than a standalone B200 instance.

View all providers
ProviderB200 offeringGPU memoryExample configurationRegionsStarting priceBilling optionsDetails
AceCloudNVIDIA B200 early accessConfirm final configurationWaitlist and workload-based sizingIndia-first; confirm deployment regionCustom quoteEarly access · enterprise quoteView provider
E2E Networks1× NVIDIA B200192 GB HBM3e32 vCPU · 400 GB RAMIndia₹671/hour · ₹4,78,384/monthOn-demand · monthly · annual · volume pricingView provider
Utho1× to multi-GPU NVIDIA B200192 GB HBM3e per GPUCustom single-GPU or cluster configurationIndiaCustom quoteOn-demand · reserved capacity · volume discountsView provider
Cyfuture AI1×, 2–4× or 8× HGX B200192 GB HBM3e per GPUSingle GPU to dedicated 8–256+ GPU clusterIndia-hosted optionsCustom quoteHourly · monthly · 6-month · annual · custom clusterView provider
Microsoft AzureND GB200 v64× Blackwell GPUs per VM2 Grace CPUs · 128 vCPU · 4-GPU VMSelected Azure regionsUse Azure Pricing CalculatorPay as you go · reservations · savings planView provider
AWSEC2 P6-B2008× B200 · 1,440 GB total HBM3e192 vCPU · 2 TiB RAM · 30 TB local NVMeSelected AWS regions, including Mumbai$98.84/instance-hour · $12.355/GPU-hourCapacity Blocks · Savings Plans · check On-Demand accessView provider

B200 listings are not always equivalent. A provider may offer one GPU, an eight-GPU HGX node, a GB200 Grace Blackwell VM or reserved cluster capacity. Compare actual GPU count, memory per GPU, topology, NVLink access, CPU, RAM, storage, networking, region, availability, commitment term and taxes before comparing the displayed price.

NVIDIA B200 FAQs

How much GPU memory does NVIDIA B200 have?

The DGX B200 datasheet lists 1,440 GB across eight B200 GPUs, which equals 180 GB of HBM3e memory per GPU.

What is the memory bandwidth of NVIDIA B200?

The datasheet lists 64 TB/s across eight GPUs, equivalent to 8 TB/s of HBM3e bandwidth per B200 GPU.

Does NVIDIA B200 support FP4?

Yes. Blackwell introduces FP4 support through its second-generation Transformer Engine and Blackwell Tensor Cores.

Is NVIDIA B200 suitable for LLM training?

Yes. B200 is designed for large distributed training, fine-tuning and inference workloads that benefit from high memory capacity, FP4 or FP8 compute and fifth-generation NVLink.

What is the NVLink bandwidth of NVIDIA B200?

DGX B200 provides 14.4 TB/s of aggregate NVLink bandwidth across eight GPUs, equivalent to 1.8 TB/s per GPU.

How is B200 different from H200?

B200 uses the newer Blackwell architecture and adds FP4 support, higher memory bandwidth and faster NVLink. H200 remains a Hopper GPU with 141 GB HBM3e and 4.8 TB/s bandwidth.

Compare B200 configurations before choosing a provider

Confirm the actual GPU count and topology, then compare memory, NVLink access, CPU, RAM, storage, networking, region and the complete hourly or monthly price. A B200 listing does not automatically mean you receive a full eight-GPU DGX B200 system.