NVIDIA B200 Cloud GPU
Compare NVIDIA B200 specifications, Blackwell capabilities and workload fit before choosing a cloud GPU provider.

NVIDIA Blackwell data-center GPU
Built for large-scale generative AI
NVIDIA B200 is a Blackwell-generation data-center GPU designed for large-model training, real-time inference and AI factory workloads. In DGX B200, each GPU provides 180 GB of HBM3e memory, 8 TB/s of memory bandwidth and fifth-generation NVLink.
- BlackwellGPU architecture
- 180 GBHBM3e memory per GPU
- 8 TB/sMemory bandwidth per GPU
- 1.8 TB/sNVLink bandwidth per GPU
Who should choose a B200 cloud GPU?
Choose B200 when your training or inference stack can benefit from Blackwell FP4, very high memory bandwidth and faster multi-GPU communication. H100 or H200 may offer better value when your model already fits comfortably and your software cannot use FP4.
NVIDIA B200 specifications
NVIDIA's supplied datasheet describes the eight-GPU DGX B200 system. The per-GPU values below are calculated from the published system totals and are shown beside the complete DGX configuration.
| Specification | B200 per GPU | DGX B200 system |
|---|---|---|
| Architecture | NVIDIA Blackwell | 8 NVIDIA Blackwell GPUs |
| GPU memory | 180 GB HBM3e | 1,440 GB HBM3e total |
| Memory bandwidth | 8 TB/s | 64 TB/s aggregate |
| FP4 Tensor Core | 18 PFLOPS sparse / 9 PFLOPS dense | 144 PFLOPS sparse / 72 PFLOPS dense |
| FP8 Tensor Core | 9 PFLOPS sparse / 4.5 PFLOPS dense | 72 PFLOPS sparse / 36 PFLOPS dense |
| NVLink bandwidth | 1.8 TB/s | 14.4 TB/s aggregate |
| NVSwitch | Connected through fifth-generation NVLink | 2 NVIDIA NVSwitch units |
| Maximum power | Up to 1,000 W per GPU | Approximately 14.3 kW system maximum |
| CPU | Depends on cloud configuration | 2 Intel Xeon Platinum 8570 processors, 112 cores total |
| System memory | Depends on cloud configuration | 2 TB, configurable to 4 TB |
| Networking | Depends on cloud provider | ConnectX-7 and BlueField-3, up to 400 Gb/s |
| Form factor | Provider and platform dependent | 10 RU DGX system |
Per-GPU memory, bandwidth, Tensor Core and NVLink values are derived by dividing NVIDIA's published eight-GPU DGX totals by eight. Dense FP8 is one-half of the published sparse value, following NVIDIA's datasheet note. Cloud instances may expose different CPU, RAM, storage, networking and GPU topology.
Blackwell advantage
Choose B200 for FP4 and higher throughput
B200 adds Blackwell Tensor Cores, second-generation Transformer Engine, FP4 support, 8 TB/s memory bandwidth and 1.8 TB/s NVLink. These features matter most for optimized large-model training and high-volume inference.
Check software readiness
Blackwell value depends on your stack
B200's strongest advantages require software that can use FP4, Blackwell-optimized kernels and the available GPU topology. H100 or H200 can remain more economical for workloads that are not compute constrained or cannot use these features.
What makes B200 different?
Open each capability for a practical explanation of its effect on cloud AI infrastructure.
Second-generation Transformer EngineFP4 acceleration for training and inference
Blackwell's second-generation Transformer Engine combines Blackwell Tensor Cores with fine-grained scaling techniques to support FP4 AI while preserving model accuracy.
180 GB of HBM3e memoryMore room for large models and KV cache
Each B200 GPU in DGX B200 provides 180 GB of HBM3e memory. This can reduce model partitioning and provide more headroom for larger batches, longer contexts and large scientific datasets.
8 TB/s memory bandwidthDesigned for data-intensive AI workloads
The DGX B200 specification publishes 64 TB/s of aggregate HBM3e bandwidth across eight GPUs, equivalent to 8 TB/s per B200 GPU.
Fifth-generation NVLinkFaster GPU-to-GPU communication
Fifth-generation NVLink provides 1.8 TB/s of interconnect bandwidth per GPU in the DGX B200 configuration, helping large distributed training and inference workloads communicate efficiently.
Confidential computingHardware-backed protection for AI workloads
Blackwell includes confidential-computing capabilities designed to protect sensitive data and AI models during training, inference and federated-learning workloads.
Reliability and serviceability engineDesigned for resilient AI infrastructure
Blackwell adds a dedicated Reliability, Availability and Serviceability engine that helps identify potential faults, improve diagnostics and reduce infrastructure downtime.
Best workloads for a B200 cloud GPU
Large-scale LLM inference
A strong fit for high-throughput inference, real-time generative AI and large mixture-of-experts models that benefit from FP4 compute and high memory bandwidth.
Foundation-model training
Designed for distributed pre-training, fine-tuning and reinforcement-learning workloads that need large memory, fast interconnects and high low-precision throughput.
Long-context and high-concurrency AI
The 180 GB memory capacity provides additional room for model weights, KV cache, larger batches and concurrent inference requests.
AI-enabled scientific computing
Useful for large simulations, data-intensive research and mixed AI-HPC pipelines that benefit from high memory capacity and multi-GPU scale.
B200 is a strong fit when…
- Your workload can use FP4 or FP8 to improve training or inference throughput.
- You need more GPU memory than an H100 or H200 configuration provides.
- Your workload benefits from 8 TB/s of HBM3e bandwidth.
- You need fifth-generation NVLink for tightly coupled multi-GPU workloads.
- You are training or serving very large transformer or mixture-of-experts models.
- You expect sustained utilization that can justify a premium Blackwell GPU.
Consider another GPU when…
B200 may be unnecessary when your models fit comfortably on Hopper, your stack cannot use FP4 or your workload does not maintain enough utilization to justify Blackwell pricing.
- Compare H200 when high memory capacity matters but FP4 does not.
- Compare H100 for established enterprise AI training and inference.
- Compare L40S for smaller inference, rendering and visual AI.
Provider comparison
Cloud providers offering NVIDIA B200 or B200-based systems
Compare direct B200 rentals, HGX nodes and rack-scale Blackwell systems by GPU count, memory, attached compute, region, starting price and purchasing model. Azure GB200 is a Grace Blackwell system rather than a standalone B200 instance.
| Provider | B200 offering | GPU memory | Example configuration | Regions | Starting price | Billing options | Details |
|---|---|---|---|---|---|---|---|
| AceCloud | NVIDIA B200 early access | Confirm final configuration | Waitlist and workload-based sizing | India-first; confirm deployment region | Custom quote | Early access · enterprise quote | View provider |
| E2E Networks | 1× NVIDIA B200 | 192 GB HBM3e | 32 vCPU · 400 GB RAM | India | ₹671/hour · ₹4,78,384/month | On-demand · monthly · annual · volume pricing | View provider |
| Utho | 1× to multi-GPU NVIDIA B200 | 192 GB HBM3e per GPU | Custom single-GPU or cluster configuration | India | Custom quote | On-demand · reserved capacity · volume discounts | View provider |
| Cyfuture AI | 1×, 2–4× or 8× HGX B200 | 192 GB HBM3e per GPU | Single GPU to dedicated 8–256+ GPU cluster | India-hosted options | Custom quote | Hourly · monthly · 6-month · annual · custom cluster | View provider |
| Microsoft Azure | ND GB200 v6 | 4× Blackwell GPUs per VM | 2 Grace CPUs · 128 vCPU · 4-GPU VM | Selected Azure regions | Use Azure Pricing Calculator | Pay as you go · reservations · savings plan | View provider |
| AWS | EC2 P6-B200 | 8× B200 · 1,440 GB total HBM3e | 192 vCPU · 2 TiB RAM · 30 TB local NVMe | Selected AWS regions, including Mumbai | $98.84/instance-hour · $12.355/GPU-hour | Capacity Blocks · Savings Plans · check On-Demand access | View provider |
B200 listings are not always equivalent. A provider may offer one GPU, an eight-GPU HGX node, a GB200 Grace Blackwell VM or reserved cluster capacity. Compare actual GPU count, memory per GPU, topology, NVLink access, CPU, RAM, storage, networking, region, availability, commitment term and taxes before comparing the displayed price.
NVIDIA B200 FAQs
How much GPU memory does NVIDIA B200 have?
The DGX B200 datasheet lists 1,440 GB across eight B200 GPUs, which equals 180 GB of HBM3e memory per GPU.
What is the memory bandwidth of NVIDIA B200?
The datasheet lists 64 TB/s across eight GPUs, equivalent to 8 TB/s of HBM3e bandwidth per B200 GPU.
Does NVIDIA B200 support FP4?
Yes. Blackwell introduces FP4 support through its second-generation Transformer Engine and Blackwell Tensor Cores.
Is NVIDIA B200 suitable for LLM training?
Yes. B200 is designed for large distributed training, fine-tuning and inference workloads that benefit from high memory capacity, FP4 or FP8 compute and fifth-generation NVLink.
What is the NVLink bandwidth of NVIDIA B200?
DGX B200 provides 14.4 TB/s of aggregate NVLink bandwidth across eight GPUs, equivalent to 1.8 TB/s per GPU.
How is B200 different from H200?
B200 uses the newer Blackwell architecture and adds FP4 support, higher memory bandwidth and faster NVLink. H200 remains a Hopper GPU with 141 GB HBM3e and 4.8 TB/s bandwidth.
Compare B200 configurations before choosing a provider
Confirm the actual GPU count and topology, then compare memory, NVLink access, CPU, RAM, storage, networking, region and the complete hourly or monthly price. A B200 listing does not automatically mean you receive a full eight-GPU DGX B200 system.