NVIDIA B200
Reserve Blackwell GPU clusters for large-scale training and inference. Plan GPU count, networking and capacity with our team.
Reserved capacity · Contact our team for a quote
Workloads
What to run on B200.
Choose for your workload’s memory needs, software support and measured runtime.
Large training jobs
Plan a cluster around your model, parallelism strategy and training schedule.
Low-precision inference
Blackwell supports FP4. Benefits depend on your model, quantization method and serving software.
Reserved capacity
Discuss deployments of 128, 256, 1,024 GPUs and beyond. Configuration and delivery depend on your requirements and available capacity.
Choose your setup
Plan your reserved cluster.
Match capacity to your model, schedule and networking requirements.
Choose your scale
Discuss 128, 256, 1,024 GPUs or a larger configuration for your workload.
Plan the network
B200 reserved clusters use NVIDIA reference architecture and InfiniBand. Confirm the exact hardware configuration and schedule with our team.
Agree on capacity
Confirm GPU configuration, reservation term and delivery schedule before committing.
Plan a reservation↗Sizing
Start with memory. Then measure performance.
Model weights are only part of the working set. Leave room for everything the job needs while it runs.
For inference
Account for model precision, context length, KV cache and concurrent requests. A model loading successfully does not tell you how much traffic it can serve.
For training
Include activations, gradients and optimizer state. Batch size, sequence length and checkpointing change memory use.
For multiple GPUs
GPU memory is not automatically pooled. Use a framework and parallelism strategy that distribute the workload across devices.
¹ Memory shown uses the NVIDIA DGX B200 reference: 1,440 GB across eight GPUs. Confirm the deployed configuration in your quote. Hardware specifications: NVIDIA B200. Software and workload affect realized performance.
Compare options
Compare memory and hourly rates.
Use this as a shortlist, then test your workload. Lower hourly pricing does not always mean a lower total job cost.
Published USD on-demand rates per GPU; storage extra. Availability and regional prices vary.
- Memory
- 141 GB HBM3e
- Architecture
- Hopper
- From / GPU / hour
- $3.99
- Memory
- 80 GB HBM3
- Architecture
- Hopper
- From / GPU / hour
- $2.69
- Memory
- 180 GB per GPU¹
- Architecture
- Blackwell
- From / GPU / hour
- Request a quote
- Memory
- 96 GB GDDR7
- Architecture
- Blackwell
- From / GPU / hour
- $1.89
- Memory
- 80 GB HBM2e
- Architecture
- Ampere
- From / GPU / hour
- $1.49
- Memory
- 40 GB HBM2
- Architecture
- Ampere
- From / GPU / hour
- $0.89
- Memory
- 24 GB GDDR6
- Architecture
- Ada Lovelace
- From / GPU / hour
- $0.44
| GPU | Memory | Architecture | From / GPU / hour |
|---|---|---|---|
| NVIDIA H200 | 141 GB HBM3e | Hopper | $3.99 |
| NVIDIA H100 | 80 GB HBM3 | Hopper | $2.69 |
| NVIDIA B200 | 180 GB per GPU¹ | Blackwell | Request a quote |
| NVIDIA RTX PRO 6000 | 96 GB GDDR7 | Blackwell | $1.89 |
| NVIDIA A100 80 GB | 80 GB HBM2e | Ampere | $1.49 |
| NVIDIA A100 40 GB | 40 GB HBM2 | Ampere | $0.89 |
| NVIDIA L4 | 24 GB GDDR6 | Ada Lovelace | $0.44 |
Before you start
Common questions.
Practical details for choosing and using this product.
How do I get started with NVIDIA B200?
Visit the reserved clusters section and contact our team to discuss GPU count, network requirements, timing and pricing.
Is storage included in the GPU rate?
Storage is billed separately. Retained storage continues to incur charges while an instance is paused. Review the full configuration price before launching.
Will my model fit on one GPU?
This GPU has 180 GB per GPU¹. Fit depends on weights, precision, framework overhead and workload state. For inference, also account for context and concurrency; for training, include activations and optimizer state.
Can I get more than eight GPUs?
For workloads spanning nodes, explore GPU clusters. Reserved capacity can be planned at 128, 256, 1,024 GPUs and beyond, subject to configuration and availability.
Get started with NVIDIA B200.
Tell us what you’re building. We’ll help you choose the GPU count, configuration and reservation term.