NVIDIA H200
Rent H200 GPUs on demand for LLM training and inference with 141 GB of memory per GPU. Choose a template, GPU VM or cluster for your workload.
From $3.99/GPU/hour · USD on-demand rate · Storage extra
H200 rental
Choose how you rent H200.
Start on demand, use interruptible capacity for restartable jobs, or plan a longer reservation with our team.
USD rates per GPU-hour. Storage is billed separately; regional prices and availability vary.
- On demand
- GPU rate
- $3.99/hr
- When to choose it
- Pay by the minute without a long-term commitment. Choose your GPU count and region in the dashboard.
- Spot
- GPU rate
- $1.99/hr
- When to choose it
- Lower-cost, interruptible capacity. Save checkpoints and use jobs that can restart; subject to availability.
- Reserved capacity
- GPU rate
- Discuss your requirements
- When to choose it
- Plan GPU count, reservation term and networking with our team for sustained or multi-node workloads.
| Option | GPU rate | When to choose it |
|---|---|---|
| On demand | $3.99/hr | Pay by the minute without a long-term commitment. Choose your GPU count and region in the dashboard. |
| Spot | $1.99/hr | Lower-cost, interruptible capacity. Save checkpoints and use jobs that can restart; subject to availability. |
| Reserved capacity | Discuss your requirements | Plan GPU count, reservation term and networking with our team for sustained or multi-node workloads. |
Compare the full cost
Review H200 rental options, ownership costs and the tradeoffs against H100 before choosing your setup.
Read the H200 pricing guide↗Plan multi-GPU inference
Understand tensor, pipeline and data parallelism before distributing a model across GPUs.
Read the inference scaling guide↗Reserve a cluster
Share your workload and schedule to confirm the right configuration and capacity.
Discuss reserved H200 capacity↗Workloads
What to run on H200.
Choose for your workload’s memory needs, software support and measured runtime.
Larger model working sets
Keep more weights, activations and inference cache on each GPU. Size your job for its precision, context length and batch size.
Training and fine-tuning
Use the additional memory for larger batches or longer sequences when your training framework supports them.
Distributed workloads
Use an H200 cluster when the job needs multiple nodes. Plan model parallelism and data movement alongside GPU count.
Choose your setup
Start with the environment you need.
Use a preconfigured container or manage your operating system in a GPU VM. The dashboard shows current GPU and region availability.
GPU Templates
Start with PyTorch, ComfyUI or another supported environment. Drivers and framework dependencies are preconfigured.
Explore templates↗GPU VMs
Get SSH and full root access when you need control over the operating system and runtime. Available configurations vary by GPU and region.
Explore VMs↗GPU clusters
Instant and reserved H200 clusters are available. Check the dashboard for current instant capacity.
Explore clusters↗Sizing
Start with memory. Then measure performance.
Model weights are only part of the working set. Leave room for everything the job needs while it runs.
For inference
Account for model precision, context length, KV cache and concurrent requests. A model loading successfully does not tell you how much traffic it can serve.
For training
Include activations, gradients and optimizer state. Batch size, sequence length and checkpointing change memory use.
For multiple GPUs
GPU memory is not automatically pooled. Use a framework and parallelism strategy that distribute the workload across devices.
Hardware specifications: NVIDIA H200. Software and workload affect realized performance.
Compare options
Compare memory and hourly rates.
Use this as a shortlist, then test your workload. Lower hourly pricing does not always mean a lower total job cost.
Published USD on-demand rates per GPU; storage extra. Availability and regional prices vary.
- Memory
- 141 GB HBM3e
- Architecture
- Hopper
- From / GPU / hour
- $3.99
- Memory
- 80 GB HBM3
- Architecture
- Hopper
- From / GPU / hour
- $2.69
- Memory
- 180 GB per GPU¹
- Architecture
- Blackwell
- From / GPU / hour
- Request a quote
- Memory
- 96 GB GDDR7
- Architecture
- Blackwell
- From / GPU / hour
- $1.89
- Memory
- 80 GB HBM2e
- Architecture
- Ampere
- From / GPU / hour
- $1.49
- Memory
- 40 GB HBM2
- Architecture
- Ampere
- From / GPU / hour
- $0.89
- Memory
- 24 GB GDDR6
- Architecture
- Ada Lovelace
- From / GPU / hour
- $0.44
| GPU | Memory | Architecture | From / GPU / hour |
|---|---|---|---|
| NVIDIA H200 | 141 GB HBM3e | Hopper | $3.99 |
| NVIDIA H100 | 80 GB HBM3 | Hopper | $2.69 |
| NVIDIA B200 | 180 GB per GPU¹ | Blackwell | Request a quote |
| NVIDIA RTX PRO 6000 | 96 GB GDDR7 | Blackwell | $1.89 |
| NVIDIA A100 80 GB | 80 GB HBM2e | Ampere | $1.49 |
| NVIDIA A100 40 GB | 40 GB HBM2 | Ampere | $0.89 |
| NVIDIA L4 | 24 GB GDDR6 | Ada Lovelace | $0.44 |
Same model, same serving stack, both GPUs: Read the measured H100 vs H200 comparison
Before you start
Common questions.
Practical details for choosing and using this product.
How do I get started with NVIDIA H200?
Choose a template or VM, select the GPU and region, and review the configuration and price in the dashboard before launch.
Is storage included in the GPU rate?
Storage is billed separately. Retained storage continues to incur charges while an instance is paused. Review the full configuration price before launching.
Will my model fit on one GPU?
This GPU has 141 GB HBM3e. Fit depends on weights, precision, framework overhead and workload state. For inference, also account for context and concurrency; for training, include activations and optimizer state.
Can I get more than eight GPUs?
For workloads spanning nodes, explore GPU clusters. Reserved capacity can be planned at 128, 256, 1,024 GPUs and beyond, subject to configuration and availability.
Get started with NVIDIA H200.
Launch a GPU instance or talk to our team about the right setup for your workload.