Blog/GPU Guides

NVIDIA H100 Price Guide 2026: GPU Costs, Cloud Pricing & Buy vs Rent

Team Jarvislabs
Team Jarvislabs

Jarvislabs

March 23, 2026·10 min read·Updated September 8, 2026
NVIDIA H100 Price Guide 2026: GPU Costs, Cloud Pricing & Buy vs Rent

Pricing checked September 8, 2026. USD rates; availability and additional charges depend on the configuration.

An NVIDIA H100 SXM with 80 GB of GPU memory costs $2.69 per GPU-hour to rent on Jarvislabs, billed by the minute. Storage is extra. Buying hardware requires a current supplier quote for the exact GPU or complete server; a card price alone does not cover a working system.

This guide separates three decisions: what H100 compute costs by the hour, how to compare cloud configurations, and when buying might make sense for your workload.

  • H100 rental: $2.69 per GPU-hour for Jarvislabs H100 SXM on-demand compute.
  • 100 GPU-hours: $269 in compute charges.
  • A 30-day month running continuously: $1,936.80 per GPU, before storage and applicable taxes.
  • Purchase comparison: use the complete installed system cost, expected utilization and ongoing operating costs.

Check H100 availability and launch →

1. How Much Does an NVIDIA H100 GPU Cost to Buy?

There is no single purchase price that covers every H100 configuration. Request a dated quote that specifies the variant, GPU count, new or used condition, warranty, host system and interconnects. A standalone PCIe card, an SXM-based HGX server and a complete DGX system are different purchases.

For the worked example below, we use $30,000 per GPU of installed capital cost as a budgeting assumption, not a current market quote. Replace it with your supplier's complete system quote divided by the GPU count.

Include the costs of the host CPUs and RAM, storage, network adapters and switches, racks, power delivery, cooling, support and installation. For multi-node training, confirm the network topology and measured communication performance before comparing prices.

2. H100 Cloud GPU Pricing: Hourly Rates Compared

H100 price per hour, day and month on Jarvislabs

All figures below use $2.69 per H100 SXM GPU-hour. They cover on-demand GPU compute only, excluding storage and applicable taxes. The monthly example assumes 30 days × 24 hours = 720 hours; it is not a reserved-capacity quote.

Running time1 H100 SXM8 H100 SXM GPUs
1 hour$2.69$21.52
10 hours$26.90$215.20
24 hours$64.56$516.48
100 hours$269.00$2,152.00
720 hours (30 days)$1,936.80$15,494.40

Compute cost = GPU count × running hours × $2.69. Retained storage continues to incur charges while an instance is paused. Check the full pricing details and the configuration shown in the dashboard before launching.

Check H100 availability and launch →

Compare the same GPU and service type

The public pricing pages checked on September 8, 2026 show these rates. This is a snapshot of the listed configurations, not a guarantee of available capacity or the cheapest rate in every region.

Provider and sourceListed configurationUSD per GPU-hourWhat to check
JarvislabsH100 SXM, 80 GB, on-demand$2.69Per-minute GPU billing; storage extra; select an available region
RunpodH100 SXM, 80 GB, Pods listing$3.49Cloud selection, region, storage and live availability
Lambda1× H100 SXM instance, 80 GB$4.29Single-GPU instance rate; included host resources and applicable taxes

Lambda also lists different per-GPU rates for multi-GPU instances. Runpod lists PCIe and NVL variants separately. Do not compare an SXM instance against a cheaper PCIe listing without accounting for the hardware difference. Serverless worker prices and managed inference services have different billing and operating models from GPU instances.

For a fair quote, record the GPU variant, GPU count, region, commitment, host resources, storage charges and billing unit. A USD price does not imply that the GPU is physically located in the United States; confirm the deployment region if data location or latency matters.

3. H100 Cost Considerations: What to Know Before Renting

The hourly rate is only one part of job cost. Measure the complete run: environment setup, downloading model weights, model loading, warm-up, useful work and time left idle before shutdown.

For inference, specify the model and precision, input and output lengths, concurrency, cache behavior, and latency target. For training, specify the dataset, sequence length, batch size, optimizer and number of steps. These determine whether the job fits and how long it runs.

Real-World Cost Benchmarks

The following are cost calculations for assumed runtimes, not measured performance benchmarks. They do not predict how long a particular model takes to train or serve.

Example allocationCalculationCompute cost
One-GPU experiment lasting 4 hours1 × 4 × $2.69$10.76
Four-GPU fine-tuning run lasting 15 hours4 × 15 × $2.69$161.40
Eight-GPU job lasting 840 hours8 × 840 × $2.69$18,076.80
One GPU serving for 3 hours/day over 30 days1 × 3 × 30 × $2.69$242.10

For measured serving results, see our vLLM, SGLang and TensorRT-LLM H100 comparison. Use the tested model, workload and configuration when interpreting its numbers. A throughput measurement from one setup is not a universal tokens-per-second guarantee.

H100 Rental vs Purchase Calculator

Start with a simple capital-cost comparison:

Rental hours to match capital cost = installed purchase cost ÷ equivalent hourly rental cost.

For an assumed $30,000 installed cost per GPU, divided by $2.69/hour, the result is approximately 11,152 running hours. At 720 running hours per month, that is 15.5 months; at 100 hours per month, it is 111.5 months. This calculation excludes ongoing ownership costs, cloud extras, financing and resale value, so it is not a complete ownership break-even forecast.

For an eight-GPU system assumed to cost $250,000 installed, equivalent GPU rental is 8 × $2.69 = $21.52 per cluster-hour. The capital-only crossing is $250,000 ÷ $21.52 ≈ 11,617 cluster-hours, or 16.1 months at 720 hours/month. Keep cluster-hours and GPU-hours separate: one hour on eight GPUs is eight GPU-hours.

A fuller comparison over your planning period is:

  • Own: installed capital cost + power + cooling + maintenance + operating staff + financing − resale value.
  • Rent: GPU runtime charges + retained storage + applicable network and other service charges.

Use the same workload, capacity and time period on both sides. Low utilization can make ownership expensive even when the system's theoretical throughput is high. Continuous demand can make buying worth investigating, but only after including operating costs and the risk of hardware becoming a poor fit.

4. H100 Alternatives: A100, H200 & Other Options

Choose hardware by memory requirements and measured cost per completed job, not just hourly price.

  • A100: consider it when it fits your workload and a lower hourly rate matters. Compare actual runtime in our H100 vs A100 guide.
  • H200: consider its larger memory for workloads constrained by model weights or KV cache. See the H200 pricing guide. Our H100 vs H200 benchmark shows measured throughput and cost per token for both.
  • RTX PRO 6000: investigate it for workloads suited to its memory capacity and software support. See available configurations.

Benchmark a representative job before reserving a large allocation. A faster GPU only saves money if its runtime improvement offsets its higher rate for your workload.

5. H100 PCIe vs SXM: Which Version Should You Choose?

The $2.69 rate in this guide is for H100 SXM. Confirm the variant and server topology when comparing another offer. SXM and PCIe form factors require compatible host systems; the server's cooling design and GPU interconnects also matter.

NVIDIA H100 SXM Specifications at a Glance

SpecificationH100 SXM
ArchitectureNVIDIA Hopper
GPU memory80 GB HBM3
Memory bandwidth3.35 TB/s
Maximum TDPUp to 700 W, configurable
FP3267 TFLOPS
TF32 Tensor Core, with sparsity989 TFLOPS
FP16 Tensor Core, with sparsity1,979 TFLOPS
FP8 Tensor Core, with sparsity3,958 TFLOPS
NVLink bandwidth900 GB/s

Source: NVIDIA H100 specifications. Theoretical Tensor Core performance with sparsity is different from ordinary FP32 throughput and from application performance. GPU memory is not automatically pooled across devices; the serving or training framework must distribute the workload.

A price history needs dated, comparable quotes. A change in GPU variant, cloud region, commitment or service type can look like a price drop even when the equivalent configuration has not become cheaper.

For planning, use the September 8 snapshot above and obtain fresh quotes before committing. Avoid assuming a fixed percentage price decline or a guaranteed price floor. Availability, competing hardware and your required deployment region can change the decision.

Frequently Asked Questions (FAQ)

How much does the NVIDIA H100 GPU cost?

Jarvislabs H100 SXM rental is $2.69 per GPU-hour, with per-minute billing and storage charged separately. Hardware purchase pricing requires a supplier quote for the exact card or installed system. The $30,000 installed-cost figure used above is an example assumption, not an advertised purchase price.

How much does an H100 cost per month?

At $2.69 per GPU-hour, one H100 SXM running for a 30-day month costs $1,936.80 in compute charges. Running for 100 hours costs $269. Storage and applicable taxes are additional; a reserved-capacity offer may use different terms.

What is the cheapest way to rent an H100 GPU?

Compare current quotes for equivalent configurations. Jarvislabs lists H100 SXM at $2.69 per GPU-hour on demand. Spot capacity may cost less but has different availability and interruption behavior. Review the current pricing options and stop unused compute; per-minute billing does not eliminate charges for a running but idle GPU.

What's the break-even point for buying vs renting H100?

With an assumed $30,000 installed cost and $2.69/hour rental, the capital-only crossing is about 11,152 running hours, or 15.5 months of continuous use at 720 hours/month. Include power, cooling, maintenance, financing, resale value and cloud extras before making a purchase decision.

How much does it cost to train or run LLaMA on H100?

Multiply measured running hours by GPU count and $2.69. Model size alone cannot establish a training budget or serving throughput. Precision, dataset size, context length, concurrency and latency requirements all matter. Use the cost examples as arithmetic templates and benchmark your configuration.

How much VRAM does the H100 have?

The H100 SXM offered here has 80 GB HBM3. Other H100 variants can differ. Leave room beyond model weights for framework overhead, activations or KV cache. A model fitting in memory does not establish how much traffic it can serve.

What is the power consumption of the NVIDIA H100 GPU?

NVIDIA specifies up to 700 W configurable TDP for H100 SXM. A complete server also consumes power for its CPUs, memory, storage, networking and cooling. TDP is not a measurement of your job's average power use.

Is H100 worth it over A100?

It depends on runtime and memory requirements. Compare cost per completed job or per million tokens at your latency target. See our H100 vs A100 comparison, then test a representative workload.

Can US customers rent an H100 on Jarvislabs?

Customers outside India are billed in USD. Review the available region, configuration and total price in the dashboard before launching. The billing currency does not specify the physical data location. Check H100 availability and launch →

Get Started

Build & Deploy Your AI in Minutes

Cloud GPU infrastructure designed specifically for AI development. Start training and deploying models today.

View Pricing