GPU comparison

H100 vs H200

Specifications and current cloud pricing, side by side.

HoppervsHopperUpdated 17 days ago

Rent the H100 by default. Both cards deliver the same 989 TFLOPS dense FP16 and 1979 TFLOPS dense FP8, so for any model whose weights and KV cache fit in 80 GB the H200 buys you nothing on compute. Switch to the H200 when you need more than 80 GB on one card or when decode throughput is the bottleneck, because it has 141 GB against 80 to 94 GB and 43% more bandwidth than the H100.

H100 from $2.50/GPU/hrH200 from $4.29/GPU/hr

Right now, from live stock

  • Cheapest right now: H100 at $2.50/hr on Hyperstack

    Deploy
  • Most providers in stock: H100 (7)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: H100 ($0.0025 per TFLOPS-hour at 989 TFLOPS)

    Deploy

Specifications Compared

SpecH100H200
TDP700W700W
VRAM80-94 GB141 GB
CUDA Cores16,89616,896
FP8 (dense)1,979 TFLOPS1,979 TFLOPS
Memory TypeHBM3HBM3e
ArchitectureHopperHopper
FP16 (dense)989 TFLOPS989 TFLOPS
Form FactorsSXM5, PCIe, NVLSXM, NVL
INT8 (dense)1,979 TOPS1,979 TOPS
InterconnectNVLink, PCIe 5.0, InfiniBandNVLink, PCIe 5.0, InfiniBand
Tensor Cores528528
FP32 Performance67 TFLOPS67 TFLOPS
FP64 Performance34 TFLOPS34 TFLOPS
Memory Bandwidth3,350 GB/s4,800 GB/s
FP8 (with sparsity)3,958 TFLOPS3,958 TFLOPS
FP16 (with sparsity)1,979 TFLOPS1,979 TFLOPS
INT8 (with sparsity)3,958 TOPS3,958 TOPS

Performance Analysis

The FP16 dense figure of 989 TFLOPS on both GPUs exceeds the FP32 figure of 67 TFLOPS by a substantial margin which supports faster matrix operations during training and inference phases. Memory bandwidth differences appear directly in batch size potential because the H200 bandwidth of 4800 GB/s surpasses the H100 bandwidth of 3350 GB/s and permits larger batches without memory stalls. Dense FP16 performance remains 989 TFLOPS on each device while sparse FP16 performance remains 1979 TFLOPS on each device so any throughput gain stems solely from the H200 memory subsystem rather than arithmetic throughput.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

H100

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
HyperstackCANADA-12$2.50$5.00Deploy
Vast.aiCzechia, CZ1$2.59—Deploy
QuantaCloudus-midwest-22$2.59$5.18Deploy
Massed Computeus-central-31$2.73—Deploy
RunPodglobal1$2.89—Deploy

7 providers in stock, 18 offers (cheapest per provider shown). All H100 offers, price history and alerts

H200

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
LyceumEurope1$4.29—Deploy
RunPodglobal1$4.59—Deploy
Vast.aiCzechia, CZ1$5.57—Deploy

3 providers in stock, 6 offers (cheapest per provider shown). All H200 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when H100 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $2.50/GPU-hr.

QuantaCloud

Comparing H-series providers? We broker across all of them.

Hopper stock changes by the hour and prices differ by provider. If you need 16+ GPUs reserved or a cluster in the next 90 days, we quote H-series or B300 inventory at partner rates: one quote, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the H100

Choose the H100 when your model and its KV cache fit within 80 GB, which covers most fine-tuning and serving jobs. You get the identical 989 TFLOPS dense FP16 and 1979 TFLOPS dense FP8 as the H200, plus the same NVLink and InfiniBand options for multi-GPU work. The H100 is also the only one of the two offered as a PCIe card, which matters if you want a single card instance rather than a full SXM node.

When to Choose the H200

The H200 suits deployments that require 141 GB of HBM3e memory to accommodate larger models or bigger batch sizes. The 4800 GB/s bandwidth supports sustained data movement that exceeds the 3350 GB/s limit of the H100 and therefore benefits memory intensive inference or training sessions that scale beyond 94 GB.

Use Cases

LLM Training
H200

The 141 GB capacity supports larger model replicas than the 80-94 GB limit while the 4800 GB/s bandwidth sustains higher throughput than 3350 GB/s.

LLM Inference
H200

The H200 memory configuration of 141 GB HBM3e accommodates longer context windows that exceed the 94 GB maximum of the H100.

Fine-tuning
H200

Batch sizes scale with the 4800 GB/s bandwidth and 141 GB capacity which surpass the corresponding 3350 GB/s and 80-94 GB figures.

Stable Diffusion
Either

Both GPUs deliver the same 989 TFLOPS dense FP16 and 1979 TFLOPS sparse FP16 so either meets requirements when model size fits in 94 GB.

Scientific Computing
H100

The H100 PCIe form factor option provides deployment flexibility when the 67 TFLOPS FP32 performance suffices and memory needs stay below 94 GB.

Frequently Asked Questions

What is the memory bandwidth difference between these two GPUs?▾

The H200 reaches 4800 GB/s bandwidth. The H100 reaches 3350 GB/s bandwidth. All other specifications including FP8 dense at 1979 TFLOPS remain identical.

Do the H100 and H200 differ in FP16 performance?▾

Both GPUs list FP16 dense performance at 989 TFLOPS and FP16 with sparsity at 1979 TFLOPS. The FP32 performance stands at 67 TFLOPS on each device.

Which GPU handles larger batch sizes during inference?▾

The H200 supports larger batch sizes because its 141 GB capacity and 4800 GB/s bandwidth exceed the 80-94 GB capacity and 3350 GB/s bandwidth of the H100.

What interconnect options exist for the H100 and H200?▾

Both GPUs support NVLink, PCIe 5.0, and InfiniBand. The H100 additionally lists an SXM5 form factor option while the H200 lists an SXM form factor option.

Which is cheaper to rent, the H100 or the H200?▾

Cloud rental prices for both the H100 and H200 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the H100 have compared to the H200?▾

The H100 has 80 to 94 GB of HBM3 memory. The H200 has 141 GB of HBM3e memory.

Can I find H100 and H200 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the H100 and the H200?▾

The H100 uses the Hopper architecture (2022) while the H200 uses Hopper (2024). Both deliver the same dense FP16 throughput (989 TFLOPS without sparsity), and the H200 has 1.4x the memory bandwidth of the H100.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the H100 and the H200. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps