Specifications Compared
| Spec | H100 | H200 |
|---|---|---|
| TDP | 700W | 700W |
| VRAM | 80-94 GB | 141 GB |
| CUDA Cores | 16,896 | 16,896 |
| FP8 (dense) | 1,979 TFLOPS | 1,979 TFLOPS |
| Memory Type | HBM3 | HBM3e |
| Architecture | Hopper | Hopper |
| FP16 (dense) | 989 TFLOPS | 989 TFLOPS |
| Form Factors | SXM5, PCIe, NVL | SXM, NVL |
| INT8 (dense) | 1,979 TOPS | 1,979 TOPS |
| Interconnect | NVLink, PCIe 5.0, InfiniBand | NVLink, PCIe 5.0, InfiniBand |
| Tensor Cores | 528 | 528 |
| FP32 Performance | 67 TFLOPS | 67 TFLOPS |
| FP64 Performance | 34 TFLOPS | 34 TFLOPS |
| Memory Bandwidth | 3,350 GB/s | 4,800 GB/s |
| FP8 (with sparsity) | 3,958 TFLOPS | 3,958 TFLOPS |
| FP16 (with sparsity) | 1,979 TFLOPS | 1,979 TFLOPS |
| INT8 (with sparsity) | 3,958 TOPS | 3,958 TOPS |
Performance Analysis
The FP16 dense figure of 989 TFLOPS on both GPUs exceeds the FP32 figure of 67 TFLOPS by a substantial margin which supports faster matrix operations during training and inference phases. Memory bandwidth differences appear directly in batch size potential because the H200 bandwidth of 4800 GB/s surpasses the H100 bandwidth of 3350 GB/s and permits larger batches without memory stalls. Dense FP16 performance remains 989 TFLOPS on each device while sparse FP16 performance remains 1979 TFLOPS on each device so any throughput gain stems solely from the H200 memory subsystem rather than arithmetic throughput.
Current On-Demand Offers
Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.
H100
| Provider | Region | GPUs | Per GPU / hr | Instance / hr | Deploy |
|---|---|---|---|---|---|
| Hyperstack | CANADA-1 | 2 | $2.50 | $5.00 | Deploy |
| Vast.ai | Czechia, CZ | 1 | $2.59 | — | Deploy |
| QuantaCloud | us-midwest-2 | 2 | $2.59 | $5.18 | Deploy |
| Massed Compute | us-central-3 | 1 | $2.73 | — | Deploy |
| RunPod | global | 1 | $2.89 | — | Deploy |
7 providers in stock, 18 offers (cheapest per provider shown). All H100 offers, price history and alerts
Notify me when H100 drops below a price
One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $2.50/GPU-hr.
QuantaCloud
Comparing H-series providers? We broker across all of them.
Hopper stock changes by the hour and prices differ by provider. If you need 16+ GPUs reserved or a cluster in the next 90 days, we quote H-series or B300 inventory at partner rates: one quote, 24h turnaround.
When to Choose the H100
Choose the H100 when your model and its KV cache fit within 80 GB, which covers most fine-tuning and serving jobs. You get the identical 989 TFLOPS dense FP16 and 1979 TFLOPS dense FP8 as the H200, plus the same NVLink and InfiniBand options for multi-GPU work. The H100 is also the only one of the two offered as a PCIe card, which matters if you want a single card instance rather than a full SXM node.
When to Choose the H200
The H200 suits deployments that require 141 GB of HBM3e memory to accommodate larger models or bigger batch sizes. The 4800 GB/s bandwidth supports sustained data movement that exceeds the 3350 GB/s limit of the H100 and therefore benefits memory intensive inference or training sessions that scale beyond 94 GB.
Use Cases
The 141 GB capacity supports larger model replicas than the 80-94 GB limit while the 4800 GB/s bandwidth sustains higher throughput than 3350 GB/s.
The H200 memory configuration of 141 GB HBM3e accommodates longer context windows that exceed the 94 GB maximum of the H100.
Batch sizes scale with the 4800 GB/s bandwidth and 141 GB capacity which surpass the corresponding 3350 GB/s and 80-94 GB figures.
Both GPUs deliver the same 989 TFLOPS dense FP16 and 1979 TFLOPS sparse FP16 so either meets requirements when model size fits in 94 GB.
The H100 PCIe form factor option provides deployment flexibility when the 67 TFLOPS FP32 performance suffices and memory needs stay below 94 GB.
Frequently Asked Questions
What is the memory bandwidth difference between these two GPUs?▾
The H200 reaches 4800 GB/s bandwidth. The H100 reaches 3350 GB/s bandwidth. All other specifications including FP8 dense at 1979 TFLOPS remain identical.
Do the H100 and H200 differ in FP16 performance?▾
Both GPUs list FP16 dense performance at 989 TFLOPS and FP16 with sparsity at 1979 TFLOPS. The FP32 performance stands at 67 TFLOPS on each device.
Which GPU handles larger batch sizes during inference?▾
The H200 supports larger batch sizes because its 141 GB capacity and 4800 GB/s bandwidth exceed the 80-94 GB capacity and 3350 GB/s bandwidth of the H100.
What interconnect options exist for the H100 and H200?▾
Both GPUs support NVLink, PCIe 5.0, and InfiniBand. The H100 additionally lists an SXM5 form factor option while the H200 lists an SXM form factor option.
Which is cheaper to rent, the H100 or the H200?▾
Cloud rental prices for both the H100 and H200 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.
How much VRAM does the H100 have compared to the H200?▾
The H100 has 80 to 94 GB of HBM3 memory. The H200 has 141 GB of HBM3e memory.
Can I find H100 and H200 GPUs available to rent right now?▾
Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.
What is the main difference between the H100 and the H200?▾
The H100 uses the Hopper architecture (2022) while the H200 uses Hopper (2024). Both deliver the same dense FP16 throughput (989 TFLOPS without sparsity), and the H200 has 1.4x the memory bandwidth of the H100.
Rent these GPUs
Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.
Related comparisons
How this page is made
- Specifications come from the NVIDIA, AMD and Intel datasheets for the H100 and the H200. Dense and sparse throughput are listed separately.
- Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
- The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
- Read how we collect the data, or report an error on this page.
Next steps
- Rent H100Every current offer by provider, daily price history and a price alert.
- Rent H200Every current offer by provider, daily price history and a price alert.
- GPU price indexHow on-demand prices for the major GPUs have moved, updated daily.
- All GPUsEvery GPU we track, with current lows and provider counts.