GPU comparison

A100 vs B200

Specifications and current cloud pricing, side by side.

AmperevsBlackwellUpdated 17 days ago

The B200 wins for the most common use case of LLM training because its 2250 TFLOPS dense FP16 and 8000 GB/s memory bandwidth exceed the corresponding A100 figures by the largest margins among the provided specifications.

A100 from $0.67/GPU/hrB200 from $6.79/GPU/hr

Right now, from live stock

  • Cheapest right now: A100 at $0.67/hr on Vast.ai

    Deploy
  • Most providers in stock: A100 (12)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: A100 ($0.0021 per TFLOPS-hour at 312 TFLOPS)

    Deploy

Specifications Compared

SpecA100B200
TDP400W1000W
VRAM40-80 GB180-192 GB
CUDA Cores6,91218,432
FP4 (dense)Not published9,000 TFLOPS
FP8 (dense)Not published4,500 TFLOPS
Memory TypeHBM2eHBM3e
ArchitectureAmpereBlackwell
FP16 (dense)312 TFLOPS2,250 TFLOPS
Form FactorsSXM4, PCIeSXM, NVL
INT8 (dense)624 TOPS4,500 TOPS
InterconnectNVLink, PCIe 4.0, InfiniBandNVLink, PCIe 6.0, InfiniBand
Tensor Cores432576
FP32 Performance19.5 TFLOPS75 TFLOPS
FP64 Performance9.7 TFLOPS37 TFLOPS
Memory Bandwidth2,039 GB/s8,000 GB/s
FP4 (with sparsity)Not published18,000 TFLOPS
FP8 (with sparsity)Not published9,000 TFLOPS
FP16 (with sparsity)624 TFLOPS4,500 TFLOPS
INT8 (with sparsity)1,248 TOPS9,000 TOPS

Performance Analysis

This difference affects training duration and inference throughput for models that rely on dense FP16 operations. Higher bandwidth supports larger batch sizes during both training and inference without memory constraints.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

A100

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiSlovenia, SI1$0.67—Deploy
LeaderGPUThe Netherlands8$0.68$5.42Deploy
ThunderComputeUSA1$1.09—Deploy
HyperstackCANADA-11$1.35—Deploy
Massed Computeus-central-32$1.35$2.70Deploy

12 providers in stock, 34 offers (cheapest per provider shown). All A100 offers, price history and alerts

B200

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
RunPodglobal1$6.79—Deploy
VERDAFIN-HEL2$7.20$14.40Deploy
Vast.ai, US4$9.38$37.50Deploy

3 providers in stock, 5 offers (cheapest per provider shown). All B200 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when A100 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.67/GPU-hr.

QuantaCloud

Comparing B-series options? Get one quote for all of them.

Skip the per-provider sales calls. Reserved and cluster B-series configurations from 16 to 1024+ GPUs with InfiniBand fabric, 3 to 12 month terms. One quote at partner rates, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the A100

The A100 suits workloads that operate within 40 to 80 GB of HBM2e memory and require 400W TDP. Organizations running established pipelines on Ampere architecture avoid migration costs when the 312 TFLOPS dense FP16 performance meets throughput targets. The 2039 GB/s memory bandwidth remains adequate for moderate batch sizes in inference scenarios.

When to Choose the B200

The B200 suits workloads that require 180 to 192 GB of HBM3e memory and deliver 2250 TFLOPS dense FP16 performance. Organizations scaling to larger models benefit from the 8000 GB/s memory bandwidth that accommodates bigger batch sizes. The 1000W TDP supports sustained operation in dense FP16 and dense FP8 configurations where the 4500 TFLOPS dense FP8 figure accelerates inference.

Use Cases

LLM Training
B200

The B200 provides 2250 TFLOPS dense FP16 and 8000 GB/s memory bandwidth that exceed A100 figures.

LLM Inference
B200

The B200 provides 4500 TFLOPS dense FP8 that exceed A100 dense FP16 figures for throughput.

Fine-tuning
B200

The B200 provides 180 to 192 GB memory capacity that supports larger fine-tuning batches than the A100.

Stable Diffusion
A100

The A100 provides 312 TFLOPS dense FP16 at 400W TDP that meets requirements without excess power draw.

Scientific Computing
A100

The A100 provides 19.5 TFLOPS FP32 at 400W TDP that suits established scientific workloads.

Frequently Asked Questions

How does FP16 dense performance compare between A100 and B200?▾

The A100 delivers 312 TFLOPS dense FP16 while the B200 delivers 2250 TFLOPS dense FP16.

What TDP values apply to A100 and B200?▾

The A100 operates at 400W TDP while the B200 operates at 1000W TDP.

Which GPU offers higher memory bandwidth?▾

The B200 offers 8000 GB/s memory bandwidth while the A100 offers 2039 GB/s memory bandwidth.

How do INT8 dense figures compare?▾

The A100 delivers 624 TOPS dense INT8 while the B200 delivers 4500 TOPS dense INT8.

What architectures do these GPUs use?▾

The A100 uses Ampere architecture while the B200 uses Blackwell architecture.

Which is cheaper to rent, the A100 or the B200?▾

Cloud rental prices for both the A100 and B200 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the A100 have compared to the B200?▾

The A100 has 40 to 80 GB of HBM2e memory. The B200 has 180 to 192 GB of HBM3e memory.

Can I find A100 and B200 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the A100 and the B200?▾

The A100 uses the Ampere architecture (2020) while the B200 uses Blackwell (2024). The B200 delivers 7.2x the dense FP16 throughput (2,250 vs 312 TFLOPS, both without sparsity) and 3.9x the memory bandwidth of the A100.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the A100 and the B200. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps