GPU comparison

B200 vs MI300X

Specifications and current cloud pricing, side by side.

BlackwellvsCDNA 3Updated 18 days ago

The B200 wins for the most common LLM workloads because its 2250 TFLOPS dense FP16 and 8000 GB/s bandwidth exceed the corresponding 1307.4 TFLOPS and 5300 GB/s figures of the MI300X while delivering higher tensor throughput per watt in low precision modes.

B200 from $6.79/GPU/hrMI300X from $2.99/GPU/hr

Right now, from live stock

  • Cheapest right now: MI300X at $2.99/hr on Hot Aisle

    Deploy
  • Most providers in stock: B200 (3)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: MI300X ($0.0023 per TFLOPS-hour at 1,307.4 TFLOPS)

    Deploy

Specifications Compared

SpecB200MI300X
TDP1000W750W
VRAM180-192 GB192 GB
CUDA Cores18,432Not published
FP4 (dense)9,000 TFLOPSNot published
FP8 (dense)4,500 TFLOPS2,614.9 TFLOPS
Memory TypeHBM3eHBM3
ArchitectureBlackwellCDNA 3
FP16 (dense)2,250 TFLOPS1,307.4 TFLOPS
Form FactorsSXM, NVLOAM
INT8 (dense)4,500 TOPS2,614.9 TOPS
InterconnectNVLink, PCIe 6.0, InfiniBandInfinity Fabric, PCIe 5.0
Tensor Cores576Not published
FP32 Performance75 TFLOPS163.4 TFLOPS
FP64 Performance37 TFLOPS81.7 TFLOPS
Memory Bandwidth8,000 GB/s5,300 GB/s
FP4 (with sparsity)18,000 TFLOPSNot published
FP8 (with sparsity)9,000 TFLOPS5,229.8 TFLOPS
FP16 (with sparsity)4,500 TFLOPS2,614.9 TFLOPS
INT8 (with sparsity)9,000 TOPS5,229.8 TOPS

Performance Analysis

Dense FP16 throughput reaches 2250 TFLOPS on the B200 versus 1307.4 TFLOPS on the MI300X so training and inference steps complete in fewer iterations on the B200 when sparsity remains disabled. With sparsity enabled the B200 reaches 4500 TFLOPS FP16 while the MI300X reaches 2614.9 TFLOPS FP16 preserving the same ordering. Memory bandwidth of 8000 GB/s on the B200 versus 5300 GB/s on the MI300X permits larger batch sizes before activation memory saturates during forward and backward passes. FP32 throughput stands at 75 TFLOPS dense on the B200 versus 163.4 TFLOPS on the MI300X indicating the MI300X retains an advantage when workloads stay in full precision.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

B200

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
RunPodglobal1$6.79—Deploy
VERDAFIN-HEL2$7.20$14.40Deploy
Vast.ai, US1$9.38—Deploy

3 providers in stock, 5 offers (cheapest per provider shown). All B200 offers, price history and alerts

MI300X

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Hot AisleMichigan1$2.99—Deploy

1 provider in stock, 2 offers (cheapest per provider shown). All MI300X offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when B200 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $6.79/GPU-hr.

QuantaCloud

Comparing B-series options? Get one quote for all of them.

Skip the per-provider sales calls. Reserved and cluster B-series configurations from 16 to 1024+ GPUs with InfiniBand fabric, 3 to 12 month terms. One quote at partner rates, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the B200

The B200 suits LLM training and inference pipelines that rely on dense FP16 at 2250 TFLOPS or FP8 at 4500 TFLOPS. Its 8000 GB/s bandwidth supports sustained large batch execution where the MI300X bandwidth of 5300 GB/s would constrain throughput.

When to Choose the MI300X

The MI300X suits scientific computing workloads that require 163.4 TFLOPS FP32 and operate within a 750 W power envelope. Its CDNA 3 architecture maintains higher full precision throughput than the 75 TFLOPS FP32 of the B200.

Use Cases

LLM Training
B200

The B200 supplies 2250 TFLOPS dense FP16 compared with 1307.4 TFLOPS on the MI300X allowing faster iteration on large batches.

LLM Inference
B200

The B200 supplies 4500 TFLOPS dense FP8 compared with 2614.9 TFLOPS on the MI300X enabling higher tokens per second at equivalent batch sizes.

Fine-tuning
B200

The B200 supplies 8000 GB/s bandwidth compared with 5300 GB/s on the MI300X supporting larger context windows during gradient updates.

Stable Diffusion
Either

Both GPUs provide 192 GB VRAM so either device accommodates typical diffusion model weights without memory overflow.

Scientific Computing
MI300X

The MI300X supplies 163.4 TFLOPS FP32 compared with 75 TFLOPS on the B200 delivering superior throughput for double precision kernels.

Frequently Asked Questions

How does B200 memory bandwidth compare with MI300X bandwidth?▾

The B200 lists 8000 GB/s while the MI300X lists 5300 GB/s so the B200 sustains higher data movement rates during training loops.

What FP16 dense performance does each GPU deliver?▾

The B200 delivers 2250 TFLOPS dense FP16 while the MI300X delivers 1307.4 TFLOPS dense FP16.

Which GPU offers higher FP32 throughput?▾

The MI300X offers 163.4 TFLOPS FP32 while the B200 offers 75 TFLOPS FP32.

How do the TDP ratings differ between B200 and MI300X?▾

The B200 carries a 1000 W TDP rating while the MI300X carries a 750 W TDP rating.

Does the B200 support sparsity in FP8 mode?▾

The B200 lists 9000 TFLOPS FP8 with sparsity and 4500 TFLOPS FP8 dense.

What interconnect options exist on each card?▾

The B200 supports NVLink and PCIe 6.0 while the MI300X supports Infinity Fabric and PCIe 5.0.

Which is cheaper to rent, the B200 or the MI300X?▾

Cloud rental prices for both the B200 and MI300X vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the B200 have compared to the MI300X?▾

The B200 has 180 to 192 GB of HBM3e memory. The MI300X has 192 GB of HBM3 memory.

Can I find B200 and MI300X GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the B200 and the MI300X?▾

The B200 uses the Blackwell architecture (2024) while the MI300X uses CDNA 3 (2023). The B200 delivers 1.7x the dense FP16 throughput (2,250 vs 1,307.4 TFLOPS, both without sparsity) and 1.5x the memory bandwidth of the MI300X.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the B200 and the MI300X. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps