GPU comparison

RTX 4090 vs A100

Specifications and current cloud pricing, side by side.

Ada LovelacevsAmpereUpdated 17 days ago

The A100 wins for the most common use case of large scale training and inference because dense FP16 performance reaches 312 TFLOPS with 40 to 80 GB HBM2e VRAM and 2039 GB/s bandwidth. The RTX 4090 remains viable only when FP32 at 82.6 TFLOPS or consumer form factor constraints dominate.

RTX 4090 from $0.51/GPU/hrA100 from $0.68/GPU/hr

Right now, from live stock

  • Cheapest right now: RTX 4090 at $0.51/hr on Vast.ai

    Deploy
  • Most providers in stock: A100 (12)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: A100 ($0.0022 per TFLOPS-hour at 312 TFLOPS)

    Deploy

Specifications Compared

SpecRTX 4090A100
TDP450W400W
VRAM24 GB40-80 GB
CUDA Cores16,3846,912
FP8 (dense)330.3 TFLOPSNot published
Memory TypeGDDR6XHBM2e
ArchitectureAda LovelaceAmpere
FP16 (dense)165.2 TFLOPS312 TFLOPS
Form FactorsPCIeSXM4, PCIe
INT8 (dense)660.6 TOPS624 TOPS
InterconnectPCIe 4.0NVLink, PCIe 4.0, InfiniBand
Tensor Cores512432
FP32 Performance82.6 TFLOPS19.5 TFLOPS
FP64 Performance1.3 TFLOPS9.7 TFLOPS
Memory Bandwidth1,008 GB/s2,039 GB/s
FP8 (with sparsity)660.6 TFLOPSNot published
FP16 (with sparsity)330.4 TFLOPS624 TFLOPS
INT8 (with sparsity)1,321.2 TOPS1,248 TOPS

Performance Analysis

Sparse FP16 performance measures 330.4 TFLOPS on the RTX 4090 against 624 TFLOPS on the A100 preserving the same relative advantage for the A100. The FP16 to FP32 ratio favors the RTX 4090 at 165.2 TFLOPS dense FP16 to 82.6 TFLOPS FP32 while the A100 shows 312 TFLOPS dense FP16 to 19.5 TFLOPS FP32. Memory bandwidth of 1008 GB/s on the RTX 4090 constrains batch sizes relative to 2039 GB/s on the A100 during inference or training sessions. Dense INT8 performance reaches 660.6 TOPS on the RTX 4090 versus 624 TOPS on the A100.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

RTX 4090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiUnited Kingdom, GB2$0.51$1.01Deploy
RunPodglobal1$0.74—Deploy
LeaderGPUThe Netherlands8$0.88$7.04Deploy

3 providers in stock, 10 offers (cheapest per provider shown). All RTX 4090 offers, price history and alerts

A100

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
LeaderGPUThe Netherlands8$0.68$5.42Deploy
Vast.aiCzechia, CZ4$0.80$3.20Deploy
ThunderComputeUSA1$1.09—Deploy
Massed Computeus-central-34$1.35$5.40Deploy
HyperstackCANADA-11$1.35—Deploy

12 providers in stock, 32 offers (cheapest per provider shown). All A100 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when RTX 4090 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.51/GPU-hr.

QuantaCloud

Comparing A100 providers? We broker across all of them.

Need 16+ A100s reserved for fine-tuning, simulation, or production inference? We quote volume pricing across multiple data center partners: one quote at partner rates, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the RTX 4090

The RTX 4090 suits single GPU deployments that require 82.6 TFLOPS FP32 performance and 24 GB GDDR6X VRAM at a TDP of 450W. PCIe interconnect on the RTX 4090 fits consumer or workstation environments where NVLink is unnecessary. Workloads that value dense FP8 at 330.3 TFLOPS or sparse INT8 at 1321.2 TOPS benefit from the RTX 4090 specifications.

When to Choose the A100

The A100 suits multi GPU configurations that require 40 to 80 GB HBM2e VRAM and 2039 GB/s memory bandwidth at a TDP of 400W. NVLink and InfiniBand interconnect options on the A100 enable scaled training or inference clusters. Dense FP16 at 312 TFLOPS and sparse FP16 at 624 TFLOPS favor the A100 for memory intensive tasks.

Use Cases

LLM Training
A100

The A100 delivers 312 TFLOPS dense FP16 and up to 80 GB HBM2e VRAM which support larger models and batches than the 24 GB on the RTX 4090.

LLM Inference
A100

The A100 provides 2039 GB/s memory bandwidth and 624 TOPS dense INT8 allowing higher throughput with larger context sizes than the RTX 4090 at 1008 GB/s.

Fine-tuning
Either

Stable Diffusion
RTX 4090

The RTX 4090 supplies 82.6 TFLOPS FP32 and 330.3 TFLOPS dense FP8 which accelerate image generation workloads in single GPU PCIe setups.

Scientific Computing
RTX 4090

The RTX 4090 achieves 82.6 TFLOPS FP32 exceeding the 19.5 TFLOPS on the A100 for precision sensitive computations.

Frequently Asked Questions

How does FP16 performance compare on RTX 4090 versus A100?▾

Dense FP16 reaches 165.2 TFLOPS on the RTX 4090. Dense FP16 reaches 312 TFLOPS on the A100. Sparse FP16 reaches 330.4 TFLOPS on the RTX 4090 and 624 TFLOPS on the A100.

Which GPU has higher memory bandwidth?▾

The A100 reaches 2039 GB/s memory bandwidth. The RTX 4090 reaches 1008 GB/s memory bandwidth.

What interconnect options exist on each GPU?▾

The RTX 4090 uses PCIe 4.0. The A100 supports NVLink, PCIe 4.0, and InfiniBand.

What are the TDP ratings for RTX 4090 and A100?▾

The RTX 4090 has a TDP of 450W. The A100 has a TDP of 400W.

Which architecture does each GPU use?▾

The RTX 4090 uses Ada Lovelace architecture. The A100 uses Ampere architecture.

Which is cheaper to rent, the RTX 4090 or the A100?▾

Cloud rental prices for both the RTX 4090 and A100 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the RTX 4090 have compared to the A100?▾

The RTX 4090 has 24 GB of GDDR6X memory. The A100 has 40 to 80 GB of HBM2e memory.

Can I find RTX 4090 and A100 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the RTX 4090 and the A100?▾

The RTX 4090 uses the Ada Lovelace architecture (2022) while the A100 uses Ampere (2020). The A100 delivers 1.9x the dense FP16 throughput (312 vs 165.2 TFLOPS, both without sparsity) and 2.0x the memory bandwidth of the RTX 4090.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the RTX 4090 and the A100. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps