GPU comparison

A100 vs RTX 3090

Specifications and current cloud pricing, side by side.

AmperevsAmpereUpdated 17 days ago

The A100 wins for the most common use case of LLM training. Its 312 TFLOPS FP16 dense performance and 2039 GB/s bandwidth exceed the corresponding 71 TFLOPS FP16 dense and 936 GB/s figures of the RTX 3090 while its 40 to 80 GB memory capacity accommodates larger models than the 24 GB capacity of the RTX 3090.

A100 from $0.68/GPU/hrRTX 3090 from $0.27/GPU/hr

Right now, from live stock

  • Cheapest right now: RTX 3090 at $0.27/hr on Vast.ai

    Deploy
  • Most providers in stock: A100 (11)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: A100 ($0.0022 per TFLOPS-hour at 312 TFLOPS)

    Deploy

Specifications Compared

SpecA100RTX 3090
TDP400W350W
VRAM40-80 GB24 GB
CUDA Cores6,91210,496
Memory TypeHBM2eGDDR6X
ArchitectureAmpereAmpere
FP16 (dense)312 TFLOPS71 TFLOPS
Form FactorsSXM4, PCIePCIe
INT8 (dense)624 TOPS284 TOPS
InterconnectNVLink, PCIe 4.0, InfiniBandNVLink, PCIe 4.0
Tensor Cores432328
FP32 Performance19.5 TFLOPS35.6 TFLOPS
FP64 Performance9.7 TFLOPSNot published
Memory Bandwidth2,039 GB/s936 GB/s
FP16 (with sparsity)624 TFLOPS142 TFLOPS
INT8 (with sparsity)1,248 TOPS568 TOPS

Performance Analysis

The A100 delivers 312 TFLOPS FP16 dense compared with 71 TFLOPS FP16 dense on the RTX 3090. This fourfold ratio in dense FP16 performance favors the A100 for training and inference steps that rely on half precision tensor operations. The RTX 3090 delivers 35.6 TFLOPS FP32 compared with 19.5 TFLOPS FP32 on the A100 so the RTX 3090 holds an advantage in single precision workloads. Memory bandwidth of 2039 GB/s on the A100 versus 936 GB/s on the RTX 3090 allows larger batch sizes during training before memory capacity limits are reached. The A100 also supplies 624 TOPS INT8 dense against 284 TOPS INT8 dense on the RTX 3090 which widens the gap for quantized inference paths.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

A100

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
LeaderGPUThe Netherlands8$0.68$5.42Deploy
Vast.aiCzechia, CZ1$1.05—Deploy
ThunderComputeUSA2$1.09$2.18Deploy
Massed Computeus-central-21$1.35—Deploy
HyperstackCANADA-11$1.35—Deploy

11 providers in stock, 36 offers (cheapest per provider shown). All A100 offers, price history and alerts

RTX 3090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiFinland, FI4$0.27$1.07Deploy
LeaderGPUThe Netherlands8$0.29$2.29Deploy
RunPodglobal1$0.50—Deploy

3 providers in stock, 12 offers (cheapest per provider shown). All RTX 3090 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when A100 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.68/GPU-hr.

QuantaCloud

Comparing A100 providers? We broker across all of them.

Need 16+ A100s reserved for fine-tuning, simulation, or production inference? We quote volume pricing across multiple data center partners: one quote at partner rates, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the A100

The A100 suits large scale LLM training because its 40 to 80 GB HBM2e memory and 2039 GB/s bandwidth support bigger batch sizes than the 24 GB and 936 GB/s of the RTX 3090. Its 312 TFLOPS FP16 dense rating exceeds the 71 TFLOPS FP16 dense rating of the RTX 3090 so training iterations complete faster under dense half precision arithmetic.

When to Choose the RTX 3090

The RTX 3090 suits scientific computing workloads that depend on FP32 throughput because its 35.6 TFLOPS FP32 rating exceeds the 19.5 TFLOPS FP32 rating of the A100. Its 350W TDP is lower than the 400W TDP of the A100 and its PCIe form factor fits standard consumer servers without SXM4 requirements.

Use Cases

LLM Training
A100

The A100 supplies 312 TFLOPS FP16 dense against 71 TFLOPS FP16 dense on the RTX 3090 and its 2039 GB/s bandwidth supports larger batches than the 936 GB/s bandwidth of the RTX 3090.

LLM Inference
A100

The A100 supplies 312 TFLOPS FP16 dense and 624 TOPS INT8 dense which exceed the 71 TFLOPS FP16 dense and 284 TOPS INT8 dense ratings of the RTX 3090.

Fine-tuning
A100

The A100 supplies 312 TFLOPS FP16 dense and 40 to 80 GB memory which exceed the 71 TFLOPS FP16 dense and 24 GB memory of the RTX 3090 for sustained fine tuning runs.

Stable Diffusion
RTX 3090

The RTX 3090 supplies 35.6 TFLOPS FP32 which exceeds the 19.5 TFLOPS FP32 of the A100 and its 350W TDP fits consumer grade deployments more readily than the 400W TDP of the A100.

Scientific Computing
RTX 3090

The RTX 3090 supplies 35.6 TFLOPS FP32 which exceeds the 19.5 TFLOPS FP32 of the A100 for workloads that rely on single precision arithmetic.

Frequently Asked Questions

How does FP16 dense performance differ between the A100 and RTX 3090?▾

The A100 reaches 312 TFLOPS FP16 dense while the RTX 3090 reaches 71 TFLOPS FP16 dense. This difference determines training speed for half precision models.

What memory bandwidth does each GPU provide?▾

The A100 provides 2039 GB/s bandwidth and the RTX 3090 provides 936 GB/s bandwidth. Higher bandwidth on the A100 supports larger batch sizes before memory stalls occur.

Which GPU offers more FP32 performance?▾

The RTX 3090 reaches 35.6 TFLOPS FP32 while the A100 reaches 19.5 TFLOPS FP32. The RTX 3090 therefore suits single precision scientific workloads better.

What TDP ratings apply to the A100 and RTX 3090?▾

The A100 carries a 400W TDP and the RTX 3090 carries a 350W TDP. The lower TDP on the RTX 3090 reduces power draw in dense consumer installations.

Which is cheaper to rent, the A100 or the RTX 3090?▾

Cloud rental prices for both the A100 and RTX 3090 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the A100 have compared to the RTX 3090?▾

The A100 has 40 to 80 GB of HBM2e memory. The RTX 3090 has 24 GB of GDDR6X memory.

Can I find A100 and RTX 3090 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the A100 and the RTX 3090?▾

The A100 uses the Ampere architecture (2020) while the RTX 3090 uses Ampere (2020). The A100 delivers 4.4x the dense FP16 throughput (312 vs 71 TFLOPS, both without sparsity) and 2.2x the memory bandwidth of the RTX 3090.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the A100 and the RTX 3090. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps