GPU comparison

A100 PCIe 40GB vs RTX 3090

Specifications and current cloud pricing, side by side.

AmperevsAmpereUpdated 15 days ago

This variant comparison is part of the A100 vs RTX 3090 family comparison.

The A100 PCIe 40GB wins for the most common AI training and inference workloads because 312 TFLOPS dense FP16 and 40 GB HBM2e exceed the corresponding RTX 3090 specifications by wide margins.

A100 PCIe 40GB from $0.67/GPU/hrRTX 3090 from $0.27/GPU/hr

Right now, from live stock

  • Cheapest right now: RTX 3090 at $0.27/hr on Vast.ai

    Deploy
  • Most providers in stock: A100 PCIe 40GB (11)

    See all offers

Specifications Compared

SpecA100 PCIe 40GBRTX 3090
TDP400W350W
VRAM40 GB24 GB
CUDA Cores6,91210,496
Memory TypeHBM2eGDDR6X
ArchitectureAmpereAmpere
FP16 (dense)312 TFLOPSNot published
Form FactorsPCIePCIe
INT8 (dense)624 TOPSNot published
InterconnectPCIe 4.0, InfiniBand; NVLink only via a bridge between card pairsNVLink
Tensor Cores432328
FP32 Performance19.5 TFLOPS35.6 TFLOPS
FP64 Performance9.7 TFLOPSNot published
Memory Bandwidth2,039 GB/s936 GB/s
FP16 (with sparsity)624 TFLOPSNot published
INT8 (with sparsity)1,248 TOPSNot published
FP16 (vendor figure, sparsity not specified)Not published35.6 TFLOPS

Performance Analysis

The A100 PCIe 40GB delivers 312 TFLOPS in dense FP16 while the RTX 3090 delivers 35.6 TFLOPS in FP16. This gap means the A100 PCIe 40GB completes dense FP16 matrix operations nearly nine times faster than the RTX 3090. The RTX 3090 records 35.6 TFLOPS in FP32 which exceeds the A100 PCIe 40GB figure of 19.5 TFLOPS in FP32. Memory bandwidth of 2039 GB/s on the A100 family supports larger batch sizes during training compared with 936 GB/s on the RTX 3090 family. The A100 family also lists 624 TOPS in dense INT8 while the RTX 3090 provides no equivalent sparsity labeled figure for direct comparison.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

A100 PCIe 40GB

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiSlovenia, SI1$0.67—Deploy
LeaderGPUThe Netherlands8$0.68$5.42Deploy
ThunderComputeUSA2$1.09$2.18Deploy
Massed Computeus-central-32$1.35$2.70Deploy
QuantaCloudus-midwest-22$1.48$2.95Deploy

11 providers in stock, 31 offers (cheapest per provider shown). All A100 PCIe 40GB offers, price history and alerts

RTX 3090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiFinland, FI4$0.27$1.07Deploy
LeaderGPUThe Netherlands8$0.29$2.29Deploy
RunPodglobal1$0.50—Deploy

3 providers in stock, 13 offers (cheapest per provider shown). All RTX 3090 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when A100 PCIe 40GB drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.67/GPU-hr.

QuantaCloud

Comparing A100 providers? We broker across all of them.

Need 16+ A100s reserved for fine-tuning, simulation, or production inference? We quote volume pricing across multiple data center partners: one quote at partner rates, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the A100 PCIe 40GB

The A100 PCIe 40GB suits workloads that require 40 GB HBM2e capacity and 312 TFLOPS dense FP16 throughput. Large model training and scientific computing benefit from the 2039 GB/s memory bandwidth and 400W TDP envelope. Multi card NVLink bridging further extends scalability for distributed jobs.

When to Choose the RTX 3090

The RTX 3090 suits FP32 centric tasks that value 35.6 TFLOPS in FP32 and a 350W TDP. Inference pipelines with modest memory needs of 24 GB GDDR6X run efficiently on this variant. Direct NVLink support aids certain consumer multi card setups.

Use Cases

LLM Training
A100 PCIe 40GB

The A100 PCIe 40GB provides 312 TFLOPS dense FP16 and 40 GB HBM2e which accommodate larger models than the RTX 3090 24 GB limit.

LLM Inference
A100 PCIe 40GB

Dense FP16 throughput of 312 TFLOPS on the A100 PCIe 40GB enables higher throughput than the 35.6 TFLOPS available on the RTX 3090.

Fine-tuning
A100 PCIe 40GB

Memory bandwidth of 2039 GB/s and 40 GB capacity on the A100 PCIe 40GB support larger batch sizes during fine tuning compared with 936 GB/s on the RTX 3090.

Stable Diffusion
RTX 3090

The RTX 3090 delivers 35.6 TFLOPS in FP32 which matches diffusion workloads and operates within a 350W TDP.

Scientific Computing
A100 PCIe 40GB

The A100 PCIe 40GB supplies 19.5 TFLOPS in FP32 together with 2039 GB/s bandwidth suited to double precision scientific codes.

Frequently Asked Questions

What are the FP16 performance figures?▾

The A100 family lists 312 TFLOPS dense FP16 and the RTX 3090 lists 35.6 TFLOPS FP16.

Which card offers higher memory bandwidth?▾

The A100 family reaches 2039 GB/s while the RTX 3090 family reaches 936 GB/s.

What TDP values apply to these variants?▾

The A100 family operates at 400W TDP and the RTX 3090 family operates at 350W TDP.

Do both cards support NVLink?▾

The A100 PCIe 40GB supports NVLink only via a bridge between card pairs while the RTX 3090 supports NVLink directly.

Which architecture do these GPUs share?▾

Both the A100 PCIe 40GB and the RTX 3090 follow the Ampere architecture released in 2020.

Which is cheaper to rent, the A100 PCIe 40GB or the RTX 3090?▾

Cloud rental prices for both the A100 PCIe 40GB and RTX 3090 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the A100 PCIe 40GB have compared to the RTX 3090?▾

The A100 PCIe 40GB has 40 GB of HBM2e memory. The RTX 3090 has 24 GB of GDDR6X memory.

Can I find A100 PCIe 40GB and RTX 3090 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the A100 PCIe 40GB and the RTX 3090?▾

The A100 PCIe 40GB uses the Ampere architecture (2020) while the RTX 3090 uses Ampere (2020). The A100 PCIe 40GB delivers 4.4x the dense FP16 throughput (312 vs 71 TFLOPS, both without sparsity) and 2.2x the memory bandwidth of the RTX 3090.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the A100 PCIe 40GB and the RTX 3090. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps