GPU comparison

A100 SXM4 40GB vs RTX 3090

Specifications and current cloud pricing, side by side.

AmperevsAmpereUpdated 22 days ago

This variant comparison is part of the A100 vs RTX 3090 family comparison.

The A100 SXM4 40GB wins for the most common large scale training use case because its 312 TFLOPS FP16 dense rating and 40 GB HBM2e capacity exceed the RTX 3090 specifications by wide margins in tensor operations.

A100 SXM4 40GB from $0.67/GPU/hrRTX 3090 from $0.27/GPU/hr

Right now, from live stock

  • Cheapest right now: RTX 3090 at $0.27/hr on Vast.ai

    Deploy
  • Most providers in stock: A100 SXM4 40GB (12)

    See all offers

Specifications Compared

SpecA100 SXM4 40GBRTX 3090
TDP400W350W
VRAM40 GB24 GB
CUDA Cores6,91210,496
Memory TypeHBM2eGDDR6X
ArchitectureAmpereAmpere
FP16 (dense)312 TFLOPSNot published
Form FactorsSXM4PCIe
INT8 (dense)624 TOPSNot published
InterconnectNVLink, PCIe 4.0, InfiniBandNVLink
Tensor Cores432328
FP32 Performance19.5 TFLOPS35.6 TFLOPS
FP64 Performance9.7 TFLOPSNot published
Memory Bandwidth2,039 GB/s936 GB/s
FP16 (with sparsity)624 TFLOPSNot published
INT8 (with sparsity)1,248 TOPSNot published
FP16 (vendor figure, sparsity not specified)Not published35.6 TFLOPS

Performance Analysis

The A100 SXM4 40GB delivers 312 TFLOPS FP16 dense and 19.5 TFLOPS FP32 while the RTX 3090 delivers 35.6 TFLOPS FP16 and 35.6 TFLOPS FP32. This FP16 dense advantage of the A100 SXM4 40GB supports larger batch sizes during training workloads that rely on half precision tensor operations. The RTX 3090 maintains parity in FP32 throughput at 35.6 TFLOPS which favors certain graphics or mixed precision inference tasks. Memory bandwidth stands at 2039 GB/s for the A100 SXM4 40GB family and 936 GB/s for the RTX 3090 family so the A100 SXM4 40GB sustains higher data movement rates that reduce stalls when handling large tensors.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

A100 SXM4 40GB

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiSlovenia, SI1$0.67—Deploy
LeaderGPUThe Netherlands8$0.68$5.42Deploy
ThunderComputeUSA1$1.09—Deploy
HyperstackCANADA-11$1.35—Deploy
Massed Computeus-central-31$1.35—Deploy

12 providers in stock, 34 offers (cheapest per provider shown). All A100 SXM4 40GB offers, price history and alerts

RTX 3090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiFinland, FI4$0.27$1.07Deploy
LeaderGPUThe Netherlands8$0.29$2.29Deploy
RunPodglobal1$0.50—Deploy

3 providers in stock, 13 offers (cheapest per provider shown). All RTX 3090 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when A100 SXM4 40GB drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.67/GPU-hr.

QuantaCloud

Comparing A100 providers? We broker across all of them.

Need 16+ A100s reserved for fine-tuning, simulation, or production inference? We quote volume pricing across multiple data center partners: one quote at partner rates, 24h turnaround.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the A100 SXM4 40GB

The A100 SXM4 40GB suits workloads that require 40 GB HBM2e capacity and 2039 GB/s bandwidth such as large model training sessions that exceed 24 GB limits. Its SXM4 form factor and InfiniBand support enable multi GPU scaling in server environments where NVLink connectivity is available.

When to Choose the RTX 3090

Its PCIe form factor integrates directly into standard consumer systems without requiring specialized server infrastructure.

Use Cases

LLM Training
A100 SXM4 40GB

The A100 SXM4 40GB supplies 312 TFLOPS FP16 dense and 40 GB HBM2e that accommodate models too large for 24 GB GDDR6X.

LLM Inference
A100 SXM4 40GB

The A100 SXM4 40GB provides 624 TOPS INT8 dense throughput that processes larger batches than the RTX 3090 at 35.6 TFLOPS FP16.

Fine-tuning
A100 SXM4 40GB

The A100 SXM4 40GB 2039 GB/s bandwidth and 40 GB capacity support fine tuning runs that exceed the RTX 3090 memory limit.

Stable Diffusion
RTX 3090

The RTX 3090 delivers 35.6 TFLOPS FP32 that handles typical diffusion workloads within its 24 GB GDDR6X capacity at lower TDP.

Scientific Computing
A100 SXM4 40GB

The A100 SXM4 40GB supplies 19.5 TFLOPS FP32 along with NVLink and InfiniBand that scale across multiple SXM4 modules.

Frequently Asked Questions

What memory types distinguish the A100 SXM4 40GB from the RTX 3090?▾

The A100 SXM4 40GB uses 40 GB HBM2e while the RTX 3090 uses 24 GB GDDR6X. These capacities appear in the variant specifications.

How do the FP16 figures compare between the two GPUs?▾

The A100 SXM4 40GB lists 312 TFLOPS FP16 dense. The RTX 3090 lists 35.6 TFLOPS FP16. Both figures come from the shared family specifications.

Which GPU offers higher memory bandwidth?▾

The A100 SXM4 40GB family provides 2039 GB/s bandwidth. The RTX 3090 family provides 936 GB/s bandwidth. These values apply to every variant in each family.

What TDP values are listed for each model?▾

The A100 SXM4 40GB family lists 400W TDP. The RTX 3090 family lists 350W TDP. No variant specific adjustments exist in the stored data.

Do both GPUs support NVLink?▾

The A100 SXM4 40GB variant includes NVLink among its interconnect options. The RTX 3090 variant also includes NVLink. Both share this interconnect in their specifications.

Which is cheaper to rent, the A100 SXM4 40GB or the RTX 3090?▾

Cloud rental prices for both the A100 SXM4 40GB and RTX 3090 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the A100 SXM4 40GB have compared to the RTX 3090?▾

The A100 SXM4 40GB has 40 GB of HBM2e memory. The RTX 3090 has 24 GB of GDDR6X memory.

Can I find A100 SXM4 40GB and RTX 3090 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the A100 SXM4 40GB and the RTX 3090?▾

The A100 SXM4 40GB uses the Ampere architecture (2020) while the RTX 3090 uses Ampere (2020). The A100 SXM4 40GB delivers 4.4x the dense FP16 throughput (312 vs 71 TFLOPS, both without sparsity) and 2.2x the memory bandwidth of the RTX 3090.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the A100 SXM4 40GB and the RTX 3090. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps