GPU comparison

A10 vs RTX 4090

Specifications and current cloud pricing, side by side.

AmperevsAda LovelaceUpdated 17 days ago

The RTX 4090 wins for the most common mixed precision inference and training workloads because its dense FP16 rating of 165.2 TFLOPS exceeds the A10 rating of 125 TFLOPS while its memory bandwidth of 1008 GB/s exceeds the A10 rating of 600 GB/s.

A10 from $0.37/GPU/hrRTX 4090 from $0.40/GPU/hr

Right now, from live stock

  • Cheapest right now: A10 at $0.37/hr on LeaderGPU

    Deploy
  • Most providers in stock: RTX 4090 (3)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: RTX 4090 ($0.0024 per TFLOPS-hour at 165.2 TFLOPS)

    Deploy

Specifications Compared

SpecA10RTX 4090
TDP150W450W
VRAM24 GB24 GB
CUDA Cores9,21616,384
FP8 (dense)Not published330.3 TFLOPS
Memory TypeGDDR6GDDR6X
ArchitectureAmpereAda Lovelace
FP16 (dense)125 TFLOPS165.2 TFLOPS
Form FactorsPCIePCIe
INT8 (dense)250 TOPS660.6 TOPS
InterconnectPCIe 4.0PCIe 4.0
Tensor Cores288512
FP32 Performance31.2 TFLOPS82.6 TFLOPS
FP64 PerformanceNot published1.3 TFLOPS
Memory Bandwidth600 GB/s1,008 GB/s
FP8 (with sparsity)Not published660.6 TFLOPS
FP16 (with sparsity)250 TFLOPS330.4 TFLOPS
INT8 (with sparsity)500 TOPS1,321.2 TOPS

Performance Analysis

Dense FP16 performance reaches 125 TFLOPS on the A10 and 165.2 TFLOPS on the RTX 4090 while sparse FP16 performance reaches 250 TFLOPS on the A10 and 330.4 TFLOPS on the RTX 4090. The FP16 to FP32 ratio equals 4.0 on the A10 and 2.0 on the RTX 4090 indicating different acceleration profiles for mixed precision training and inference workloads. Memory bandwidth of 600 GB/s on the A10 versus 1008 GB/s on the RTX 4090 directly limits maximum batch sizes in memory bound operations such as large model inference.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

A10

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
LeaderGPUThe Netherlands10$0.37$3.71Deploy
Lambda Labsus-east-11$1.29—Deploy

2 providers in stock, 3 offers (cheapest per provider shown). All A10 offers, price history and alerts

RTX 4090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiIreland, IE1$0.40—Deploy
RunPodglobal1$0.74—Deploy
LeaderGPUThe Netherlands8$0.88$7.04Deploy

3 providers in stock, 10 offers (cheapest per provider shown). All RTX 4090 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when A10 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.37/GPU-hr.

QuantaCloud

Comparing providers? We broker across all of them.

Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the A10

The A10 suits environments that prioritize a TDP of 150W over peak throughput. Its dense INT8 rating of 250 TOPS combined with the lower power envelope supports sustained operation in dense server deployments where cooling capacity remains constrained.

When to Choose the RTX 4090

The RTX 4090 suits workloads that require the higher dense FP16 rating of 165.2 TFLOPS and the higher memory bandwidth of 1008 GB/s. Its FP32 rating of 82.6 TFLOPS enables faster single precision scientific kernels compared with the 31.2 TFLOPS rating of the A10.

Use Cases

LLM Training
RTX 4090

The RTX 4090 delivers dense FP16 performance of 165.2 TFLOPS compared with 125 TFLOPS on the A10.

LLM Inference
RTX 4090

The RTX 4090 provides memory bandwidth of 1008 GB/s which supports larger batch sizes than the 600 GB/s bandwidth of the A10.

Fine-tuning
RTX 4090

The RTX 4090 supplies sparse FP16 performance of 330.4 TFLOPS compared with 250 TFLOPS on the A10.

Stable Diffusion
RTX 4090

The RTX 4090 supplies dense FP16 performance of 165.2 TFLOPS and FP32 performance of 82.6 TFLOPS.

Scientific Computing
RTX 4090

The RTX 4090 supplies FP32 performance of 82.6 TFLOPS compared with 31.2 TFLOPS on the A10.

Frequently Asked Questions

How does FP32 performance differ between the A10 and the RTX 4090?▾

The A10 delivers 31.2 TFLOPS of FP32 performance. The RTX 4090 delivers 82.6 TFLOPS of FP32 performance.

What TDP values apply to the A10 and the RTX 4090?▾

The A10 lists a TDP of 150W. The RTX 4090 lists a TDP of 450W.

Do both GPUs support the same interconnect?▾

Both the A10 and the RTX 4090 use PCIe 4.0 interconnect.

How does memory bandwidth compare?▾

The A10 provides 600 GB/s of memory bandwidth. The RTX 4090 provides 1008 GB/s of memory bandwidth.

Which is cheaper to rent, the A10 or the RTX 4090?▾

Cloud rental prices for both the A10 and RTX 4090 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the A10 have compared to the RTX 4090?▾

The A10 has 24 GB of GDDR6 memory. The RTX 4090 has 24 GB of GDDR6X memory.

Can I find A10 and RTX 4090 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the A10 and the RTX 4090?▾

The A10 uses the Ampere architecture (2021) while the RTX 4090 uses Ada Lovelace (2022). The RTX 4090 delivers 1.3x the dense FP16 throughput (165.2 vs 125 TFLOPS, both without sparsity) and 1.7x the memory bandwidth of the A10.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the A10 and the RTX 4090. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps