GPU comparison

RTX 4090 vs V100

Specifications and current cloud pricing, side by side.

Ada LovelacevsVoltaUpdated 17 days ago

The RTX 4090 wins for the most common use case of dense FP16 workloads. Its 165.2 TFLOPS dense FP16 rating exceeds the V100 125 TFLOPS figure by a clear margin while the higher FP32 rating of 82.6 TFLOPS adds further advantage.

RTX 4090 from $0.40/GPU/hrV100 from $0.19/GPU/hr

Right now, from live stock

  • Cheapest right now: V100 at $0.19/hr on VERDA

    Deploy
  • Most providers in stock: RTX 4090 and V100 (3 each)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: V100 ($0.0016 per TFLOPS-hour at 125 TFLOPS)

    Deploy

Specifications Compared

SpecRTX 4090V100
TDP450W300W
VRAM24 GB16-32 GB
CUDA Cores16,3845,120
FP8 (dense)330.3 TFLOPSNot published
Memory TypeGDDR6XHBM2
ArchitectureAda LovelaceVolta
FP16 (dense)165.2 TFLOPS125 TFLOPS
Form FactorsPCIeSXM2, PCIe
INT8 (dense)660.6 TOPSNot published
InterconnectPCIe 4.0NVLink, PCIe 3.0
Tensor Cores512640
FP32 Performance82.6 TFLOPS15.7 TFLOPS
FP64 Performance1.3 TFLOPS7.8 TFLOPS
Memory Bandwidth1,008 GB/s900 GB/s
FP8 (with sparsity)660.6 TFLOPSNot published
FP16 (with sparsity)330.4 TFLOPSNot published
INT8 (with sparsity)1,321.2 TOPSNot published

Performance Analysis

Dense FP16 performance reaches 165.2 TFLOPS on the RTX 4090 compared to 125 TFLOPS on the V100. This ratio of 1.32 times higher dense FP16 on the RTX 4090 supports faster training iterations and inference throughput in compatible frameworks. FP32 figures show the RTX 4090 at 82.6 TFLOPS versus 15.7 TFLOPS on the V100 which translates to over five times the throughput for precision sensitive computations. Memory bandwidth of 1008 GB/s on the RTX 4090 versus 900 GB/s on the V100 permits modestly larger batch sizes during model execution before memory constraints appear. Sparsity enabled FP16 on the RTX 4090 reaches 330.4 TFLOPS while the V100 provides no sparsity figure for direct comparison.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

RTX 4090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiIreland, IE1$0.40—Deploy
RunPodglobal1$0.74—Deploy
LeaderGPUThe Netherlands8$0.88$7.04Deploy

3 providers in stock, 10 offers (cheapest per provider shown). All RTX 4090 offers, price history and alerts

V100

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
VERDAFIN-HEL1$0.19—Deploy
Oribeauharnois4$0.83$3.32Deploy
Paperspaceny21$2.30—Deploy

3 providers in stock, 75 offers (cheapest per provider shown). All V100 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when RTX 4090 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.40/GPU-hr.

QuantaCloud

Comparing providers? We broker across all of them.

Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the RTX 4090

The RTX 4090 suits workloads that demand the 82.6 TFLOPS FP32 rating or the 165.2 TFLOPS dense FP16 rating. Its 24 GB memory capacity and 1008 GB/s bandwidth further support larger data sets in single GPU configurations.

When to Choose the V100

The V100 fits environments where the 300W TDP rating reduces power draw relative to the 450W TDP of the RTX 4090. NVLink interconnect availability on the V100 also enables multi GPU scaling in systems that require that specific fabric.

Use Cases

LLM Training
RTX 4090

The RTX 4090 provides 165.2 TFLOPS dense FP16 which exceeds the V100 125 TFLOPS dense FP16 rating.

LLM Inference
RTX 4090

Dense FP16 throughput at 165.2 TFLOPS on the RTX 4090 supports higher tokens per second than the V100 125 TFLOPS.

Fine-tuning
RTX 4090

FP32 performance of 82.6 TFLOPS on the RTX 4090 enables more efficient gradient updates than the V100 15.7 TFLOPS.

Stable Diffusion
RTX 4090

The RTX 4090 memory bandwidth of 1008 GB/s exceeds the V100 900 GB/s for image generation batches.

Scientific Computing
V100

The V100 TDP of 300W lowers power requirements compared to the RTX 4090 450W TDP in sustained runs.

Frequently Asked Questions

How does dense FP16 performance compare between the RTX 4090 and V100?▾

The RTX 4090 reaches 165.2 TFLOPS dense FP16 while the V100 reaches 125 TFLOPS dense FP16. This difference favors the RTX 4090 for tensor operations that rely on dense FP16 arithmetic.

What FP32 performance do the RTX 4090 and V100 deliver?▾

The RTX 4090 delivers 82.6 TFLOPS FP32. The V100 delivers 15.7 TFLOPS FP32.

Which GPU has higher memory bandwidth?▾

The RTX 4090 provides 1008 GB/s memory bandwidth. The V100 provides 900 GB/s memory bandwidth.

How do TDP ratings differ for the RTX 4090 and V100?▾

The RTX 4090 carries a 450W TDP rating. The V100 carries a 300W TDP rating.

What interconnect options exist on each GPU?▾

The RTX 4090 uses PCIe 4.0. The V100 supports NVLink along with PCIe 3.0.

Which is cheaper to rent, the RTX 4090 or the V100?▾

Cloud rental prices for both the RTX 4090 and V100 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the RTX 4090 have compared to the V100?▾

The RTX 4090 has 24 GB of GDDR6X memory. The V100 has 16 to 32 GB of HBM2 memory.

Can I find RTX 4090 and V100 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the RTX 4090 and the V100?▾

The RTX 4090 uses the Ada Lovelace architecture (2022) while the V100 uses Volta (2017). The RTX 4090 delivers 1.3x the dense FP16 throughput (165.2 vs 125 TFLOPS, both without sparsity) and 1.1x the memory bandwidth of the V100.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the RTX 4090 and the V100. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps