GPU comparison

P100 vs RTX 5090

Specifications and current cloud pricing, side by side.

PascalvsBlackwellUpdated 24 days ago

The RTX 5090 wins for the most common use case of LLM inference because its 209.5 TFLOPS dense FP16 and 1792 GB/s bandwidth deliver substantially higher throughput than the 18.7 TFLOPS dense FP16 and 732 GB/s bandwidth of the P100 while still supporting the same dense arithmetic mode.

P100 from $0.09/GPU/hrRTX 5090 from $0.87/GPU/hr

Right now, from live stock

  • Cheapest right now: P100 at $0.09/hr on Vast.ai

    Deploy
  • Most providers in stock: RTX 5090 (3)

    See all offers
  • Best $ per TFLOPS FP16 dense right now: RTX 5090 ($0.0041 per TFLOPS-hour at 209.5 TFLOPS)

    Deploy

Specifications Compared

SpecP100RTX 5090
TDP250W575W
VRAM16 GB32 GB
CUDA Cores3,58421,760
FP4 (dense)Not published1,676 TFLOPS
FP8 (dense)Not published419 TFLOPS
Memory TypeHBM2GDDR7
ArchitecturePascalBlackwell
FP16 (dense)18.7 TFLOPS209.5 TFLOPS
Form FactorsSXM2, PCIePCIe
INT8 (dense)Not published838 TOPS
InterconnectNVLink, PCIe 3.0PCIe 5.0
Tensor CoresNot published680
FP32 Performance9.3 TFLOPS104.8 TFLOPS
FP64 Performance4.7 TFLOPS1.6 TFLOPS
Memory Bandwidth732 GB/s1,792 GB/s
FP4 (with sparsity)Not published3,352 TFLOPS
FP8 (with sparsity)Not published838 TFLOPS
FP16 (with sparsity)Not published419 TFLOPS
INT8 (with sparsity)Not published1,676 TOPS

Performance Analysis

The FP16 to FP32 ratio remains 2 times on both GPUs because the P100 lists 18.7 TFLOPS dense FP16 against 9.3 TFLOPS FP32 and the RTX 5090 lists 209.5 TFLOPS dense FP16 against 104.8 TFLOPS FP32. This ratio indicates that both cards double throughput when workloads stay in dense FP16 rather than FP32 during training or inference steps. Memory bandwidth of 1792 GB/s on the RTX 5090 versus 732 GB/s on the P100 permits larger batch sizes before memory stalls occur in data heavy operations. When sparsity is applied the RTX 5090 reaches 419 TFLOPS in FP16 with sparsity which exceeds the dense FP16 figure of the P100 by more than 22 times and therefore accelerates sparse matrix workloads in inference pipelines.

Current On-Demand Offers

Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.

P100

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiTexas, US1$0.09—Deploy

1 provider in stock, 3 offers (cheapest per provider shown). All P100 offers, price history and alerts

RTX 5090

ProviderRegionGPUsPer GPU / hrInstance / hrDeploy
Vast.aiAlberta, CA1$0.87—Deploy
RunPodglobal1$1.19—Deploy
LeaderGPUThe Netherlands1$2.40—Deploy

3 providers in stock, 4 offers (cheapest per provider shown). All RTX 5090 offers, price history and alerts

Which GPU to watchWatch the price of

Notify me when P100 drops below a price

One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.09/GPU-hr.

QuantaCloud

Comparing providers? We broker across all of them.

Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.

No waitlist24hr quote turnaroundInfiniBand fabric

Compare real-time pricing across 25+ providers

When to Choose the P100

The P100 suits environments that prioritize a 250W TDP over higher throughput because power constraints limit total system draw. Its NVLink interconnect supports direct GPU to GPU transfers at scales where PCIe 3.0 alone would constrain multi card clusters. Older code bases written for Pascal architecture also run without recompilation on the P100 while newer Blackwell features remain unused.

When to Choose the RTX 5090

The RTX 5090 suits workloads that require 209.5 TFLOPS dense FP16 or 104.8 TFLOPS FP32 because these figures exceed the corresponding P100 numbers by more than 11 times. Its 32 GB GDDR7 memory and 1792 GB/s bandwidth accommodate larger models or batch sizes than the 16 GB HBM2 and 732 GB/s on the P100. PCIe 5.0 interconnect further reduces host transfer latency when data movement dominates execution time.

Use Cases

LLM Training
RTX 5090

The RTX 5090 provides 209.5 TFLOPS dense FP16 compared with 18.7 TFLOPS dense FP16 on the P100 allowing faster iteration through large training runs.

LLM Inference
RTX 5090

The RTX 5090 supplies 209.5 TFLOPS dense FP16 and 1792 GB/s bandwidth which exceed the 18.7 TFLOPS dense FP16 and 732 GB/s bandwidth of the P100 and therefore support higher token rates.

Fine-tuning
RTX 5090

The RTX 5090 lists 32 GB GDDR7 memory against 16 GB HBM2 on the P100 enabling larger adapter modules during parameter updates.

Stable Diffusion
RTX 5090

The RTX 5090 reaches 419 TFLOPS FP8 dense which surpasses any published figure on the P100 and accelerates diffusion sampling steps.

Scientific Computing
P100

The P100 operates at a 250W TDP which reduces total energy draw compared with the 575W TDP of the RTX 5090 in sustained floating point workloads.

Frequently Asked Questions

How does FP32 performance compare between these two GPUs?▾

The P100 delivers 9.3 TFLOPS FP32 and the RTX 5090 delivers 104.8 TFLOPS FP32. The RTX 5090 therefore provides more than 11 times the FP32 throughput of the P100.

Which GPU offers higher memory bandwidth?▾

The RTX 5090 lists 1792 GB/s memory bandwidth while the P100 lists 732 GB/s memory bandwidth. The higher figure on the RTX 5090 supports larger data transfers per second during kernel execution.

Does the P100 support NVLink?▾

The P100 includes NVLink interconnect alongside PCIe 3.0. The RTX 5090 supports only PCIe 5.0 without NVLink.

What are the TDP ratings for each card?▾

The P100 carries a 250W TDP and the RTX 5090 carries a 575W TDP. System designers must account for the more than double power requirement on the RTX 5090.

Which is cheaper to rent, the P100 or the RTX 5090?▾

Cloud rental prices for both the P100 and RTX 5090 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.

How much VRAM does the P100 have compared to the RTX 5090?▾

The P100 has 16 GB of HBM2 memory. The RTX 5090 has 32 GB of GDDR7 memory.

Can I find P100 and RTX 5090 GPUs available to rent right now?▾

Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.

What is the main difference between the P100 and the RTX 5090?▾

The P100 uses the Pascal architecture (2016) while the RTX 5090 uses Blackwell (2025). The RTX 5090 delivers 11.2x the dense FP16 throughput (209.5 vs 18.7 TFLOPS, both without sparsity) and 2.4x the memory bandwidth of the P100.

Rent these GPUs

Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.

Related comparisons

How this page is made

  • Specifications come from the NVIDIA, AMD and Intel datasheets for the P100 and the RTX 5090. Dense and sparse throughput are listed separately.
  • Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
  • The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
  • Read how we collect the data, or report an error on this page.

Next steps