Specifications Compared
| Spec | L4 | L40S |
|---|---|---|
| TDP | 72W | 350W |
| VRAM | 24 GB | 48 GB |
| CUDA Cores | 7,424 | 18,176 |
| FP8 (dense) | 242 TFLOPS | 733 TFLOPS |
| Memory Type | GDDR6 | GDDR6 |
| Architecture | Ada Lovelace | Ada Lovelace |
| FP16 (dense) | 121 TFLOPS | 362.05 TFLOPS |
| Form Factors | PCIe | PCIe |
| INT8 (dense) | 242 TOPS | 733 TOPS |
| Interconnect | PCIe 4.0 | PCIe 4.0 |
| Tensor Cores | 232 | 568 |
| FP32 Performance | 30.3 TFLOPS | 91.6 TFLOPS |
| FP64 Performance | 0.5 TFLOPS | 1.4 TFLOPS |
| Memory Bandwidth | 300 GB/s | 864 GB/s |
| FP8 (with sparsity) | 485 TFLOPS | 1,466 TFLOPS |
| FP16 (with sparsity) | 242 TFLOPS | 733 TFLOPS |
| INT8 (with sparsity) | 485 TOPS | 1,466 TOPS |
Performance Analysis
The L4 delivers 121 TFLOPS in dense FP16 while the L40S reaches 362.05 TFLOPS in dense FP16. The L40S therefore supplies three times the dense FP16 throughput of the L4. In FP32 the L40S provides 91.6 TFLOPS against the L4 figure of 30.3 TFLOPS. Higher memory bandwidth of 864 GB/s on the L40S versus 300 GB/s on the L4 supports larger batch sizes during both training and inference. The L4 FP8 dense performance of 242 TFLOPS remains lower than the L40S FP8 dense performance of 733 TFLOPS. Sparsity multiplies both GPUs by the same factor yet preserves the same relative gap between them. Real world training benefits from the L40S dense FP16 advantage while inference on smaller models can fit within the L4 24 GB capacity at reduced power draw.
Current On-Demand Offers
Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.
L4
| Provider | Region | GPUs | Per GPU / hr | Instance / hr | Deploy |
|---|---|---|---|---|---|
| RunPod | global | 1 | $0.49 | — | Deploy |
| Scaleway | Warsaw, Poland (WAW2) | 8 | $0.91 | $7.27 | Deploy |
2 providers in stock, 7 offers (cheapest per provider shown). All L4 offers, price history and alerts
L40S
| Provider | Region | GPUs | Per GPU / hr | Instance / hr | Deploy |
|---|---|---|---|---|---|
| Packet.ai | Reading, United Kingdom (UK-1) | 1 | $0.92 | — | Deploy |
| Massed Compute | us-central-2 | 4 | $0.97 | $3.88 | Deploy |
| QuantaCloud | us-midwest-1 | 1 | $1.09 | — | Deploy |
3 providers in stock, 15 offers (cheapest per provider shown). All L40S offers, price history and alerts
Notify me when L4 drops below a price
One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.49/GPU-hr.
QuantaCloud
Comparing providers? We broker across all of them.
Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.
When to Choose the L4
The L4 suits deployments that must stay within a 72W power envelope. Its 24 GB GDDR6 memory and 300 GB/s bandwidth handle moderate inference loads that do not require the 362.05 TFLOPS dense FP16 of the L40S. Users select the L4 when PCIe slot power limits or thermal constraints rule out the 350W TDP of the L40S.
When to Choose the L40S
The L40S fits workloads that need 48 GB GDDR6 memory and 864 GB/s bandwidth to support large batch sizes. Its 362.05 TFLOPS dense FP16 and 91.6 TFLOPS FP32 enable faster training and inference than the L4 equivalents.
Use Cases
The L40S supplies 362.05 TFLOPS dense FP16 and 48 GB memory to accommodate larger models and batches than the L4 121 TFLOPS dense FP16 and 24 GB.
The L40S 733 TFLOPS FP8 dense and 864 GB/s bandwidth support higher throughput than the L4 242 TFLOPS FP8 dense and 300 GB/s.
The L40S 91.6 TFLOPS FP32 and 48 GB capacity handle fine-tuning workloads that exceed L4 30.3 TFLOPS FP32 and 24 GB limits.
Both GPUs run Stable Diffusion yet the L40S 48 GB memory permits larger resolutions while the L4 72W TDP suits power limited setups.
The L40S 91.6 TFLOPS FP32 exceeds the L4 30.3 TFLOPS FP32 and therefore accelerates compute heavy scientific codes.
Frequently Asked Questions
What are the FP16 dense ratings?▾
The L4 lists 121 TFLOPS dense FP16 and the L40S lists 362.05 TFLOPS dense FP16. The L40S therefore provides three times the dense FP16 throughput.
How do the power draws compare?▾
The L4 TDP is 72W and the L40S TDP is 350W. Systems must accommodate nearly five times the power for the L40S.
Which GPU has higher memory bandwidth?▾
The L40S reaches 864 GB/s while the L4 reaches 300 GB/s. The L40S bandwidth supports larger batches during inference.
Do both GPUs use the same interconnect?▾
Both the L4 and L40S employ PCIe 4.0. No difference exists in interconnect technology between the two cards.
What FP32 performance do they offer?▾
The L4 provides 30.3 TFLOPS FP32 and the L40S provides 91.6 TFLOPS FP32. The L40S delivers three times the FP32 throughput.
Which is cheaper to rent, the L4 or the L40S?▾
Cloud rental prices for both the L4 and L40S vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.
How much VRAM does the L4 have compared to the L40S?▾
The L4 has 24 GB of GDDR6 memory. The L40S has 48 GB of GDDR6 memory.
Can I find L4 and L40S GPUs available to rent right now?▾
Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.
What is the main difference between the L4 and the L40S?▾
The L4 uses the Ada Lovelace architecture (2023) while the L40S uses Ada Lovelace (2023). The L40S delivers 3.0x the dense FP16 throughput (362.05 vs 121 TFLOPS, both without sparsity) and 2.9x the memory bandwidth of the L4.
Rent these GPUs
Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.
Related comparisons
How this page is made
- Specifications come from the NVIDIA, AMD and Intel datasheets for the L4 and the L40S. Dense and sparse throughput are listed separately.
- Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
- The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
- Read how we collect the data, or report an error on this page.
Next steps
- Rent L4Every current offer by provider, daily price history and a price alert.
- Rent L40SEvery current offer by provider, daily price history and a price alert.
- GPU price indexHow on-demand prices for the major GPUs have moved, updated daily.
- All GPUsEvery GPU we track, with current lows and provider counts.