Specifications Compared
| Spec | L40 | RTX 4090 |
|---|---|---|
| TDP | 300W | 450W |
| VRAM | 48 GB | 24 GB |
| CUDA Cores | 18,176 | 16,384 |
| FP8 (dense) | 362 TFLOPS | 330.3 TFLOPS |
| Memory Type | GDDR6 | GDDR6X |
| Architecture | Ada Lovelace | Ada Lovelace |
| FP16 (dense) | 181.05 TFLOPS | 165.2 TFLOPS |
| Form Factors | PCIe | PCIe |
| INT8 (dense) | 362 TOPS | 660.6 TOPS |
| Interconnect | PCIe 4.0 | PCIe 4.0 |
| Tensor Cores | 568 | 512 |
| FP32 Performance | 90.5 TFLOPS | 82.6 TFLOPS |
| FP64 Performance | Not published | 1.3 TFLOPS |
| Memory Bandwidth | 864 GB/s | 1,008 GB/s |
| FP8 (with sparsity) | 724 TFLOPS | 660.6 TFLOPS |
| FP16 (with sparsity) | 362.1 TFLOPS | 330.4 TFLOPS |
| INT8 (with sparsity) | 724 TOPS | 1,321.2 TOPS |
Performance Analysis
Dense FP16 performance stands at 181.05 TFLOPS on the L40 and 165.2 TFLOPS on the RTX 4090. The FP16 to FP32 ratio equals 2 on both cards because 181.05 divided by 90.5 yields 2 and 165.2 divided by 82.6 yields 2. This ratio indicates that mixed precision training or inference can double throughput relative to pure FP32 execution on either card. Sparse FP16 performance reaches 362.1 TFLOPS on the L40 and 330.4 TFLOPS on the RTX 4090. Memory bandwidth of 864 GB per second on the L40 versus 1008 GB per second on the RTX 4090 determines the maximum sustainable batch size when activation tensors exceed on chip cache capacity. Higher bandwidth on the RTX 4090 supports larger batches in bandwidth bound kernels while the L40 larger 48 GB capacity supports models whose weights alone exceed 24 GB.
Current On-Demand Offers
Cheapest secure on-demand offer per provider that is reported in stock right now, per GPU per hour; multi-GPU instances show the whole-instance price alongside. Deploy opens the DeployGPU console for providers on the platform, otherwise the provider's own site. Stock is rechecked every minute.
L40
| Provider | Region | GPUs | Per GPU / hr | Instance / hr | Deploy |
|---|---|---|---|---|---|
| ThunderCompute | USA | 2 | $0.79 | $1.58 | Deploy |
| RunPod | global | 1 | $0.82 | — | Deploy |
| Massed Compute | us-central-3 | 2 | $0.86 | $1.72 | Deploy |
| QuantaCloud | us-midwest-2 | 4 | $0.94 | $3.75 | Deploy |
| Hyperstack | CANADA-1 | 1 | $1.00 | — | Deploy |
5 providers in stock, 19 offers (cheapest per provider shown). All L40 offers, price history and alerts
Notify me when L40 drops below a price
One email when the cheapest in-stock on-demand price per GPU-hour falls below your threshold. Cheapest right now: $0.79/GPU-hr.
QuantaCloud
Comparing providers? We broker across all of them.
Stop tab-switching between pricing pages. Tell us what you need, 16+ GPUs reserved or cluster capacity, and we return one quote at partner rates within 24 hours.
When to Choose the L40
The L40 suits workloads that require 48 GB of memory such as large language model inference with batch sizes that fit only in the larger frame buffer. Its 300 W thermal design power also fits installations where power delivery or cooling capacity remains constrained to 300 W per card. Dense FP8 performance of 362 TFLOPS further benefits quantized inference pipelines that stay within the 48 GB limit.
When to Choose the RTX 4090
Its INT8 dense performance of 660.6 TOPS exceeds the L40 figure of 362 TOPS and therefore accelerates integer quantized inference when memory capacity of 24 GB proves sufficient.
Use Cases
The L40 supplies 48 GB of memory which accommodates larger model states than the 24 GB available on the RTX 4090.
The L40 supplies 48 GB of memory which accommodates larger model states than the 24 GB available on the RTX 4090.
Both cards deliver comparable dense FP16 performance of 181.05 TFLOPS and 165.2 TFLOPS so selection depends on whether memory capacity or bandwidth dominates the workload.
The RTX 4090 supplies 1008 GB per second memory bandwidth which exceeds the 864 GB per second on the L40 and therefore sustains higher throughput in bandwidth bound diffusion kernels.
The L40 lists a 300 W thermal design power which reduces total system power draw compared with the 450 W rating of the RTX 4090.
Frequently Asked Questions
What is the FP32 performance of each card?▾
The L40 delivers 90.5 TFLOPS in FP32. The RTX 4090 delivers 82.6 TFLOPS in FP32.
Which card has higher memory bandwidth?▾
The RTX 4090 provides 1008 GB per second of memory bandwidth. The L40 provides 864 GB per second of memory bandwidth.
What are the dense FP16 figures?▾
The L40 lists 181.05 TFLOPS dense FP16. The RTX 4090 lists 165.2 TFLOPS dense FP16.
How do the TDP ratings compare?▾
The L40 carries a 300 W thermal design power. The RTX 4090 carries a 450 W thermal design power.
Do both cards use the same interconnect?▾
Both the L40 and the RTX 4090 use PCIe 4.0 interconnect.
Which is cheaper to rent, the L40 or the RTX 4090?▾
Cloud rental prices for both the L40 and RTX 4090 vary by provider, configuration, and availability. This page shows live pricing from 25+ providers updated every 60 seconds. Scroll to the Live Cloud Pricing section to compare current rates.
How much VRAM does the L40 have compared to the RTX 4090?▾
The L40 has 48 GB of GDDR6 memory. The RTX 4090 has 24 GB of GDDR6X memory.
Can I find L40 and RTX 4090 GPUs available to rent right now?▾
Yes. This page shows real-time availability across 25+ cloud GPU providers. The Live Cloud Pricing section displays only in-stock offers with current pricing.
What is the main difference between the L40 and the RTX 4090?▾
The L40 uses the Ada Lovelace architecture (2023) while the RTX 4090 uses Ada Lovelace (2022). The L40 delivers 1.1x the dense FP16 throughput (181.05 vs 165.2 TFLOPS, both without sparsity) while the RTX 4090 has 1.2x the memory bandwidth of the L40.
Rent these GPUs
Each GPU page lists every current on-demand offer by provider, per GPU-hour, updated every minute.
Related comparisons
How this page is made
- Specifications come from the NVIDIA, AMD and Intel datasheets for the L40 and the RTX 4090. Dense and sparse throughput are listed separately.
- Prices and availability are live: every offer shown is an in-stock, on-demand listing from a provider we track, rechecked every minute.
- The written comparison (overview, performance notes, when to choose each card, verdict, use cases and FAQ) was drafted with an AI model from the specification table above and passed an automated check that rejects any figure not in that table. It was last generated on . It contains no prices; those are always read live.
- Read how we collect the data, or report an error on this page.
Next steps
- Rent L40Every current offer by provider, daily price history and a price alert.
- Rent RTX 4090Every current offer by provider, daily price history and a price alert.
- GPU price indexHow on-demand prices for the major GPUs have moved, updated daily.
- All GPUsEvery GPU we track, with current lows and provider counts.