A GPU-hour is one GPU rented for one hour. An eight-GPU node running for one hour is eight GPU-hours, and so is one GPU running for eight hours. It is the unit almost every GPU cloud prices in, so if you can count GPU-hours you can estimate any job.
Here are current prices in that unit. Every price on this site is normalised to one GPU for one hour.
"Current" means an in-stock, on-demand offer, seen in the last 15 minutes, from a provider whose stock we track live. Stock is rechecked about every minute.
Per GPU or per instance: the normalisation trap
Providers publish prices in two ways. Some quote per GPU. Lambda's price list, for example, is per GPU per hour even for its eight-GPU instances. Others quote per instance. Replicate lists its multi-GPU hardware as bundles of two, four or eight GPUs, each with one price for the whole bundle. Hyperscaler instances work the same way: the AWS p5.48xlarge and the Google Cloud a3-highgpu-8g are each one machine with eight H100 GPUs and one hourly price.
Put those two numbers next to each other and the per-instance price looks eight times worse than it is. The fix is one division.
Say provider A lists an eight-GPU node at a price of X per hour, and provider B lists the same GPU at Y per GPU per hour. The numbers to compare are X ÷ 8 and Y. If you need all eight GPUs, you can also go the other way and compare X with 8 × Y. Either works. Mixing them does not.
That is why a comparison site has to normalise, and this one does. An eight-GPU instance in our data is shown at its price divided by eight, next to single-GPU offers, so every row in every table is a price per GPU-hour.
Two things survive normalisation, and you should keep them in mind:
- You may not be able to buy one. A per-GPU price taken from an eight-GPU node is only available if you rent all eight. The minimum purchase is the instance, not the GPU.
- The rest of the machine differs. A GPU-hour comes with some share of CPU, RAM and disk, and the share varies by provider. On Vast.ai, the docs say the GPU is exclusive to you while CPU, RAM and storage are a proportional share of the host.
The time unit needs the same care. Serverless platforms often quote per second, and Baseten quotes per minute. Multiply a per-second price by 3,600, or a per-minute price by 60, to get a price per GPU-hour. For example, Modal listed an H100 at $0.001097 per second on 21 September 2026. Times 3,600, that is about $3.95 per GPU-hour. Hold that thought, because Modal's number comes back in the section on what a price leaves out.
Billing increments by provider
The price is per hour. The meter usually is not. The increment decides what a short job costs: on a one-minute meter, a 20-second job pays for a minute.
Each row below is what the provider's own documentation or price page said on 21 September 2026.
| Provider and product | Metering as documented |
|---|---|
| RunPod Pods | "billed by the second for compute and storage" |
| RunPod Serverless | From when a worker starts until it fully stops, "rounded up to the nearest second". Start-up time and the idle timeout (default 5 seconds) are billed |
| Vast.ai | "Billed by the second for actual usage" |
| Lambda on-demand instances | "billed in one-minute increments", for as long as the instance runs, "regardless if they're actively being used" |
| Lambda 1-Click Clusters | Priced per GPU per hour, "billed in weekly increments" |
| Modal | Priced per second |
| Cerebrium | "billed by the second" |
| Beam | "billed by the millisecond" |
| Replicate | Priced per second |
| Koyeb | "accounted per second", with billing "rounded up to the nearest unit" |
| Baseten | "down to the minute" |
| Hugging Face Inference Endpoints | Shown by the hour, but "the actual cost is calculated by the minute" |
| Google Cloud Run GPUs | The whole instance lifetime, "with a minimum of 1 minute", rounded up to the nearest 100 milliseconds |
| Hetzner dedicated GPU servers | Hourly, up to a monthly maximum |
| Latitude.sh bare metal | Hourly, monthly or yearly |
The increment matters most for short, bursty work such as serving an API. For a training run that lasts hours, the difference between a second and a minute is noise. For short jobs, also check whether start-up time is billed: on RunPod Serverless it is, as the table shows.
One more rule to know before your first launch: RunPod's docs say you need at least one hour's worth of credits for your chosen configuration to deploy an on-demand pod. Per-second billing does not mean a per-second deposit.
What a GPU-hour price leaves out
The GPU-hour price covers the GPU while the instance runs. These are the usual extras.
Storage, including while stopped. Stopping an instance stops the GPU charge, not the disk charge. Vast.ai's docs say "Storage charges continue even when instances are stopped. Delete instances completely to cease storage billing." On 21 September 2026 RunPod charged $0.10 per GB per month for a running pod's volume and $0.20 per GB per month once the pod is stopped, so parking a pod doubles its storage rate. Lambda has no stop at all: instances "can only be launched, restarted, or terminated", and its persistent filesystems bill "as long as a filesystem exists, even if it's not mounted to an instance".
Data egress. Some providers charge nothing: RunPod documents "no fees for data ingress or egress", and Lambda says "you are not charged for ingress or egress". On Vast.ai, "Data transfer costs vary by host and include both upload and download traffic". Hyperscaler egress is a topic of its own. Our data egress reference has the details.
CPU and RAM on some serverless platforms. This is the one that catches people. Modal's price page lists CPU cores and memory as separate per-second lines, charged in addition to the GPU. Cerebrium, Beam and Google Cloud Run GPUs also price CPU and memory as separate lines. Koyeb goes the other way: its instance price includes vCPU, RAM and disk. So the $3.95 Modal figure above is a GPU-only price. Modal's page also applies a multiplier of 1.15 to 1.75 times for choosing a region and 3 times for non-preemptible execution, so the headline is the floor, not the bill.
Idle time you asked for. Keeping a serverless worker warm is billed. RunPod's docs say "Active workers incur charges continuously, including when idle."
Tax. Lambda's price list says prices are "plus applicable sales tax/VAT/GST".
The estimating formula
GPU-hours = number of GPUs × hours running
Compute cost = GPU-hours × price per GPU-hour
Hours means wall-clock hours the instance is up and billed, including setup, downloads and the time you spend debugging. It does not mean hours of useful compute.
A worked example. You plan to fine-tune on four GPUs, and a test run suggests six hours. That is 4 × 6 = 24 GPU-hours. Call the live price per GPU-hour P. The compute cost is 24 × P. Take P from the table at the top of this page for your GPU, and check that the provider behind it rents that GPU in groups of four.
Then add the extras:
- Storage: gigabytes × the monthly rate × the fraction of a month you keep the disk. Count the days after the job too, until you delete it.
- Egress: gigabytes you download × the provider's rate, which may be zero.
- A margin for reruns. First runs fail. Budget for a second attempt.
A bigger example, to show how fast the unit grows: an eight-GPU node for three days is 8 × 72 = 576 GPU-hours. At that scale, reserved and cluster pricing start to matter, and our GPU cluster page is the place to look.
For an always-on workload, the month is the natural unit. A month is 730 hours, so one GPU running all month is 730 GPU-hours. This table does the multiplication at the cheapest current H100 price:
| GPU | $/GPU-hr | Per day (×24) | Per month (×730) | Provider |
|---|---|---|---|---|
| H100 | $2.50 | $60.00 | $1,825 | Hyperstack |
If a monthly figure is what you are planning around, read about on-demand, spot and reserved pricing before you pay the on-demand rate for 730 hours in a row.
Five-line checklist
- Is the price per GPU or per instance? Divide by the GPU count.
- Is it per hour, per minute or per second? Convert to per hour.
- What is the billing increment and the minimum charge?
- What keeps billing when you stop: storage, warm workers, anything else?
- What is outside the price: egress, CPU and RAM, tax?
Answer those five and any two offers become comparable. Then the rule is simple: count your GPU-hours, multiply by the normalised price, add storage and egress, and pick the lowest total from a provider that has the GPU in stock. Start from the live H100 offers, the full provider list, or the daily GPU price index if you want to see how the price per GPU-hour has moved.
Sources
Source pages were retrieved with automated research tools on the dates shown and each figure was traced back to its source before publishing. Rental prices on this page are not typed: they are read live from GPUperhour's own data.
- RunPod pod pricing, accessed 21 September 2026
- RunPod Serverless pricing, accessed 21 September 2026
- RunPod Serverless endpoint configurations, accessed 21 September 2026
- Vast.ai pricing guide, accessed 21 September 2026
- Vast.ai instances overview, accessed 21 September 2026
- Vast.ai instance pricing, accessed 21 September 2026
- Lambda pricing, accessed 21 September 2026
- Lambda billing documentation, accessed 21 September 2026
- Lambda, creating and managing instances, accessed 21 September 2026
- Modal pricing, accessed 21 September 2026
- Cerebrium pricing, accessed 21 September 2026
- Cerebrium documentation, introduction, accessed 21 September 2026
- Beam pricing, accessed 21 September 2026
- Replicate pricing, accessed 21 September 2026
- Koyeb pricing, accessed 21 September 2026
- Baseten pricing, accessed 21 September 2026
- Hugging Face Inference Endpoints pricing documentation, accessed 21 September 2026
- Google Cloud Run pricing, accessed 21 September 2026
- Hetzner GEX131 product page, accessed 21 September 2026
- Latitude.sh pricing, accessed 21 September 2026
- AWS EC2 Capacity Blocks pricing, p5.48xlarge, accessed 21 September 2026
- Google Cloud Dynamic Workload Scheduler pricing, a3-highgpu-8g, accessed 21 September 2026