Use on-demand until a GPU is busy for most of the month. Use spot only for jobs that checkpoint and can be stopped with seconds of warning. Reserve only the capacity you would have run anyway. That is GPU as a service in three sentences. The rest of this page covers what each model bills you for, how much warning a spot GPU gives you, and what a reservation binds you to.
Pick a model by your situation
| Your situation | Model to use | The catch |
|---|---|---|
| Prototyping, debugging, bursty inference, anything shorter than a few weeks | On-demand | Highest rate, and capacity can be sold out when you need it |
| Batch training, sweeps and offline inference that save checkpoints | Spot, preemptible or interruptible | Stopped with seconds of notice, and no service level agreement |
| A steady service or training programme that runs most hours for months | Reserved or committed use | You owe the money whether you use the GPU or not, often up front |
| A large run with a fixed start date, lasting days or weeks | Capacity block or calendar reservation | Paid in full up front, and AWS allows no cancellation |
| Price-sensitive single-GPU work that tolerates uneven hosts | Marketplace | Reliability, transfer charges and maximum rental length vary by listing |
| Multi-node training for a quarter or longer | Reserved cluster contract | Take-or-pay terms, often multi-year |
The gap between the models moves daily. The blocks below show the live spread between the cheapest on-demand, spot and reserved listing we track for the same GPU, starting with the H100.
| Pricing type | Cheapest listed $/GPU-hr | vs on-demand | Provider | Listing | Listings seen |
|---|---|---|---|---|---|
| On-demand | $2.44 | baseline | CoreWeave | 8 GPUs | 36 from 14 providers |
| Spot or preemptible | $0.63 | 74% lower | Vast.ai | 1 GPU | 9 from 6 providers |
| Reserved | $5.54 | 127% higher | Lambda Labs | 512 GPUs, 1-week minimum term | 12 from 1 provider |
The same spread for the A100, the older data centre card.
| Pricing type | Cheapest listed $/GPU-hr | vs on-demand | Provider | Listing | Listings seen |
|---|---|---|---|---|---|
| On-demand | $0.68 | baseline | LeaderGPU | 8 GPUs | 43 from 10 providers |
| Spot or preemptible | $1.24 | 84% higher | Massed Compute | 8 GPUs | 1 from 1 provider |
| Reserved | no reserved listing in the last 24 hours | ||||
And for the RTX 4090, a consumer card.
These are listed prices, not stock-checked ones. A line reading "no listing in the last 24 hours" means no provider we track published that model for that GPU in that window, not that nobody sells it. Reserved pricing is often quote-only: Lambda's pricing page says "Contact us for reserved capacity at our lowest prices".
On-demand
You pay for the time an instance exists, you can end it whenever you like, and nobody can take the GPU mid-job. Runpod's wording is that resources "cannot be displaced by other users". The billing unit differs. Runpod Pods are "billed by the second for compute and storage". Vast.ai charges GPU compute "per second while your instance is running". Lambda bills "in one-minute increments", for as long as an instance runs, "regardless if they're actively being used". (All read on 21 September 2026.)
What goes wrong is rarely the hourly rate.
- Stopped does not mean free. Runpod bills a stopped Pod's volume at twice the running storage rate. Vast.ai says "Storage charges continue even when instances are stopped." Lambda has no stop at all: instances "can only be launched, restarted, or terminated".
- You may not get the GPU back. A stopped Runpod Pod stays tied to its machine. If someone else rents that GPU meanwhile, your Pod "cannot start with a GPU".
- Capacity can be sold out. SemiAnalysis reported on 2 April 2026 that "On-Demand GPU rental capacity is sold out across all GPU types". On-demand is a price, not a guarantee of supply.
- The rate itself moves. AWS announced on 5 June 2025 a reduction of "up to 45 percent" on its NVIDIA GPU instances. From 1 June 2025, On-Demand prices fell 33% on P4d and P4de (A100), 44% on P5 (H100) and 25% on P5en (H200), with Savings Plan reductions of 25% to 45% for purchases after 4 June. The on-demand side of any reserved comparison can drop by that much in one announcement.
Our daily GPU price index tracks how on-demand rates drift.
Spot, preemptible and interruptible
A spot GPU, a preemptible GPU and an interruptible GPU are the same trade: spare capacity sold cheaply, which the provider can take back. AWS and Azure call it Spot. Google Cloud calls it Spot VMs and used to call it preemptible. Vast.ai calls it interruptible.
What differs is the warning and what happens to the machine. These are the vendors' own documented terms, read on 21 September 2026.
| Provider | Discount the vendor states | Warning | What happens to the instance |
|---|---|---|---|
| AWS Spot | "up to a 90% discount" against On-Demand | Interruption notice two minutes before | Stopped or terminated. No two minute warning if you use hibernation |
| Google Cloud Spot VMs | "up to 91%" for many machine types, GPUs and TPUs | Shutdown period is "best effort and up to 30 seconds" | No SLA. No maximum runtime (legacy preemptible VMs were capped at 24 hours) |
| Azure Spot | Not captured in our research | Evicted "with 30-seconds notice" | Deallocated by default, with disks still billed, or deleted. No SLA |
| Vast.ai interruptible | Docs say 50 to 80% savings | None documented in the pages retrieved | Stopped, "killing running processes", and it "may wait long to resume" |
Google Cloud's notice wording has changed. The current documentation gives a best effort shutdown period of up to 30 seconds, plus an optional 120 second preemption notice that is in Preview and defaults to 0 seconds. Google also says Spot prices "can change up to once every day".
Azure lets you set a maximum price. A value of -1 means you are never evicted for price.
Vast.ai's interruptible tier is an auction. Its docs say "You set a bid price" and the instance "Can be stopped by higher bids", while "On-demand instances always have highest priority". An outbid instance is paused, not deleted, so its storage keeps billing.
Runpod is a case we could not settle. A 5 second SIGTERM warning before SIGKILL is attributed to Runpod's Pod management docs, but we have it only second hand, and Runpod's pricing documentation on 21 September 2026 listed only on-demand and savings plans. Before planning around Runpod spot Pods, confirm they are still sold. Lambda's instance price list showed no spot tier on 21 September 2026.
Reserved and committed use
A reserved GPU trades flexibility for a lower rate. On demand vs reserved GPU is a question of how many hours you will use.
The hyperscalers sell one and three year terms. AWS Savings Plans commit you to "a consistent amount of usage (measured in $/hour)", with savings of "up to 66%" on Compute Savings Plans and "up to 72%" on EC2 Instance Savings Plans. Google Cloud committed use discounts offer "up to 55% discount off on-demand prices for most GPU types" and up to 65% for some. Azure Reserved VM Instances advertise "up to 72%" on one or three year terms. These are ceilings, read on 21 September 2026, and only Google's is stated for GPUs specifically.
Specialist clouds sell shorter terms. Runpod savings plans are "3 or 6 months upfront", "non-refundable", with "fixed expiration dates". Vast.ai reserved instances advertise "up to 50% discount based on commitment length", are prepaid, and have "Credits locked to the specific instance". Its API docs add that the "Maximum discount is typically 40%".
Large contracts are different again. CoreWeave's annual report for 2025, filed on 2 March 2026, describes its committed contracts as take-or-pay, "requiring payment regardless of the level of utilization", with initial terms that generally run one to six years and often involve prepayment. SemiAnalysis wrote in April 2026 that most of the rental market runs on contracts of six months or longer. Our GPU cluster pricing page covers reserved multi-node capacity.
A reserved product is not always a discount. On 21 September 2026, Lambda published per-GPU rates for its 1-Click Clusters (minimum 16 GPUs, terms from two weeks to one year) that were higher than its own on-demand 8-GPU instance rate, for both H100 and B200. Compare before assuming a reservation is cheaper.
Locking in a price can go either way. The SemiAnalysis index of one year H100 contracts rose almost 40% between October 2025 and March 2026, which rewarded early commitments. The AWS cut of June 2025 did the opposite.
Capacity blocks and calendar reservations
You book GPUs for a fixed window and pay for the whole window.
AWS Capacity Blocks for ML reserve GPU instances "on a future date to support your short duration machine learning (ML) workloads". Per AWS's documentation on 21 September 2026, a start time can be up to eight weeks ahead, a block can hold up to 64 instances, "The reservation fee is charged up front", and "Capacity Block cancellations aren't allowed". As a dated example, AWS listed a p5.48xlarge block in US East at $41.528 per hour, which it states as $5.191 per H100. AWS updates these prices "regularly based on trends in supply and demand", next in October 2026, so treat the figure as a snapshot.
One trap: AWS's docs say blocks end at 11:30 UTC, with instance termination starting at 11:00 UTC. Write your last checkpoint before then.
Google Cloud's Dynamic Workload Scheduler has two modes. Flex-start VMs "are charged based on usage, in the same way as on-demand instances". Calendar mode VMs "are tied to a reservation in which you pay for the full duration" and come only in 8 GPU shapes. Google's next price update is due in November 2026.
Marketplaces
A marketplace does not own the GPUs. Vast.ai's docs describe "a marketplace model where hosts set their own prices", billed by the second.
The risks sit in the listing. Vast.ai's own docs say higher reliability scores "typically correlate with higher prices". On-demand rentals there have a maximum duration set by the host, and "Expired instances may be deleted 48 hours after expiration." Host verification is "fully automated" and reflects machine health, not a security audit. "After spending credits, there are absolutely no refunds." Runpod's Community Cloud is a similar peer-to-peer layer, with reliability rated "Variable" in Runpod's docs.
Data transfer is the cost nobody budgets. On Vast.ai, "Data transfer costs vary by host and include both upload and download traffic." Runpod and Lambda both state they charge no ingress or egress fees. Our data egress reference compares the rest.
Checkpointing: the price of using spot
A spot discount is only real if the job survives being stopped. Each interruption costs the time since your last checkpoint, plus the restart, plus any wait for capacity to return.
- Checkpoint on a timer, not on the warning. Thirty seconds on Azure or Google Cloud is not enough to rely on for a large checkpoint. Treat the notice as a chance to stop cleanly. AWS recommends polling for its two minute notice every 5 seconds.
- Write to storage that outlives the instance. On Runpod, a container disk is lost when the Pod stops, a volume disk lasts until the Pod is deleted, and network volumes persist independently. Azure's default Deallocate policy keeps your disks and keeps billing them. Moving checkpoints to another provider is egress.
- Make resume automatic. The job should find the latest checkpoint and continue unattended. Test it by killing the job yourself.
- Set the interval from measurement. Time one checkpoint write. Pick an interval that keeps writing a small share of run time, and shorten it if interruptions are frequent. We give no number because the right interval varies with your model and storage.
- Keep the fragile parts on-demand. Serving endpoints with users waiting do not belong on spot.
What a commitment really commits you to
Ask these before you sign or prepay anything.
- What exactly is committed? A spend per hour (AWS Savings Plans), a specific instance (Vast.ai locks credits to it), or named hardware for a term?
- Is it take-or-pay? CoreWeave's filing says its committed contracts typically are.
- How much is paid up front, and is any refundable? Runpod savings plans are non-refundable and AWS Capacity Blocks cannot be cancelled. Lambda's billing docs say refunds come as service credits that expire 12 months after issue.
- Can you cancel, shrink, or move to a newer GPU mid-term? If the contract is silent, assume not.
- What happens if capacity is late or hardware fails? Get the service level and credits in writing.
- When exactly does it start and end, and how is it invoiced? Lambda bills 1-Click Clusters in weekly increments and gives ten days to pay for a reservation.
- What else is billed? Storage while idle, and egress.
- What is your break-even utilisation? At a 40% discount, the reservation beats on-demand only if you would otherwise have run more than 60% of the committed hours.
The rule
Start on-demand and measure real utilisation for a month. Move interruption-tolerant jobs to spot once automatic resume is tested. Reserve only the baseline that was busy for more than your break-even share of hours, on the shortest term that earns most of the discount. Leave peaks on on-demand or spot. Buy a capacity block only when the start date is fixed and the code is ready to run. For a long reservation, also run the rent vs buy calculator.
Sources
Source pages were retrieved with automated research tools on the dates shown and each figure was traced back to its source before publishing. Rental prices on this page are not typed: they are read live from GPUperhour's own data.
Every source was accessed on 21 September 2026. A date in an entry is the publication, filing or last-update date.
- Runpod docs: Pod pricing
- Runpod docs: Pods overview
- Runpod docs: Choose a Pod
- Runpod docs: Manage Pods
- Runpod docs: zero GPU Pods on restart
- Vast.ai docs: instance pricing
- Vast.ai docs: pricing guide
- Vast.ai docs: rental types FAQ
- Vast.ai docs: instance types
- Vast.ai docs index (API note on maximum reserved discount)
- Vast.ai docs: host verification
- Vast.ai docs: billing
- Lambda pricing
- Lambda docs: billing
- Lambda docs: creating and managing instances
- SemiAnalysis: The Great GPU Shortage, rental capacity, published 2 April 2026
- AWS blog: up to 45% price reduction for EC2 NVIDIA GPU instances, published 5 June 2025
- AWS docs: Spot Instance interruption notices
- AWS: EC2 Spot
- Google Cloud docs: Spot VMs
- Microsoft Learn: Azure Spot Virtual Machines, updated 24 June 2026
- AWS: Savings Plans pricing
- Google Cloud docs: committed use discounts
- Azure: Reserved VM Instances
- CoreWeave annual report for 2025 (Form 10-K), filed 2 March 2026
- AWS docs: Capacity Blocks for ML
- AWS: Capacity Blocks pricing
- Google Cloud: Dynamic Workload Scheduler pricing