Serverless GPU wins when your GPU would sit idle most of the day, and a dedicated rental wins when it would not. The dividing line is one ratio: the dedicated hourly price divided by the serverless per-hour equivalent for the same GPU. If your workload keeps the GPU busy for a smaller fraction of the time than that ratio, go serverless. If it is busier, rent.
The decision rule
A dedicated GPU costs its hourly price whether it is busy or idle. A serverless GPU costs its per-second rate only while a worker is running. So:
Break-even busy fraction = dedicated hourly price ÷ serverless per-hour equivalent
Serverless per-hour equivalent = published price per second × 3,600 (or price per minute × 60)
Below the break-even fraction, serverless is cheaper. Above it, dedicated is cheaper. We do not give a fixed percentage, because both sides of the ratio move. Serverless list prices for the same GPU differ by more than a factor of two between vendors, as the table below shows, and dedicated prices change daily. Work it out with today's numbers. The method takes two minutes and is set out step by step further down.
One consequence is worth holding on to. The more a vendor charges per second relative to the dedicated market, the lower the break-even falls, and the quieter your workload must be for serverless to pay.
Published serverless GPU prices, 21 September 2026
Every price below was read from the vendor's own pricing page on 21 September 2026 and is shown in the unit the vendor publishes. The last column is our arithmetic, not a vendor figure, unless it says "published".
| Vendor | GPU | Published price and unit | Per-hour equivalent |
|---|---|---|---|
| Modal | L4 | $0.000222 per second | $0.7992 (arithmetic) |
| Modal | A100 80 GB | $0.000694 per second | $2.4984 (arithmetic) |
| Modal | H100 | $0.001097 per second | $3.9492 (arithmetic) |
| Modal | H200 | $0.001261 per second | $4.5396 (arithmetic) |
| RunPod Serverless, flex worker | 24 GB class (L4, A5000, 3090) | $0.69 per hour | $0.69 (published) |
| RunPod Serverless, flex worker | A100 80 GB | $2.72 per hour | $2.72 (published) |
| RunPod Serverless, flex worker | H100 80 GB (PRO) | $4.79 per hour | $4.79 (published) |
| RunPod Serverless, flex worker | H200 141 GB | $5.93 per hour | $5.93 (published) |
| Replicate | A100 80 GB | $0.001400 per second | $5.04 (published, matches arithmetic) |
| Replicate | H100 | $0.001525 per second | $5.49 (published, matches arithmetic) |
| Baseten | L4 | $0.01414 per minute | $0.8484 (arithmetic, × 60) |
| Baseten | A100 80 GiB | $0.06667 per minute | $4.0002 (arithmetic, × 60) |
| Baseten | H100 80 GiB | $0.10833 per minute | $6.4998 (arithmetic, × 60) |
| Cerebrium | L4 | $0.000222 per second | $0.7992 (arithmetic) |
| Cerebrium | A100 80 GB | $0.000583 per second | $2.0988 (arithmetic) |
| Cerebrium | H100 | $0.000944 per second | $3.3984 (arithmetic) |
| Cerebrium | H200 | $0.001166 per second | $4.1976 (arithmetic) |
| Beam | RTX 4090 | $0.000191667 per second | $0.69 (arithmetic, matches Beam's calculator) |
| Beam | H100 | $0.000972 per second | $3.4992 (arithmetic) |
| Google Cloud Run | L4, no zonal redundancy | $0.0001867 per second | $0.6721 (arithmetic) |
| Koyeb | H100 | $2.50 per hour | $2.50 (published) |
| fal | H100 | $4.50 per hour list price | $4.50 (published) |
These rows are not like for like. Read the notes before you compare them.
- Modal. The headline rates are for preemptible execution in any region. Modal's page lists "Non-preemptible execution 3x base prices", and says choosing a region multiplies base prices by 1.15 to 1.75. CPU and memory are billed on top of the GPU rate. The Starter plan includes "$30 / month free compute".
- RunPod. The pricing page shows flex workers per hour with a per-second toggle, and the docs say billing is "rounded up to the nearest second". Active worker prices are no longer published. The docs table lists active workers as "Discounts available through sales inquiry". Older articles that quote a fixed active-worker discount are out of date.
- Replicate. On public models "you only pay for the time it's active processing your requests". On private models and deployments you pay for all the time instances are online, including setup and idle time. Many official models are billed per output instead of per second.
- Baseten. Baseten publishes per-minute prices: "Only pay for the compute you use, down to the minute." Its page has an hourly toggle, but we could not read the hourly figures from the page itself, and the hourly numbers in circulation look derived. Quote the per-minute price. New accounts "come with credits"; no amount is stated.
- Cerebrium and Beam. Both bill CPU and memory separately. Both list a $0 plan billed on usage, and neither pricing page states a free credit amount.
- Cloud Run. The figure is the default region view, Tier 1. CPU and memory are extra, and an L4 needs a "minimum of 4 CPU and 16 GiB of memory". No committed use discount applies to the GPU line.
- Koyeb and fal. Both publish per hour. Koyeb's instance price includes vCPU, RAM and disk and is "accounted per second". fal also shows negotiated "as low as" rates, which we leave out because they are not list prices.
- Vast.ai Serverless has no list price. Vast's docs say it charges "the same price as Vast.ai's non-Serverless GPU instances", which hosts set.
How the meter runs
Serverless billing has four parts, and vendors differ on each.
The increment. RunPod rounds up to the second. Beam says it is "billed by the millisecond". Baseten bills to the minute. Cloud Run rounds to 100 milliseconds but bills each instance for "a minimum of 1 minute". Hugging Face Inference Endpoints show hourly prices and calculate cost by the minute.
The cold start. RunPod bills from when a worker starts, which covers container start and model load. Baseten bills "the time your model is actively deploying, scaling up or down, or making predictions." Replicate bills setup time on private models and deployments, and not on public models. Beam says: "We don't charge for the time to spin up a server or load your container image." Vast.ai bills GPU time for workers in the Loading state.
The idle tail. After a request finishes, the worker stays up for a while in case another arrives, and you pay for that time. RunPod's idle timeout defaults to 5 seconds. Modal's scaledown_window can be set "anywhere between two seconds and twenty minutes", and "you will be billed for any resources used while the container is idle". fal's keep_alive defaults to 60 seconds. Koyeb's default idle period for GPU instances is 5 minutes. Baseten's default scale-down delay is 15 minutes.
Idle and active workers. A minimum worker count above zero is a dedicated GPU by another name. RunPod: "Active workers incur charges continuously, including when idle." fal says a min_concurrency above zero "Costs money even with zero traffic". Replicate deployments with minimum instances bill while idle. If you need one always-on worker to hide cold starts, compare that worker against a dedicated rental at 100 percent, not against the break-even.
The live dedicated floor
This is the other half of the ratio: the cheapest dedicated on-demand price we track right now for the GPUs in the table above.
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| L4 | $0.90 | Scaleway | 1 |
| A100 | $0.67 | Vast.ai | 9 |
| H100 | $2.50 | Hyperstack | 9 |
| H200 | $3.43 | QuantaCloud | 7 |
These are whole GPUs billed for every hour the instance exists. Billing is finer than the hour: RunPod says Pods are "billed by the second" and Lambda bills "in one-minute increments". So a dedicated GPU does not have to run all day, every day. You can start it for the working day and end it at night. It does mean you manage that yourself, and you pay while it waits for requests. For the differences between on-demand, spot and reserved rates, see GPU pricing models. Per-provider prices for the H100, A100 and L4 are on their rent pages, and movement over time is on the GPU price index.
Working the break-even
Take the H100 on Modal as the example.
- Get the serverless per-hour equivalent. Modal publishes $0.001097 per second for an H100 (21 September 2026). Multiply by 3,600 and you get $3.9492 per busy hour. This is arithmetic on Modal's figure.
- Read the dedicated price. Call it D. Take it from the H100 row of the live table above. Use today's number from the table. Do not use a number you remember or one typed into an article.
- Divide. D ÷ 3.9492 is the break-even busy fraction.
- Turn it into hours. Multiply the fraction by 24 for busy hours per day, or by 730 for busy hours per month.
- Compare it with your traffic. Add up the time per day your workers would be billed, including cold starts and idle tails where the vendor charges for them, and convert it to hours. If that total is below the break-even hours, serverless costs less.
The shape of the result is easy to see without a price. If the dedicated price is half the serverless equivalent, break-even is 50 percent, or 12 busy hours a day. If it is a quarter, break-even is 25 percent, or 6 hours. A chatbot for a small team that is busy 2 hours a day sits well below either line. A batch pipeline that saturates the GPU for 20 hours a day sits well above both.
Now adjust for what you would really buy. If you need Modal's non-preemptible execution, the rate is 3 times base, so the H100 equivalent becomes $11.8476 per hour (arithmetic) and the break-even fraction falls to a third of what it was. Swap vendors and the denominator changes again: RunPod publishes $4.79 per hour for an H100 flex worker and Replicate publishes $5.49, so the same dedicated price gives a lower break-even on each. Run the division for each vendor on your shortlist.
What the math leaves out
Cold starts are a latency cost before they are a money cost. Vendor figures describe container or instance start: Modal "about one second", RunPod "sub-200ms FlashBoot cold starts" on its marketing page with no guaranteed figure in the docs, Cerebrium 1 to 3 seconds, Cloud Run "under 5 seconds". None of these include downloading the model and loading it into GPU memory, which is usually the larger part. The only end-to-end example we found is Google's: a time to first token of about 19 seconds for a gemma3:4b model from zero. Replicate says a cold boot "can take several minutes" in some cases. If your users cannot wait, you will keep a worker warm, and the break-even no longer describes your bill.
Model load time grows with model size. Google's 19 second example is a 4B model. Larger models take longer to download and load. We have no sourced load-time figures for them, so measure yours. The LLM VRAM calculator tells you how much has to be loaded.
Minimum workers. Covered above. Compare an always-on worker's per-hour equivalent with the dedicated price in the live table. Whenever the break-even fraction is below 100 percent, an always-on serverless worker at list rates is the dearer of the two.
CPU, memory and storage. Modal, Cerebrium, Beam and Cloud Run add CPU and memory charges to the GPU rate. RunPod Serverless charges for container disk and network volumes. On the dedicated side, RunPod and Vast.ai both keep billing storage for stopped instances.
Egress. Beam states "There are no egress or bandwidth fees on any plan." Koyeb charges "$0.04/GB transferred" beyond its allowance. RunPod says Pods have "no fees for data ingress or egress". Image and video generation moves far more data than text. Our data egress reference lists dedicated providers' fees.
Engineering time. A serverless platform scales workers up and down for you. With a dedicated GPU you build that or go without. Dedicated gives you a fixed machine and full control of the serving stack. Neither appears in the ratio.
The rule in practice
Run the division with today's numbers. If your billed seconds put you under half the break-even, choose serverless. The saving is large enough that the items left out of the ratio will not reverse it. If you are over the break-even, or you need a warm worker around the clock, rent a dedicated GPU and pick the cheapest suitable one from the live table. Between the two, start serverless, log billed seconds for two weeks, and redo the sum with real data. If you are not sure which GPU the model needs, start with our best GPU for LLM guide. If the workload is small enough for a free tier, see free cloud GPUs and credits.
Sources
Source pages were retrieved with automated research tools on the dates shown and each figure was traced back to its source before publishing. Rental prices on this page are not typed: they are read live from GPUperhour's own data.
All accessed 21 September 2026.
- Modal pricing
- Modal docs: cold start
- RunPod pricing
- RunPod docs: Serverless pricing
- RunPod docs: endpoint configurations
- RunPod Serverless product page
- RunPod docs: Pod pricing
- Replicate pricing
- Replicate docs: billing
- Replicate docs: how Replicate works
- Baseten pricing
- Baseten docs: cold starts
- Cerebrium pricing
- Beam pricing
- fal pricing
- fal docs: scale your application
- Koyeb pricing
- Koyeb docs: scale to zero
- Google Cloud Run pricing
- Google Cloud Run docs: GPU configuration
- Google Cloud blog: Cloud Run GPUs generally available
- Vast.ai docs: Serverless pricing
- Vast.ai docs: pricing
- Lambda docs: billing
- Hugging Face docs: Inference Endpoints pricing