The short version: Pick serverless when your GPU would sit idle most of the day; pick a dedicated instance when it wouldn't. Serverless platforms (Modal, RunPod Serverless, Baseten, Beam/Fal, Replicate) bill per second and scale to zero, so idle costs nothing — but their per-second rate runs roughly 1.5–3x a cheap on-demand hourly rate, and every cold start adds seconds to minutes of latency. A dedicated or reserved pod bills the full hour regardless and is always warm. The tipping point is duty cycle: below roughly 30–60% of the day busy, serverless wins; above it, you're paying the premium on hours you'd have used anyway, and the pod wins.
The deciding variable is duty cycle, not sticker price#
Everyone opens the vendor pricing page and compares hourly numbers. That's the wrong comparison. A serverless platform and a dedicated pod are billing two different things: serverless charges for work done, a pod charges for time held. The only number that reconciles them is how many hours a day the GPU is actually busy — the duty cycle.
At 100% duty cycle, serverless is strictly worse: you're paying a premium rate for every hour, and you get cold-start risk on top. At 5% duty cycle, serverless is a rout: the pod bills 24 hours to do 1.2 hours of work. The interesting question is where the lines cross.
The break-even math#
The formula is one line:
break-even duty cycle = dedicated hourly rate / serverless effective hourly rate
Plug in verified August 2026 numbers. A specialty on-demand H100 runs about $2–4/hr (RunPod ~$1.99, GMI ~$2.00, Lambda ~$3.99 — see the price map). Serverless H100 lands at ~$3.95/hr effective on Modal ($0.001097/s), ~$4.55/hr on RunPod Serverless, ~$5.49/hr on Replicate, ~$6.50/hr on Baseten. So the serverless premium over a cheap on-demand card is roughly 2x to 3.3x, and over a mid-priced one (~$3.99 Lambda) closer to 1x–1.6x. That's the "1.5–3x" band, and it puts break-even between about 33% and 60% of the day.
Worked example. You run an agent backend that needs an H100 and is genuinely busy 6 hours a day (25% duty cycle).
- Dedicated at $2.50/hr: billed 24 hrs = $60/day ≈ $1,800/mo, GPU idle 18 hrs/day.
- Serverless (Modal, ~$3.95/hr) for 6 busy hrs: ~$23.70/day ≈ $711/mo.
Serverless saves ~$1,100/mo — you're not renting the 18 idle hours. Now push the workload to 16 busy hours/day (67%). Dedicated is still ~$1,800/mo (flat). Serverless becomes 16 × $3.95 × 30 ≈ $1,896/mo. The pod is now cheaper, and it never cold-starts. Break-even here sits at 24 × ($2.50 / $3.95) ≈ 15.2 busy hours ≈ 63%. Swap in a $1.99 on-demand card and break-even drops to ~50%.
The cold-start tradeoff#
Cost isn't the only axis — latency is the other. Scale-to-zero means the first request after idle has to spin up a container and load weights. That's anywhere from sub-200ms (RunPod FlashBoot, Modal memory snapshots) to 20–60 seconds for a true cold container pulling a large model into VRAM.
Two consequences. First, you may be billed for the cold start: RunPod Serverless bills the init window, and Replicate/Baseten private deployments bill setup and online time — so a 30-second cold boot on a 10-second job can triple that call's cost. (Replicate public models are the exception; cold starts there are free.) Second, the fix — keep-warm (min-replica ≥ 1) — quietly converts serverless back into a per-hour instance at a premium rate. Keep-warm is the right call for latency-sensitive p99, but price it as dedicated-plus, not serverless.
Serverless at a glance#
- RunPod Serverless — cheapest floor, FlashBoot cold starts, bills init time. Best for very spiky traffic.
- Modal — Python-first, snapshotting to cut cold starts, ~$3.95/hr effective H100. Best for bursty jobs and batch.
- Baseten / Replicate — managed model endpoints, lowest ops; pricier per hour and private deploys bill idle. Best when you don't want to run infra.
- Beam / Fal — lightweight, fast-boot serverless aimed at short generative calls.
Dedicated at a glance#
A reserved pod (CoreWeave, Lambda, Nebius, a RunPod on-demand pod) is the play for steady, high-utilization serving and training. No cold starts, predictable latency, and the cheapest per-hour dollar once you're keeping the card busy. The cost is that you pay for every idle hour and, on committed contracts, you're locked in. Our CoreWeave vs Lambda vs Nebius comparison covers picking one.
The decision#
Estimate your real duty cycle for a week. Under ~40% — spiky inference, a demo, an agent fleet with bursty traffic — go serverless and let it scale to zero. Over ~60% — a steady endpoint, a training run, anything you'd keep warm anyway — rent the pod. In between, let cold-start tolerance break the tie: if seconds of first-hit latency are fine, serverless; if not, the pod. For the per-task view of the same tradeoff on agent workloads, see what it costs to run a coding agent. Prices move weekly — the ratios and the duty-cycle rule don't.



