If you rent GPUs by the hour, the single most expensive mistake in 2026 isn't picking the wrong card — it's paying a hyperscaler's rate for a card a specialty cloud rents for a fifth of the price. The spread is now roughly 5–7× for the identical H100. Here's the published map, and the small number of things that decide where you actually belong on it.
If you read one line: on-demand H100 is ~$2–4/GPU-hour on specialty clouds versus a ~$14/hr hyperscaler median; the decision that saves you the most money isn't which provider — it's whether your GPU is busy enough to rent one at all.
The price map (published on-demand rates, early August 2026)#
Every figure below is a published on-demand rate, gathered the first week of August 2026 from public pricing trackers. GPU pricing moves weekly with supply, region, and commitment — read this for the shape of the market, then confirm on the provider's own page before you commit a dollar.
- H100 (80GB) — the workhorse. Roughly $1.99–$4/hr on specialty clouds: ~$1.99 at RunPod, ~$2.00 at GMI Cloud, ~$3.99 at Lambda. Spot dips to about $1.20/hr where offered. Hyperscaler median: ~$13.96/hr — the same card, 5–7× the money.
- H200 (141GB) — more memory, bigger models on one card. About $2.60–$6.31/hr on-demand: ~$2.60 at GMI, ~$5.29 at Lambda, ~$5.50 at Nebius, ~$6.31 at CoreWeave. The full cross-provider range runs roughly $2.29 to $13.78/hr.
- B200 (Blackwell) — the frontier card. Published on-demand is ~$4.99/hr at Lambda, ~$5.50/hr at Nebius (SXM6), and roughly $6.50/hr at CoreWeave (whose B200 is largely contract-oriented, not a public on-demand rate). Spot B200 has appeared near $2.12/hr.
The same H200, run continuously for a month, costs about $2,664 more at CoreWeave (~$6.31/hr) than at GMI (~$2.60/hr). That's not a rounding error — that's the difference between two providers offering the identical silicon.
Why the spread is so wide#
It isn't margin gouging — it's three different businesses wearing the same "GPU cloud" label.
- Discount specialty clouds (GMI, RunPod, Vast.ai) optimize for cheap, no-commitment, by-the-hour access. You get the lowest sticker and the fewest guarantees.
- Reserved/contract clouds (CoreWeave, and to a degree Nebius) are built for large, committed allocations. Their on-demand rate is deliberately high because on-demand isn't their product — reservations are.
- Hyperscalers (AWS, GCP, Azure, Oracle) bundle the GPU with a full platform, compliance surface, and enterprise support. You pay ~$14/hr for an H100 because you're not really buying the H100 — you're buying everything around it.
For a solo founder or small team renting one or two cards, the discount column is almost always the right one. The hyperscaler premium only pays off when you genuinely need the surrounding platform.
The three numbers that actually decide your bill#
Sticker price is the distraction. These three decide what you pay:
- Utilization (duty cycle). A rented GPU bills 24/7 whether or not it's inferring. At 100% utilization, $2/hr is cheap. At 10%, you're paying $2/hr for a card that's idle 90% of the time — and a per-token API would have cost you a fraction. Below roughly 40–50% duty cycle, stop renting and call an API. (We work the break-even the other direction in Rent a GPU or Call an API?.)
- Commitment tier. On-demand is the most expensive way to rent. If your load is steady, a monthly or annual reservation on the same card can cut the rate substantially — the trade is flexibility for price.
- The card, matched to the model. Don't rent a B200 to serve a 70B model that fits comfortably on an H100. Match memory and throughput to the workload; see B200 vs H200 vs H100 for LLM inference for which card each model class actually needs, and our CoreWeave vs Lambda vs Nebius breakdown for a closer look at the three you'll compare most.
The takeaway#
The GPU market in 2026 rewards two decisions and punishes their absence. First, skip the hyperscaler on-demand rate unless you're buying the platform around the card — the specialty clouds rent the same silicon for a fifth to a seventh of the price. Second, rent by utilization, not by sticker — a cheap GPU you can't keep busy is more expensive than an API, every time. Get those two right and the exact provider is a rounding error.



