If you rent GPUs by the hour, the single most expensive mistake in 2026 is still the same one: paying a hyperscaler's rate for a card a specialty cloud rents for a fifth of the price. The spread is ~5–7× for the identical H100, and a month of price moves didn't close it. What did change: the Blackwell floor finally dropped below $4, the Grace-Blackwell superchips started renting by the hour — and, against the trend, AWS raised its prices. Here's the September map.

If you read one line: on-demand H100 is ~$2–4/GPU-hour on specialty clouds versus a ~$10–13/hr hyperscaler median; the B200's on-demand floor just fell under $4; and the decision that saves you the most money still isn't which provider — it's whether your GPU is busy enough to rent one at all.

The price map (published on-demand rates, early September 2026)#

Every figure below is a published on-demand rate, gathered the first week of September 2026 from public pricing pages and trackers. GPU pricing moves weekly with supply, region, and commitment — read this for the shape of the market, then confirm on the provider's own page before you commit a dollar.

The headline of the month is a split screen: the cheap end got cheaper (B200 floor $4.99 → $3.70) while the expensive end got more expensive (AWS raised H200 15%). The market isn't moving one direction — it's spreading apart.

What changed since August#

If you read last month's map, here's the delta in four lines:

  1. H100 floor: flat. ~$2 at GMI/Spheron, ~$1.49 at the Vast marketplace floor. But brand-name specialty on-demand firmed — Lambda's SXM rate climbed toward $3.99, so RunPod's ~$2.69 is now a more representative "cheap-but-reliable" number than August's headline $1.99.
  2. H200: stable to softer. GMI still ~$2.60, Nebius eased to ~$4.50, CoreWeave ~$6.31. No dramatic move.
  3. B200: floor down, availability up. The Blackwell on-demand floor fell from ~$4.99 to ~$3.70–4.00 and the provider count roughly doubled. This is where a founder's dollar moved the most.
  4. The hyperscalers went UP. AWS raised H200 capacity-block prices ~15% on January 4, 2026 (Data Center Dynamics) — its first GPU price increase in roughly two decades. The premium for renting from a hyperscaler is widening, not narrowing.

The through-line ties back to where the compute money is going: capital keeps pouring into new data-center capacity (Crusoe just raised $3B at a $30B valuation), and that new supply lands first and cheapest on the neoclouds — which is exactly why their floor keeps dropping while the hyperscalers, selling committed capacity and platform, can raise a specific part's price into tight demand.

Why the spread is so wide#

It isn't margin gouging — it's three different businesses wearing the same "GPU cloud" label.

For a solo founder or small team renting one or two cards, the discount column is almost always the right one. The hyperscaler premium only pays off when you genuinely need the surrounding platform. See our CoreWeave vs Lambda vs Nebius breakdown for a closer look at the three you'll compare most.

The three numbers that actually decide your bill#

Sticker price is the distraction. These three decide what you pay:

  1. Utilization (duty cycle). A rented GPU bills 24/7 whether or not it's inferring. At 100% utilization, $2/hr is cheap; at 10%, you're paying $2/hr for a card that's idle nine-tenths of the time, and a per-token API would have cost a fraction. Below roughly 40–50% duty cycle, stop renting and call an API. We work the break-even the other direction in Rent a GPU or Call an API?, and the LLM VRAM & self-host calculator tells you whether a model even fits the card before you rent it.
  2. Commitment tier. On-demand is the most expensive way to rent. If your load is steady, a monthly or annual reservation on the same card can cut the rate substantially — the trade is flexibility for price. (The AWS increase applied to committed capacity blocks, so even reservations aren't a one-way street anymore — confirm the rate before you sign.)
  3. The card, matched to the model. Don't rent a B200 to serve a 70B model that fits comfortably on an H100 — even at a sub-$4 floor. Match memory and throughput to the workload; B200 vs H200 vs H100 for LLM inference covers which card each model class actually needs.

The takeaway#

The GPU market in September 2026 rewards two decisions and punishes their absence. First, skip the hyperscaler on-demand rate unless you're buying the platform around the card — the gap is 5–7× and, after January, widening. Second, rent by utilization, not by sticker — a cheap GPU you can't keep busy is more expensive than an API, every time, and that's as true of a $3.70 B200 as it was of a $2 H100. Get those two right and the exact provider is a rounding error.