The short version: If your product runs on DeepSeek, your inference bill changed this morning. DeepSeek's new peak/off-peak API pricing takes effect today, Aug 16: V4 Flash output jumps from $0.28 to $1.32 per million tokens at peak (a 136% to 371% increase depending on the window), and the new V4 Pro tier runs up to 14x the old Flash rate. That lands three days after Google halved Gemini 3.7 Flash. Same week, opposite directions. And xAI's Grok 4.6 slid in at $2/$6 to undercut the frontier on price. Here's what to re-check before you renew anything.

1. DeepSeek's price hike is live today — and it's steep#

On Aug 14, 2026, DeepSeek announced new pricing that takes effect today, Aug 16, moving to a peak/off-peak model the company says will "allocate resources more reasonably." The headline number: V4 Flash output rises from $0.28 to $1.32 per million tokens at peak, with an off-peak rate of $0.66 — a 136% to 371% increase depending on which window you hit. Input tokens climb from $0.14 to as much as $0.44 at peak (around $0.22 off-peak), a 57% to 214% jump. Alongside the increase, DeepSeek launched a higher-end V4 Pro tier priced up to 14x the old V4 Flash rate. Fortune framed the overall change as the API getting several times more expensive — a hard turn for the provider that built its name on being the cheap one.

What it means: This is a re-pricing event, and if you run DeepSeek in production you should treat it as one today, not next sprint. Three moves, in order. First, audit which of your calls truly need peak-hour capacity — most agent and batch work (embeddings refreshes, document ingestion, offline evals, overnight summarization) can shift to the off-peak window and roughly halve the new rate. Second, cap spend and wire alerts, because a 3-4x increase compounds fastest on the high-volume pipelines you stopped watching the moment they started working. Third, actually re-run the bake-off below — the cheapest provider for your workload last week may not be this week.

2. Grok 4.6 undercuts the frontier on price — but not on the agent loop#

Three days before the DeepSeek hike, xAI shipped Grok 4.6 on Aug 12, 2026. On the third-party Artificial Analysis Intelligence Index it scores 61 — matching GPT-5.6 Sol and trailing Claude Fable 5 by a single point — while pricing at $2 per million input tokens and $6 per million output, roughly 60% below GPT-5.6 Sol's $5/$30. It keeps the 500K-token context window and ships everywhere at once: Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare, with a Cursor-driven coding integration front and center. The honest caveat from the benchmarks: it's strongest on knowledge work and legal-style reasoning and weakest on terminal use — which is exactly the muscle an autonomous coding agent leans on hardest.

What it means: Grok 4.6 is a genuine price-to-intelligence play, and for reasoning-heavy knowledge work — research synthesis, drafting, analysis, classification — it belongs in your bake-off today. But don't mistake a cheap reasoning model for a cheap coding agent: its weak spot is the agent loop itself. If you're choosing what to run your coding workflow on, our AI coding agent ranking puts the harnesses in order, and the best LLM for coding roundup covers the model layer underneath — the two comparisons to read before you commit a team to one stack.

3. The real signal: pricing now moves both ways at once#

Line up the week and the pattern is unmistakable. On Aug 13, Google cut Gemini 3.7 Flash to $0.75/$3.75 per million — half its predecessor. On Aug 16, DeepSeek's up-to-4x hike takes effect. Inside four days, two major providers moved their prices in opposite directions. The comfortable assumption of the last two years — that inference gets cheaper forever, so you can defer the pricing question — is now false for at least one major provider. What replaces it is less convenient: per-token economics diverge by vendor, and you have to price each provider on its own curve. Keep a live bake-off, re-check the provider on your critical path every few weeks rather than every quarter, and remember that the underlying cost floor is set by hardware — our GPU rental price map tracks the H100/H200/B200 rates that ultimately decide how low any of these API prices can go.

Also on the wire#

The reason a provider raises prices into strong demand is usually the same reason: margin, ahead of the public markets everyone in this industry is now marching toward. Anthropic is on track for its first operating profit — a reported $559M in operating income on $10.9B of Q2 revenue, up ~130% quarter over quarter — even as the company cautions it won't sustain profitability. OpenAI, meanwhile, filed a confidential S-1 back in June that still hasn't surfaced publicly, with a listing reportedly targeting north of $1 trillion against an $852B last private round — while still losing roughly $1.22 for every $1 it earns. We covered the Anthropic IPO expectations earlier this week; the DeepSeek hike is the same story from the cost side. For a founder, the takeaway is simple: your model vendors are optimizing for margin and market debut, which historically means firmer pricing and less patience for below-cost tiers. Lock in any annual commitment you're happy with before that clock runs out.


Every figure in this edition is dated and linked, with at least two independent outlets per item. DeepSeek's exact per-window rates are drawn from its pricing announcement as reported by Engadget, InfoWorld, and Fortune; Grok 4.6's Artificial Analysis score is a third-party benchmark, and its "weak on terminal use" characterization reflects that same independent testing rather than a vendor claim.