The short version: Three moves landed inside 48 hours and they all cut the same way — running AI coding and agent workloads just got faster and cheaper. OpenAI previewed an "Ultrafast" tier that runs its flagship GPT-5.6 Sol at a claimed 750 tokens/sec on Cerebras hardware, up to 14x its Standard tier. Google shipped Gemini 3.7 Flash for coding and agents at half the price of 3.6 Flash — but only through year-end. And Zhipu dropped GLM-5.3, claiming the top open-weights coding model on its own benchmarks. If you priced your agent stack more than a month ago, that number is now stale. Here's what to re-check before the intro deals expire.

1. OpenAI's "Ultrafast" pushes its flagship to real-time speed on Cerebras#

On Aug 13, 2026, OpenAI and Cerebras Systems previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol — OpenAI's most capable model — at a vendor-claimed 750 output tokens per second, which OpenAI describes as up to 14x faster than its Standard tier. The tier runs the same model quality at Cerebras' wafer-scale speed, and OpenAI frames it for live and near-production tasks: voice, customer support, commerce, developer agents, financial research, and security response. The catch is that it launched as a limited preview open only to a select group of customers, with no published price, no confirmed general-availability date, and no model ID string yet.

What it means: If latency is what's blocking your voice or real-time agent product from feeling usable, this is your signal that frontier-quality inference at conversational speed is arriving — so join the preview waitlist and prototype the interaction now. But because there's no price and no GA date, do not rewrite your unit economics around it; a token that's 14x faster is worthless to a team of one if it turns out to cost 14x more. Keep your current model in production and treat Ultrafast as an R&D lane until the pricing sheet exists. For where the underlying compute costs are actually headed, our GPU rental price map tracks the H100/H200/B200 rates that ultimately set these API floors.

2. Google halves Gemini 3.7 Flash — but the discount has an expiry date#

Google launched Gemini 3.7 Flash on Aug 13, 2026, calling it its most intelligent workhorse model yet for coding and agents. The introductory price is $0.75 per million input tokens and $3.75 per million output tokens — roughly half the cost of Gemini 3.6 Flash, which shipped just three weeks earlier. Those rates are scheduled to rise to $1.50 input / $7.50 output on Jan 1, 2027. Google says the model posts gains over 3.6 Flash across software engineering, knowledge work, and web development, and its Gemini Spark personal agent now runs on 3.7 Flash in over 160 countries as of the same day.

What it means: The introductory pricing is a time-limited arbitrage, and a solo founder should treat it as one. Run Gemini 3.7 Flash against whatever model currently powers your coding assistant or agent loop this week — not next quarter — and if it wins on quality-per-dollar, schedule your high-volume, non-latency-sensitive batch jobs (embeddings refreshes, doc processing, eval runs) to lean on it through the back half of 2026 while the rate is halved. Then set a calendar reminder for December to re-price against the January hike. If you're deciding which model to standardize your coding workflow on, our best LLM for coding roundup and the Claude Code auto-mode default breakdown are the two comparisons worth reading alongside this.

3. China's GLM-5.3 claims the open-weights coding crown — with an asterisk#

Zhipu AI (Z.ai) released GLM-5.3 on Aug 14, 2026 through its GLM Coding Plan, claiming it's the strongest open-weights coding model available. The 743-billion-parameter model shares the same base as GLM-5.2, with the improvement coming entirely from extended post-training; on Zhipu's own evaluations it jumped to 28.3 on Terminal-Bench 3.0 from GLM-5.2's 4.6, and ranked first among open models on Terminal-Bench 3.0 and Agents' Last Exam — while still trailing closed models like GPT-5.6 Sol and Fable 5 on that same test. Two important caveats: the benchmarks are self-reported with no independently reproduced set yet, and the actual open weights are promised roughly two weeks after release rather than at launch. It arrives on the heels of Alibaba's Qwen3.8-Max open weights — a 2.4-trillion-parameter MoE with 95B active — landing on Hugging Face on Aug 12, so the open coding tier is thickening by the week.

What it means: The open-weights coding tier is closing on the frontier fast enough that a founder running any self-hosted or cost-sensitive coding pipeline should keep a live bake-off going rather than settling once a year. Add GLM-5.3 to that test queue — but don't rip out a working production pipeline on the strength of a vendor's own numbers. Wait for the weights to actually ship and for independent Terminal-Bench results before you migrate, and if you're weighing self-hosting economics, our CoreWeave vs Lambda vs Nebius comparison covers where a 743B model can actually run affordably.

Also on the wire#

The week's biggest financial headline is one we covered yesterday, and it kept moving: multiple outlets reported Aug 13-14 that Anthropic investors expect an October IPO at $2 trillion or more — which would be the largest listing in history, eclipsing SpaceX's June 2026 debut — with backers reportedly projecting $100-120 billion annualized revenue by year-end, up from the roughly $47 billion Anthropic disclosed in May. Treat the valuation as reported, not confirmed: investors told reporters executives have not finalized a target, even privately. For a founder, the signal isn't the headline number — it's that your primary model vendor may soon answer to public markets, which historically means firmer pricing and less patience for below-cost tiers. Lock in any annual commitments you're happy with before that clock starts.


Every figure in this edition is dated and linked to its source, with at least two independent outlets per item; the Anthropic IPO valuation is investor expectation, not a filing, and is marked "reported." Self-reported vendor benchmarks (GLM-5.3) are flagged as such and await independent replication.