In 48 hours this week the ceiling and the floor both moved. OpenAI shipped GPT-6 Astra — the first model it has ever rated "Critical" for cyber capability — as a gated preview. Google's Gemini 3.8 Flash and Microsoft's MAI-Transcribe-2 reset the workhorse and transcription floors on price. Here's the whole edition in one screen:

The through-line for a team of one: prototype against the ceiling, route the bulk of your traffic to the floor — and price your 2027 plan on the post-promo numbers, because two of this week's three cheap deals reset upward on Jan 1. Here's what each changes.

1. OpenAI's GPT-6 Astra is the first model it has ever rated "Critical" for cyber#

On Sept 3, 2026, OpenAI began rolling out GPT-6 Astra, which CEO Sam Altman told CNBC represents a "new capability level." The headline for founders isn't the benchmark score — it's the safety designation. Astra is the first model OpenAI has ever placed in the "Critical" tier of its Preparedness Framework, after finding it can autonomously discover previously unknown security weaknesses and build functional exploits against well-defended systems without a human directing every step. Because of that, access is staged: vetted participants in an application-based cybersecurity program get it first, then ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, and AWS. Reported API list pricing is roughly $10 per 1M input tokens and $50 per 1M output (cached input ~$1, batch ~half, a "Fast" mode at ~2x), against a context window near 1.05M tokens — figures worth confirming on OpenAI's own pricing page before you wire them into a model.

What it means: Read this as two events, not one. First, a new capability ceiling — a model this strong at software engineering and reasoning is worth prototyping against for the handful of tasks that genuinely need the frontier, while you keep everyday traffic on cheaper tiers (the same discipline we laid out in the AI coding agent ranking and the agent model price map). Second, and more urgent: the same class of model that can autonomously find and exploit vulnerabilities is now shipping, and attackers get their own versions. That makes this week the right time to revisit the boring hygiene — least-privilege agent permissions, dependency and MCP-server review, and PR poisoning defenses. If you run coding agents against your repo, our guide to hardening your repo against poisoned PRs is the checklist to run now, and the wave of agent-security funding we covered yesterday is the market pricing in exactly this risk.

2. Gemini 3.8 Flash is the cheap workhorse — until Jan 1#

On Sept 2, 2026, Google launched Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output — introductory pricing through Dec 31, 2026. Google positions it as its most capable "workhorse" model, tuned for long-horizon coding and autonomous agents, and says it beats its own 3.7 Flash on every published benchmark and Claude Opus 5 on three. It's live in Google Antigravity, AI Studio, the Gemini API and Android Studio, plus the Gemini app, AI Mode in Search, and Sheets. The catch is in the pricing footnote: on Jan 1, 2027, the rate doubles to $1.50 / $7.50.

What it means: For cost-sensitive agent and coding workloads, Flash is an obvious place to route traffic today — near-frontier capability at a fraction of a flagship's per-token cost. But "cheap" here has an expiry date. If you standardize an agent product on 3.8 Flash this quarter, your token bill for the same usage is 2x higher four months from now, and that lands right as your Q1 2027 numbers get scrutinized. Put the doubling in your forecast, and keep your routing layer model-agnostic so you can rebalance in December — the same "route on cost-per-completed-task, not sticker price" logic we walked through when OpenAI cut Luna and the ranking barely moved and in the budget-tier price-war breakdown.

3. Microsoft's MAI-Transcribe-2 puts transcription at 10 cents an hour#

Also on Sept 3, 2026, Microsoft AI released MAI-Transcribe-2 at $0.10 per hour of audio — an introductory rate through the end of 2026. Microsoft claims it ranks #1 on the FLEURS benchmark across 60 languages at about a 5.2% average word error rate, adds speaker diarization, configurable styles and word-level timestamps, and processes long recordings 5-10x faster than Gemini 3.5 Transcribe, GPT-Transcribe and ElevenLabs' Scribe v2. It's in public preview via Azure AI Foundry.

What it means: Transcription is a quiet cost line in a surprising number of products — meeting-notes tools, call and sales-call analytics, podcast and media workflows, voice interfaces. A credible ten-cents-an-hour option resets what's economically buildable: at that price, features that couldn't clear their own transcription cost suddenly can. Two cautions keep it honest. The benchmark and speed claims are Microsoft's own, so pilot on your real audio before you believe the WER on your accents and domains. And the $0.10 rate is introductory with no published 2027 price — so keep the transcription step in your pipeline provider-swappable, and don't design a business whose margins only survive at ten cents. This is the same "build on a floor, but don't bet the company on a promo" caution that runs through the GPU rental price map: cheap compute is a tailwind, not a foundation.

Also on the wire#

The pattern under all three moves is worth naming on its own: the cheapest numbers announced this week are the ones with the shortest shelf life. Gemini 3.8 Flash doubles on Jan 1; MAI-Transcribe-2's intro rate ends with the year; even Astra's Fast mode carries a 2x multiplier. Vendors are competing hardest on the first impression of price while reserving the right to reprice in Q1. For a solopreneur, the defense is structural, not clever: keep every model and API call behind a thin routing layer you control, so switching a provider is a config change, not a rewrite — the difference between the founders who ride each price war and the ones who get repriced by it. It's the same lesson the local-LLM-for-coding route makes concrete: the cheapest per-token cost of all is the one you host yourself and nobody can double on January 1.


Every figure in this edition is dated and linked to a primary or major-outlet source, with independent corroboration across multiple outlets per story. GPT-6 Astra's API pricing (~$10/$50 per 1M tokens) and ~1.05M context window are reported figures widely cited across outlets; confirm them against OpenAI's own pricing page before committing. MAI-Transcribe-2's FLEURS ranking, 5.2% WER, and speed multipliers are Microsoft's own published claims. Gemini 3.8 Flash's benchmark comparisons are Google's own. The Jan 1, 2027 price changes for Gemini 3.8 Flash, and the end-of-2026 expiry of MAI-Transcribe-2's introductory rate, are as announced by each vendor.