Good morning. Three releases landed in roughly 48 hours, and they point the same direction: capability keeps sliding down-market and on-device. The hosted cheap tier got cheaper, open embeddings got small enough for a phone, and a 125-billion-parameter model got small enough for a gaming PC. For a team of one, the floor cost of doing real AI work just dropped on three fronts at once — and each cut has a footnote that eats part of the headline. Here's what matters.

1. Claude Haiku 5.5: $0.10 a million — read the footnote before you migrate#

Anthropic shipped Claude Haiku 5.5 (model id claude-haiku-5-5) at $0.10 per million input tokens and $0.50 per million output for prompts under 100K tokens — against Haiku 4.5's $1/$5. That reads as a ~90% cut. It isn't quite.

Two footnotes do the quiet work. First, the tokenizer changed: Anthropic's own docs say the same input text now produces about 30% more tokens on Haiku 5.5 than on Haiku 4.5 (Simon Willison measured ~1.25x on one long prompt; other tests ran higher on chat-style text). Because you're billed per token, a 90%-lower per-token price buys less than 90% once the token count climbs — which is exactly why Anthropic's headline figure is "~75% cheaper on average," not 90%. Second, there's a cliff: cross 100K tokens in a prompt and the rate jumps 5x, to $0.50/$2.50. Anthropic notes ~90% of Haiku 4.5 requests sat under that line, so most workloads stay in the cheap tier — but an agent that re-sends a big context every turn can wander over it.

The upside is real where it counts for agents: on Anthropic's own numbers, Haiku 5.5 posts large computer-use and terminal gains over 4.5 (OSWorld 2.1 offline subset 72.4% vs 15.7%; Terminal-Bench 4.0 39.2% vs 0.0%). Treat those as vendor-reported and not like-for-like across labs — Artificial Analysis's independent Terminal-Bench number came in lower, around 33%.

What it changes for you: if you run high-volume, low-stakes work — classification, extraction, summarization, cheap agent loops — re-pricing onto Haiku 5.5 is a genuine margin win. But do it on cost-per-completed-task over your own text, not the sticker, because the tokenizer gives some of the discount back. One operational gotcha before you flip the switch: computer-use now needs the computer_toolset_20260801 toolset, and non-default temperature/top_p/top_k values return a 400. It lands at the same $0.10/$0.50 as OpenAI's GPT-6 Luna, which turns "which cheap model?" into a real decision — we worked that one all the way through here.

2. EmbeddingGemma 2: multimodal retrieval that fits on a phone#

Google open-sourced EmbeddingGemma 2, a 740M-parameter, Apache-2.0 embedding model that maps text, code, images, audio, and video into a single shared space. It carries an 8K context window, Matryoshka outputs you can truncate from 768 dimensions down to 128, and a modular design that lets you load only the encoders a task needs — so a text-only build runs in a couple hundred megabytes of RAM and the full multimodal footprint still fits under about 0.6GB on a recent phone.

What it changes for you: this is private, on-device retrieval with no per-call API bill and no data leaving the device. "Search your stuff" — a user's documents, screenshots, voice notes, clips — becomes a feature you can ship without a hosted vector bill or a privacy review, under a license that lets you sell what you build. The benchmark claims are Google's own and untested in the wild as of this writing, so eval it on your own corpus before you rip out whatever you embed with today. But the direction is unambiguous: the embedding layer just stopped being a line item for a large class of apps.

3. Strata runs a 125B model on a 12GB gaming GPU — with a quality asterisk#

An independent, MIT-licensed engine called Strata (v0.1.39, dated Oct 4) claims to run Qwen3.8-Flash-Next — a 125-billion-parameter mixture-of-experts model — on a single 12GB consumer GPU. The trick is expert offloading: it keeps the hottest experts in VRAM, holds the full expert set in system RAM, pushes a lookup table to SSD, and uses a small draft model to speculate tokens the big model verifies. It installs with one script on Windows or Linux and serves OpenAI- and Anthropic-compatible APIs on localhost.

The asterisk is quality. The throughput figures in circulation (tens to ~120 tokens/sec) are the developer's claims, the footprint leans on aggressive 2–3-bit quantization, and early community notes flag accuracy trade-offs — plus a first-launch load that can freeze the machine for a minute or two. So: promising, not proven.

What it changes for you: the line between "model I can only rent" and "model I can run" keeps moving toward the desktop. If a frontier-adjacent open model can serve from a gaming PC — even at reduced quality — then local is a real cost hedge for privacy-sensitive or high-volume inference, not a hobbyist stunt. Pair it with a proper read on which open weights you're actually allowed to ship (the license-and-hardware map we published yesterday) and a view on what a 16GB card can realistically do before you commit a workload to it.


The one line to take into the week: the cost floor for serious AI dropped on the hosted, the open, and the local tier in the same 48 hours — but every one of those cuts shipped with a footnote (a tokenizer, an unproven benchmark, a quality trade). The discipline that wins this cycle is the same as last: price the workload, not the headline. Cheaper is only cheaper once you've measured it on your own tokens.

We left two widely-shared items out of this dispatch — a reported nine-figure agent-startup raise and a reported multi-billion-dollar compute financing — because we couldn't confirm the specifics against a primary source in time. No fabricated facts in The Wire: where we couldn't verify, we didn't print. Yesterday's edition is here.