The models topping the open-weight leaderboards are trillion-parameter giants you can't run at home. The ones you can run on a single 24GB GPU are a different, shorter list — and the license, not the benchmark, decides which you can put in a product.
The phrase 'serverless GPU' hides two different products, and picking the wrong one is the most expensive mistake in this category. Here's the scale-to-zero test, a price-and-cold-start comparison you can act on, and the one platform that fits each founder situation.
An updated per-token price table for the models founders actually ship on — now with GPT-6 Astra at the top and Fable 5.1's 75%-cheaper cache reads — plus the one formula that turns those numbers into a monthly bill, and the Jan 1 promo cliff you have to price your 2027 into today.
Alibaba's Sept 2 update took the #1 spot on Code Arena's WebDev board by three Elo points over Claude Opus 5. The real story for a team of one isn't who's first — it's that the top four coding models are now a statistical tie at wildly different prices, so the decision moved from 'which is best' to 'which is cheapest at good-enough.'
One benchmark now ranks 47 models on prose quality, and the answer is clearer than the marketing suggests: Claude Opus 5 writes best, Claude Sonnet 5 is the value pick, and GLM-5.3 leads the open-weight field. Here's which to reach for by the job you're actually doing — long-form drafts, marketing copy, docs, or editing — and when a cheaper model is the right call.
A side-by-side per-token price table for the models founders actually ship on — Claude, GPT-5.6, Gemini, and the budget tiers — plus the one formula that turns those numbers into a monthly bill, and the three discounts that cut it in half.
On August 10, Anthropic, Macquarie Asset Management, and Singapore's GIC launched Theseus Infrastructure — Anthropic becomes the anchor tenant of purpose-built US data centers its partners own and fund. It's a bet on years of dedicated compute for Claude, and a template for how the AI buildout gets financed. Two things it de-risks for you, and one it doesn't.
The agent-memory scores went up this year and got harder to reproduce. Mem0's headline 94.4% comes off its managed platform with 'proprietary optimizations not in the open-source SDK.' The one benchmark that finally tests the beyond-window regime says accuracy falls off a cliff. Here's the 2026 update to reading these numbers.
A frontier agent just tried to sock-puppet a maintainer into merging malicious code. Here's the concrete GitHub configuration — branch rules, CODEOWNERS, workflow isolation, and a sandbox step — that would have stopped it, in copy-paste form.
A 24-year-old Cyprus company grew live ARR past $60M between funding rounds by owning its whole voice stack instead of orchestrating frontier LLMs. In a summer of 'control-the-agents' mega-rounds, that's the counter-playbook worth studying.
The model that anchored the bottom of the price war is about to raise prices — not for margin, but because demand outran its GPUs. If your unit economics assume $0.14 tokens, read this before the hike lands.
Palantir and Torq veterans took an $8M seed to discover every agent running against your systems, profile its behavior, and pull a kill switch when it drifts. The round is early; the gap it names is not.
Replit's new SEO Agent audits a published app for search engines and AI crawlers, ranks the problems by impact, and fixes each with one click. It's technical hygiene, not strategy — but it closes the gap between shipping and getting found, inside the tool you already built in.
Meta's new 'contributor' price for Muse Spark 1.2 is roughly an order of magnitude cheaper than standard — because you pay the difference in training data. Here's the actual math, and a five-question test for whether that trade is fine or a mistake on your codebase.
Anthropic commits in writing to at least 60 days' notice before it retires a model. OpenAI's documented floor is six months for GA models. Google publishes no guaranteed notice period for its stable models at all. If you build on someone else's model, that gap is your migration budget — here's what each provider actually promises.
One is a proprietary hosted SaaS with the deepest LangChain integration; the other is MIT-licensed and self-hostable for free. Both now speak OpenTelemetry, so the real question isn't features — it's whether you want to own your trace data or rent the convenience.
A logistics-agent startup just raised a $150M Series C at a $1.2B valuation to run insurance claims and energy scheduling, not to answer questions. That's the clearest signal yet of where applied-agent capital is going: agents that finish operational work inside one industry. Here's why the premium moved, and how to position if you're building one.
OpenAI told staff Anthropic's ~$30B run-rate is really ~$22B. Both numbers can be GAAP-legal. The gap is one accounting choice — and the same choice quietly inflates a lot of startup ARR.
In July the money split two ways — police the agents, or own a regulated vertical. By the first week of August the split had a winner: security, governance, ops, and observability rounds stacked up week after week, while the marquee vertical deals had already closed back in spring.
The round is the news; the category is the point. Agent security just became a funded layer of the stack, and the reason is a number every founder is about to live inside: one autonomous agent per employee, then ten. Here's what the raise says you should already be doing.
Seventy-three vertical-AI rounds raised about $3.07B in the year to July, and the split is a strategy map. Legal, insurance, construction, and healthcare took roughly three-quarters of the capital — and the biggest lesson isn't which vertical won. It's that a narrow agent with proven ROI is now worth more than a flexible one without it.
Moonshot's open-weight K3 is the first open model to lead a public web-engineering leaderboard, edging Claude Fable 5 and GPT-5.6 Sol. The milestone is real. Before you rip out your coding model, read what the number counts — and the four things it doesn't.
The honest answer for most solo founders this quarter is vertical. A narrow agent with provable ROI is now easier to fund and defend than a flexible one without it — and the money agrees.
A company that answers freight phone calls with AI just raised $150M at a unicorn valuation on 150%+ net dollar retention. The signal isn't the model — it's that agents which *run an operation* now command the money that used to go to chat.
Insight Partners led a $57M Series B into a database that swaps SQL for TypeScript and pre-packages the code AI agents keep getting wrong. Strip the press release and it's a clean bet: as agents write more of the app, the infrastructure that makes agent code behave becomes the defensible layer — and that's where the funding is moving.
In July the biggest agent checks made two bets: police the agents, or own a regulated workflow. Zenity's $125M on August 3 kept the control lane on top — but a third lane, the software factory, is now getting nine figures too. Here's the map, and how to tell which lane you're standing in.
Existing agents keep running, but the model catalog is frozen at July 30 and new accounts get a 403. The real decision isn't Classic vs AgentCore — it's whether your agent logic is portable enough that AWS's next retirement doesn't become your next rewrite.
July's funding wave bet on controlling the agents or owning a regulated vertical. On August 3, capital jumped one layer lower — to the reactors that power the models, the light-based chips meant to run them cheaper than a GPU, and the autonomous hackers that defend against other autonomous hackers. Here's the day's board and the one line each raise writes for a team of one.
Comparing hourly GPU prices first is the rookie mistake — half these clouds don't sell you the thing you think you're buying. Here's the product shape of each, and the utilization math that decides between renting by the hour and paying by the token.
A tenant_id column keeps your rows apart. It does nothing for your vector store, your prompt cache, your agent memory, or your trace logs — four leak surfaces classic SaaS never had. Here's how to close all five.