The short version: Baseten closed a $1.5B Series F this summer at a valuation reported up to $13B — after being worth $5B in January. It says it now serves over a billion inference calls a day across 87 clusters on 18 clouds. The headline number isn't the story. The story is that an inference-only company is now worth $13B, which means "serve open models well" has become its own fundable infrastructure category — and that gives you a real third option for where you run models. Here's the build-vs-buy call it changes.
The takeaway up front#
For most of the last two years the choice was binary: call a closed API (OpenAI, Anthropic) and pay their per-token price, or rent GPUs and run vLLM or SGLang yourself. Baseten's raise is the clearest marker yet that a third layer has hardened in between — managed open-model inference — and it's now big enough that capital treats it as a category, not a feature. Your job isn't to pick a side. It's to route each workload to the cheapest layer that meets its quality bar.
What actually happened#
Baseten announced its Series F on June 22, 2026: $1.5B, at a valuation reported up to $13B, co-led by Altimeter Capital, Conviction, and Spark Capital, with Sands Capital and Wellington Management joining. That lands roughly five months after a $300M Series E that valued the company at $5B — a more-than-doubling of valuation in a single quarter of a year.
The operating numbers are the part founders should read closely. Baseten says it serves more than a billion inference calls a day across 87 clusters on 18 cloud providers, and third-party trackers put its annualized run-rate at roughly $200M in December 2025 rising to ~$600M by March 2026. Its named customers — Cursor, Mercor, OpenEvidence — are companies whose entire product economics ride on inference cost, and Baseten cites savings of up to ~30% versus closed-source APIs on the workloads that fit.
You can't read a $13B inference-only valuation as anything except this: serving other people's models is now a standalone business. Not a loss-leader a cloud runs to sell GPUs, not a feature a model lab tacks on — a category with its own leaders, its own competitive dynamics, and its own war chest.
Why a founder who isn't buying anything should care#
Because it changes the shape of a decision you make every week: where do I run this model call?
The old framing was two options. The real framing now is three:
- Closed API — OpenAI, Anthropic. Frontier quality, zero ops, instant start. You pay the lab's per-token price and you live inside their capability and their roadmap.
- Managed open-model inference — Baseten, Fireworks, Together, and peers. They host an open model (Llama, Qwen, DeepSeek, gpt-oss) on GPUs they run, behind an endpoint that's usually OpenAI-compatible, so you call it like any API. You get much of the cost advantage of open weights without running the server.
- Self-host on rented GPUs — CoreWeave, Lambda, Nebius, RunPod. Full control of the metal and the serving stack. Cheapest only if you keep the GPUs busy.
The middle option is the one this raise legitimizes. For a large slice of real work — classification, extraction, embeddings, transcription, retrieval, and a growing share of coding and chat — an open model is now good enough, and a managed endpoint lets you capture that without becoming an infrastructure team.
The call: price it three ways, per workload#
Here's the discipline. Don't decide once for the whole company. Take one workload at a time and price it three ways:
- The closed API you're using now — your current per-token bill.
- A managed open-model endpoint for a comparable open model at your token volume.
- Renting a GPU and serving the model yourself — but honestly, at your real request rate and utilization, not at 100% imagined load.
That third number is where most founders fool themselves. A rented H100 is only cheap if it's busy; the moment traffic is spiky or low, an idle GPU is the most expensive option on the board. The break-even between self-hosting and paying per token is a utilization threshold you can actually compute — we walk the math in Rent a GPU or Call an API, and what it actually costs to rent an H100, H200, or B200 gives you the current price map to plug in.
The common landing spot for a founder isn't a single answer. It's a split: a closed frontier API on the hard reasoning path where quality is non-negotiable, and managed open-model inference on the high-volume, cost-sensitive path where a good open model clears the bar. Same product, two layers, routed by what each request is worth.
Where this fits the bigger money story#
This isn't a lone data point. The summer's funding pattern has been capital piling into the layers underneath the models — control planes, runtime governance, and now inference itself — rather than into another chat app. A $13B inference company is that thesis in its purest form: whoever wins the next model war, someone has to serve the winner fast and cheap across clouds, and that job is more durable than any single set of weights.
The risk that pattern creates for you as a customer is lock-in — and the mitigation is baked into the choice. Because the model is open and the endpoint is standard, you can move it. When you shop the managed layer, that portability is the feature to protect: pick the vendors that keep the door open, and keep your own workloads priced and ready to re-route.
Bottom line: you don't need to care that Baseten is worth $13B. You need to care that its valuation means a real third option now exists between "closed API" and "run your own GPUs" — and that the cheapest place to run any given workload is now a question you should answer per workload, this week, with a spreadsheet instead of a habit.



