If you run an agent at any volume, the model underneath it is your single largest recurring bill, and this month the two most obvious places to put that workload both got refreshed. OpenAI shipped GPT-5.6 Luna on July 9 as the cheap tier of its Sol/Terra/Luna family; Google shipped Gemini 3.6 Flash on July 21 with an output price cut. They are aimed at exactly the same buyer — the founder running high-volume, latency-sensitive, "good enough" inference — and picking between them is one of the more consequential routing decisions you'll make this quarter.
Here's the short answer, up front: they are the same on capability, so choose on price and ecosystem. Default to Luna to spend less; pick Flash if you're already on Google's stack or you're latency-bound.
The capability gap is a rounding error#
The instinct is to reach for benchmarks and pick the "smarter" model. Don't bother — there's nothing to pick. On the independent Artificial Analysis Intelligence Index, which blends reasoning, knowledge, math, and coding into one number, Luna scores 51 and Gemini 3.6 Flash scores 50. That is a tie in any honest reading. Notably, Gemini 3.6 Flash landed on the same index score as the 3.5 Flash it replaces — the 3.6 release bought you lower cost and higher speed, not more intelligence.
So the entire decision collapses onto two axes that actually differ: what you pay, and how fast it streams.
Price: Luna wins, and it compounds#
Luna is priced at $1.00 per million input tokens and $6.00 per million output. Gemini 3.6 Flash is $1.50 / $7.50. That's roughly 33% cheaper on input and 20% cheaper on output for a model that scores a point higher on intelligence.
For near-identical capability, Luna is the cheaper token — and an agent spends tokens by the hundred-thousand.
The margin looks small per call and large per month. An agent loop makes hundreds of model calls per task; a support or research agent handling real traffic burns tens of millions of tokens a week. At that scale a 20–33% unit-cost difference is the difference between two line items you'd actually notice. If your only constraint is the bill, this is settled.
Speed: Flash wins, and latency-bound work feels it#
The counterweight is throughput. Gemini 3.6 Flash streams around 280 output tokens per second — the fastest of any model Artificial Analysis tracks — against roughly 200 tok/s for Luna. For a user-facing agent where someone is watching the response render, or a batch job where wall-clock time is the cost, that 40% throughput edge is real and felt.
One honest caveat on both: at high reasoning effort, time-to-first-token on each of these models runs high. Tokens-per-second describes the stream once it starts, not how long the user waits for the first character. Measure end-to-end latency on your own prompts before you treat Flash as categorically "faster" — it's faster once it's going.
The tiebreakers: output ceiling and which cloud you're in#
Two smaller factors decide the close calls. Luna allows 128K tokens of output per response versus Flash's ~65K, so if you generate long single-shot artifacts — a full report, a big code file — Luna's ceiling is the safer floor. And the boring one that decides most real deployments: the cloud you already live in. If your data, auth, and billing are on Vertex AI, Flash's 20% premium is often cheaper than the integration and second-vendor overhead of adding OpenAI. If you're already on the Responses API, the reverse holds.
The rule#
Same brain, different price and speed. Default to GPT-5.6 Luna to minimize the token bill. Switch to Gemini 3.6 Flash when you're already on Google's stack, when throughput is the binding constraint, or when you need Gemini-native grounding and multimodal in the same call. Don't switch expecting a smarter agent — that's not what either release delivered.
This is the same demand-side price war that's been squeezing the cheap tier all month — we mapped the wider routing picture in the model price-drop routing map and put Flash head-to-head with the open-weight challenger in Gemini 3.6 Flash vs Kimi K3. For the full week's moves, see the Founder's Wire for the week of July 28.



