The short answer: for raw writing quality in 2026, Claude Opus 5 is the best LLM for writing — it ranks #1 on the only large, independent prose benchmark that publishes its numbers. But "best" splits by the job you're doing: Claude Sonnet 5 is the value pick for everyday drafts and marketing copy, and GLM-5.3 is the strongest open-weight writer if you need to self-host. Here's the evidence, then the pick by task.
What the one objective source actually says#
Most "best LLM for writing" lists are vibes. There is one large, independent, reproducible benchmark that scores prose specifically: the lechmazur LLM Creative Story-Writing Benchmark, results updated August 23, 2026. It's worth understanding before you trust the ranking:
- 47 models each write short fiction that must incorporate 10 required story elements (character, object, concept, attribute, action, method, setting, timeframe, motivation, tone).
- Separate evaluator models read matched pairs of stories and pick the better one, across 67,178 judgments.
- Scores are relative — zero sits near the middle of the set, so a score is "how far above or below the field," not an absolute grade.
The top of the board:
| Rank | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (xhigh) | 4.2 |
| 2 | GLM-5.3 (max) | 3.5 |
| 3 | Claude Fable 5 (high) | 3.2 |
| 4 | Kimi K3 | 2.9 |
| 5 | GPT-5.6 Sol (xhigh) | 2.8 |
| 6 | GPT-5.5 (xhigh) | 2.7 |
Two things to read carefully. First, Claude Opus 5 leads by a clear margin — 4.2 vs. 3.5 for the runner-up is a real gap on this scale. Second, and more useful to a founder: the ordering is not "bigger and pricier is better." The top open-weight model (GLM-5.3) beats every GPT tier here, and some older flagship snapshots rank below cheaper mid-tier models. Which brings us to the caveat that should shape how you use any of this.
The caveat that matters more than the ranking#
This benchmark measures creative fiction, judged by other models, on relative scores. That's the strongest objective signal available for prose — but most founder writing isn't fiction. It's a launch post, a docs page, a cold email, ad copy, an investor update. For those:
- The brief you write (audience, voice, length, what to avoid) and one editing pass move the output more than switching from the value model to the flagship.
- The reasoning-effort tier you request changes quality visibly — the benchmark's own leaders are almost all running at high or "xhigh" effort. (If that lever is new to you, see reasoning effort vs. thinking budget.)
- A second independent benchmark, EQ-Bench Creative Writing, uses a different method (32 prompts, Claude Sonnet 4.6 as the judge) and broadly agrees that the Claude flagships and the top Chinese open models lead — useful as corroboration, not a second decimal.
So: pick a strong model, then spend your effort on the instructions and the edit. Here's the pick by job.
The pick, by the writing you actually do#
Best overall quality — Claude Opus 5. #1 on the benchmark, with a 1M-token input context that swallows a whole product spec or manuscript. Reach for it when the prose is the product: the launch essay, the landing page, the fundraising narrative. List rate is roughly $5 in / $25 out per million tokens (via the community LiteLLM cost map; confirm on the provider's page).
Best everyday value — Claude Sonnet 5. The same 1M-token context at roughly $2 in / $10 out per million — about a third of the flagship's price. This is the right default for the daily volume of a founder's writing: drafts, docs, marketing copy, summaries. You will not notice the quality difference on most of it, and you'll notice the bill.
Best open-weight / self-host — GLM-5.3 (Zhipu). #2 overall (3.5) and the top open model — the pick when you must self-host, keep data on your own hardware, or avoid a US API. Be clear-eyed: it's a frontier-size model, so "local" means a serious GPU, not a laptop. Genuinely small local models write well below this tier.
Best premium creative specialist — Claude Fable 5. A dedicated creative line, #3 (3.2), for the rare piece where prose quality is the whole point and budget is secondary — it's the priciest model here.
Best for high-volume cheap copy — a budget tier. When you need many first drafts rather than one perfect one — product descriptions, ad variants, internal notes — a cheap OpenAI or Gemini Flash tier, or an inexpensive open model, trades polish for a large cost cut. The full current price table across providers is in our LLM API pricing comparison; the method for reading a provider's pricing page is here.
Best for editing / rewriting — Claude Opus 5 or Sonnet 5. This one is a judgment call, not a benchmark result: no large editing-specific benchmark publishes numbers. Claude models are strong, well-calibrated prose evaluators (Sonnet is even used as the judge in EQ-Bench), which is exactly the skill editing needs — cutting, tightening, matching a voice.
The one-line rule#
Default to Claude Sonnet 5 for the daily writing, keep Claude Opus 5 for the pieces where the prose carries the weight, and reach for GLM-5.3 only when self-hosting is a hard requirement. Then stop model-shopping and start writing better briefs — that's the lever the benchmark can't rank, and the one that actually changes what you ship. Writing code instead of prose? That's a different leaderboard — see the AI coding-agent ranking.



