The formula is one line: monthly cost = requests × (input_tokens × input_price + output_tokens × output_price) ÷ 1,000,000. Input and output are priced separately because output costs roughly 4–5× more per token than input on every frontier model — so your output-to-input ratio moves the bill more than the headline input price does. To use the formula you need two things: your token counts (convert words with 1 token ≈ 4 characters ≈ 0.75 words) and the per-1M-token prices. Below are verified September-2026 numbers, two worked examples, and the two discounts that cut the bill by an order of magnitude. If you just want a number for your workload, skip to the interactive calculator and plug it in.
Here's the whole method in one screen:
- Split the bill in two. Count input tokens (your prompt, context, and any resent history) and output tokens (what the model writes) separately, price each at its own rate, add them. Output is the expensive half.
- Turn words into tokens. ~1.33 tokens per word, or 100 tokens ≈ 75 words, ~4 characters per token. The newest Claude models (4.7+) tokenize ~30% heavier, so estimate on the exact model you'll ship.
- Multiply by volume. Cost per request × requests per month = your run rate. This is where a small per-request number becomes a real invoice.
- Then apply the two levers. Prompt caching bills a resent prefix at ~10% of the input price; the Batch API takes 50% off non-urgent work. They stack.
The one formula, spelled out#
Every LLM API bill is the same shape:
cost per request = (input_tokens × input_price ÷ 1,000,000)
+ (output_tokens × output_price ÷ 1,000,000)
monthly cost = cost per request × requests per month
Prices are quoted per million tokens (per 1M, or "MTok"). That's the whole thing. The only reason estimating feels hard is that the two inputs — token counts and prices — are moving targets, so let's pin both down.
Step 1: turn your workload into tokens#
You rarely know your token counts up front; you know roughly how much text goes in and comes out. Convert with the standard rule of thumb, straight from Anthropic's pricing docs: 1 token ≈ 4 characters ≈ 0.75 words in English. Inverted, that's ~1.33 tokens per word, or 100 tokens ≈ 75 words.
So a 500-word support reply is ~665 tokens; a 2,000-word doc you paste in as context is ~2,660 tokens. Two caveats that matter for a real estimate: the ratio shifts with language and content (code and JSON tokenize differently from prose), and the newest Claude models — 4.7 and later — use a tokenizer that produces about 30% more tokens for the same text. That improves quality but raises your count and your cost, so estimate on the tokenizer of the exact model you plan to ship, not an older one.
Step 2: use verified prices#
Prices change weekly across providers, so the only safe move is to read them off the provider's own page the day you commit budget. Here are Anthropic's, verified from the official pricing page on September 27, 2026, per 1M tokens:
The one genuinely useful price story this quarter: Sonnet 5's $2 / $10 rate — first announced as introductory — is now the standard price. The increase to $3/$15 that was on the calendar for September 1, 2026 did not happen, so the everyday workhorse got quietly cheaper than planned. (For a fuller cross-provider table including OpenAI, Gemini and DeepSeek, see our September LLM API pricing comparison — but confirm any number on the provider's page before you build a forecast on it.)
Step 3: worked examples#
Example A — a simple request, no caching. A summarization endpoint on Sonnet 5 ($2 in / $10 out) sends 8,000 input tokens and gets back 2,000 output tokens:
- Input: 8,000 × $2 ÷ 1,000,000 = $0.016
- Output: 2,000 × $10 ÷ 1,000,000 = $0.020
- Per request: $0.036. At 50,000 requests/month → $1,800/month.
Notice the output half ($0.020) beats the input half ($0.016) despite being a quarter of the tokens — that's the 5× output multiplier at work. Cut the answer length before you cut the prompt.
Example B — a multi-turn agent, with caching. An agent on Opus 5.5 ($4 in / cache read $0.20 / $20 out) resends a 20,000-token system prompt and context on every turn. That resent prefix is the trap: uncached, it's 20,000 × $4 ÷ 1,000,000 = $0.08 every single turn, before the model writes a word.
Turn on prompt caching and that prefix is billed at the cache-read rate — 5% of input on Opus 5.5, i.e. $0.20/1M:
- Cache read: 20,000 × $0.20 ÷ 1,000,000 = $0.004 per turn — twenty times cheaper.
- One-time 5-minute cache write (1.25× input, $5/1M): 20,000 × $5 ÷ 1,000,000 = $0.10, paid once.
So the write costs about the same as one uncached turn, and every turn after that runs at $0.004 instead of $0.08. Over a 10-turn session the prefix drops from $0.80 to ~$0.14. For any agent that resends context — which is most of them — caching is the biggest single lever on the bill.
Step 4: the two discounts, and when they stack#
- Prompt caching — a cache read costs 10% of the input price (5% on Opus 5.5, 2.5% on Fable 5.1). A 5-minute cache write is 1.25× input, a 1-hour write is 2×, so caching pays for itself after one reuse (5-min) or two (1-hour). Use it for any stable system prompt, tool schema, or document you send more than once.
- Batch API — 50% off both input and output for asynchronous, non-time-sensitive jobs (bulk classification, offline enrichment, evals). On Sonnet 5 that's $1/$5 instead of $2/$10.
- They stack. A cache-friendly bulk job run through the Batch API gets both discounts at once — the cheapest way to move a large, non-urgent workload.
Don't do this by hand — use the calculator#
The math above is simple, but a real forecast has caching ratios, batch splits, tool-call overhead, and a model choice per task — and doing that in a spreadsheet is how estimates drift. We built the interactive tool so you don't have to:
- LLM API cost calculator — plug in your input/output tokens, request volume, model and caching, and get a monthly number, with providers side by side.
- Agent cost calculator — for multi-turn agents, where cumulative resent context (not the per-call price) dominates the bill.
- All calculators — including VRAM and latency estimators if you're weighing an open model you host yourself against the API.
The takeaway is small and it saves real money: estimate the two halves of the bill separately, remember output is the expensive one, and cache anything you send twice. Get those three right and your API line item stops being a surprise.



