The formula is one line: monthly cost = requests × (input_tokens × input_price + output_tokens × output_price) ÷ 1,000,000. Input and output are priced separately because output costs roughly 4–5× more per token than input on every frontier model — so your output-to-input ratio moves the bill more than the headline input price does. To use the formula you need two things: your token counts (convert words with 1 token ≈ 4 characters ≈ 0.75 words) and the per-1M-token prices. Below are verified September-2026 numbers, two worked examples, and the two discounts that cut the bill by an order of magnitude. If you just want a number for your workload, skip to the interactive calculator and plug it in.

Here's the whole method in one screen:

The one formula, spelled out#

Every LLM API bill is the same shape:

cost per request = (input_tokens  × input_price   ÷ 1,000,000)
                 + (output_tokens × output_price  ÷ 1,000,000)

monthly cost     = cost per request × requests per month

Prices are quoted per million tokens (per 1M, or "MTok"). That's the whole thing. The only reason estimating feels hard is that the two inputs — token counts and prices — are moving targets, so let's pin both down.

Step 1: turn your workload into tokens#

You rarely know your token counts up front; you know roughly how much text goes in and comes out. Convert with the standard rule of thumb, straight from Anthropic's pricing docs: 1 token ≈ 4 characters ≈ 0.75 words in English. Inverted, that's ~1.33 tokens per word, or 100 tokens ≈ 75 words.

So a 500-word support reply is ~665 tokens; a 2,000-word doc you paste in as context is ~2,660 tokens. Two caveats that matter for a real estimate: the ratio shifts with language and content (code and JSON tokenize differently from prose), and the newest Claude models — 4.7 and later — use a tokenizer that produces about 30% more tokens for the same text. That improves quality but raises your count and your cost, so estimate on the tokenizer of the exact model you plan to ship, not an older one.

Step 2: use verified prices#

Prices change weekly across providers, so the only safe move is to read them off the provider's own page the day you commit budget. Here are Anthropic's, verified from the official pricing page on September 27, 2026, per 1M tokens:

The one genuinely useful price story this quarter: Sonnet 5's $2 / $10 rate — first announced as introductory — is now the standard price. The increase to $3/$15 that was on the calendar for September 1, 2026 did not happen, so the everyday workhorse got quietly cheaper than planned. (For a fuller cross-provider table including OpenAI, Gemini and DeepSeek, see our September LLM API pricing comparison — but confirm any number on the provider's page before you build a forecast on it.)

Step 3: worked examples#

Example A — a simple request, no caching. A summarization endpoint on Sonnet 5 ($2 in / $10 out) sends 8,000 input tokens and gets back 2,000 output tokens:

Notice the output half ($0.020) beats the input half ($0.016) despite being a quarter of the tokens — that's the 5× output multiplier at work. Cut the answer length before you cut the prompt.

Example B — a multi-turn agent, with caching. An agent on Opus 5.5 ($4 in / cache read $0.20 / $20 out) resends a 20,000-token system prompt and context on every turn. That resent prefix is the trap: uncached, it's 20,000 × $4 ÷ 1,000,000 = $0.08 every single turn, before the model writes a word.

Turn on prompt caching and that prefix is billed at the cache-read rate — 5% of input on Opus 5.5, i.e. $0.20/1M:

So the write costs about the same as one uncached turn, and every turn after that runs at $0.004 instead of $0.08. Over a 10-turn session the prefix drops from $0.80 to ~$0.14. For any agent that resends context — which is most of them — caching is the biggest single lever on the bill.

Step 4: the two discounts, and when they stack#

Don't do this by hand — use the calculator#

The math above is simple, but a real forecast has caching ratios, batch splits, tool-call overhead, and a model choice per task — and doing that in a spreadsheet is how estimates drift. We built the interactive tool so you don't have to:

The takeaway is small and it saves real money: estimate the two halves of the bill separately, remember output is the expensive one, and cache anything you send twice. Get those three right and your API line item stops being a surprise.