Token Counter
Estimate token counts across GPT-4, Claude, Llama, Gemini, and other LLMs. Paste text, see counts per model side-by-side. Files never leave your browser.
| Model | Tokens | Input $/1M | Cost (this text) |
|---|
Why token counts vary by model
Every interaction with an LLM is metered in tokens — sub-word units that the model's tokenizer carves from your text. Tokens drive both context-window limits ("does this fit?") and pricing ("how much will this cost?"). The exact count depends on the model's tokenizer, which you usually don't have at hand. This tool gives you a fast estimate for every major model side-by-side, plus the dollar cost for the input text against each model's published per-token price.
Budget, cost, and context-window planning
- Sizing a system prompt to make sure it (plus expected user input and headroom for output) fits the context window.
- Estimating the API cost of a batch job before you run it — paste 100 representative inputs and multiply.
- Comparing prompt-engineering iterations: did the new prompt actually get shorter, or did you just feel like it did?
- Sanity-checking that a "this is too long" error wasn't caused by hidden whitespace, BOM markers, or copy-paste noise.
Heuristic accuracy vs real tokenizers
- These are heuristics, not the real tokenizers. tiktoken (OpenAI), Anthropic's tokenizer, and SentencePiece (Llama, Gemini, Mistral) each carve text differently. For English prose, our estimates land within ±5%. Code, dense JSON, and CJK text drift to ±10% or worse.
- Why we don't ship the real tokenizers. tiktoken alone is ~1 MB of WASM + data files; loading it just to count tokens would bloat the page tenfold. If you need exact counts (e.g. you're hitting a hard 8k context limit), run
tiktokenin Python locally or call the model's/v1/tokenizeendpoint. - What we get right. Relative ordering (which model uses more tokens for the same text) is reliable. Cost rankings are usually accurate. Order-of-magnitude estimates ("is this 500 or 5000 tokens?") are dead-on.
Pricing and the output multiplier
- Costs are input prices per 1M tokens as of 2025. Output tokens are typically more expensive — multiply by 3–5× for a worst-case output estimate.
- Llama 3 shows zero cost because the typical deployment is self-hosted. Hosted offerings (Together, Groq, Fireworks) charge $0.20–$1 per 1M depending on size.
- Prices change. Check the provider's pricing page before relying on these numbers for a real budget.
Tokens, words, whitespace, and CJK
- Tokens ≠ words. An English word averages 1.3 tokens; "antidisestablishmentarianism" is roughly 7. Code and structured text tokenise much higher per character.
- CJK text is dense. Each Chinese / Japanese / Korean character can be its own token, so 1000 chars ≈ 1000 tokens — much more expensive per "character" than English.
- Hidden characters add up. Pasted text with zero-width joiners, NBSPs, or BOMs gets counted too. Use the Unicode Inspector tool if your token count looks suspiciously high.
- System + user + assistant tokens compound. The "context window" budget includes every message in the conversation. Don't size your input against the raw limit; leave 30–50% headroom for replies.
Sizing a batch job
Paste a 500-word prompt and the tool estimates roughly 650–700 tokens — English runs about 1.3 tokens per word — then multiplies by each model's published input price so you can see the cost per model side by side. Multiply that by the 100 inputs in your batch job and you know the bill before you spend a cent.
Accuracy, output billing, and whitespace traps
How accurate is the count? It's a heuristic, not the model's real tokenizer, so treat it as sizing rather than billing truth — usually within 10–20% of what tiktoken or SentencePiece would report.
Why do models disagree on the count? Each model carves text with its own tokenizer, so the same string splits into a different number of sub-word units — and therefore a different price.
Does this include output tokens? No — it measures your input text only. Output is billed separately and usually at a higher rate, so leave headroom in your estimate.
Do spaces and line breaks count? Yes — whitespace becomes tokens too. That's why a stray BOM marker or copy-paste noise can quietly inflate a count and trip a context-window limit.
Chars, tokens, words — the rough conversions that matter for cost
LLMs bill by token, not word, and a token is roughly a common word-chunk of about four characters of English. That gives usable back-of-envelope conversions, but the ratio shifts hard by content type:
| Content | Rough chars/token | 100 tokens ≈ |
|---|---|---|
| Plain English prose | ~4 | ~75 words |
| Code / JSON | ~3 (denser) | Far fewer “words”; braces & punctuation each cost a token |
| Rare words, other languages, emoji | <2 | One emoji or CJK glyph can be several tokens |
Two traps: (1) whitespace and punctuation are tokens too, so minifying JSON genuinely cuts cost; (2) the same prompt tokenises differently across model families — treat any “words × 1.3” estimate as a ceiling to budget against, not an exact bill.
Context window budgeting in practice
A common mistake is sizing only the user prompt against the context limit. In reality, the window must hold the system prompt, all prior conversation turns, tool-call results, and the model's own output. Paste your system prompt here first to see its baseline cost, then add representative user messages to estimate the per-turn overhead. Reserve at least 30% of the window for the model's reply — a tight budget forces truncated or degraded output even when the input technically fits.