Token Counter

Estimate token counts across GPT-4, Claude, Llama, Gemini, and other LLMs. Paste text, see counts per model side-by-side. Files never leave your browser.

ModelTokensInput $/1MCost (this text)
Estimates only. Real tokenizers (tiktoken, SentencePiece, etc.) may differ by ±10%. For exact counts, run the model's tokenizer locally.

Why token counts vary by model

Every interaction with an LLM is metered in tokens — sub-word units that the model's tokenizer carves from your text. Tokens drive both context-window limits ("does this fit?") and pricing ("how much will this cost?"). The exact count depends on the model's tokenizer, which you usually don't have at hand. This tool gives you a fast estimate for every major model side-by-side, plus the dollar cost for the input text against each model's published per-token price.

Budget, cost, and context-window planning

Heuristic accuracy vs real tokenizers

Pricing and the output multiplier

Tokens, words, whitespace, and CJK

Sizing a batch job

Paste a 500-word prompt and the tool estimates roughly 650–700 tokens — English runs about 1.3 tokens per word — then multiplies by each model's published input price so you can see the cost per model side by side. Multiply that by the 100 inputs in your batch job and you know the bill before you spend a cent.

Accuracy, output billing, and whitespace traps

How accurate is the count? It's a heuristic, not the model's real tokenizer, so treat it as sizing rather than billing truth — usually within 10–20% of what tiktoken or SentencePiece would report.

Why do models disagree on the count? Each model carves text with its own tokenizer, so the same string splits into a different number of sub-word units — and therefore a different price.

Does this include output tokens? No — it measures your input text only. Output is billed separately and usually at a higher rate, so leave headroom in your estimate.

Do spaces and line breaks count? Yes — whitespace becomes tokens too. That's why a stray BOM marker or copy-paste noise can quietly inflate a count and trip a context-window limit.

Chars, tokens, words — the rough conversions that matter for cost

LLMs bill by token, not word, and a token is roughly a common word-chunk of about four characters of English. That gives usable back-of-envelope conversions, but the ratio shifts hard by content type:

ContentRough chars/token100 tokens ≈
Plain English prose~4~75 words
Code / JSON~3 (denser)Far fewer “words”; braces & punctuation each cost a token
Rare words, other languages, emoji<2One emoji or CJK glyph can be several tokens

Two traps: (1) whitespace and punctuation are tokens too, so minifying JSON genuinely cuts cost; (2) the same prompt tokenises differently across model families — treat any “words × 1.3” estimate as a ceiling to budget against, not an exact bill.

Context window budgeting in practice

A common mistake is sizing only the user prompt against the context limit. In reality, the window must hold the system prompt, all prior conversation turns, tool-call results, and the model's own output. Paste your system prompt here first to see its baseline cost, then add representative user messages to estimate the per-turn overhead. Reserve at least 30% of the window for the model's reply — a tight budget forces truncated or degraded output even when the input technically fits.