BPE tokenizers don't care about words. Here's what they actually count — and why code, JSON, non-English text, and repeated field names all cost more than you expect.

You paste a paragraph into GPT-4 or Claude, and you'd swear it's 150 words. The token counter says 210. Add a JSON wrapper and you're at 270. Rename the fields from terse to readable — "ts" to "timestamp", "uid" to "user_id" — and you're at 310. None of that felt like adding text. But you just doubled the bill.

The gap between how long something looks and what it costs in tokens is real, measurable, and predictable. It's not a bug or a pricing trick — it's how Byte Pair Encoding (BPE) works, and every major model provider charges for what the tokenizer sees, not what your users see. Here's what the tokenizer actually counts.

1. BPE isn't a word splitter

The standard intuition — "a token is roughly a word" — holds for plain English prose and nothing else. BPE was originally a data compression algorithm. It builds a vocabulary of subword units by repeatedly merging the most common adjacent byte pairs in the training corpus. Common English words like "the", "is", "you" end up as single tokens. Less common strings get split.

"Byte pair encoding (BPE), also known as digram coding, is a simple form of data compression in which the most common pair of consecutive bytes of data is replaced with a byte that does not occur within that data. A variant of the algorithm is used in large language model tokenization."

— Wikipedia: Byte pair encoding (CC BY-SA 4.0)

In practice, this means the tokenizer's split points are determined by training-corpus frequency, not linguistic boundaries. "The" is one token because it appears constantly. "shouldn't" is two tokens (shouldn + 't) because the apostrophe-prefixed contraction is less common as a unit. "tokenization" might be three. "GPT-4" is four — G, PT, -, 4 — because that exact string wasn't common enough in training data to earn a single-token slot.

Want to know what yours actually costs? The token counter handles the major model tokenizers and gives you the exact count for any string before you commit to it.

2. Code identifiers blow up the ratio

Plain English prose averages roughly 1.3 tokens per word. Code is often 2–3× worse.

Consider a typical Python variable: openai_api_key. Humans read that as three words. The tokenizer sees open, ai, _api, _key — four tokens for three "words". max_retry_attempts lands at three tokens, which looks fine, but only because those fragments happen to be frequent. x_request_id_header? Five tokens.

CamelCase isn't reliably better. UserAuthenticationService splits as User, Authentication, Service — clean, 1:1. But getUserByExternalId becomes six tokens because the internal casing changes where BPE cuts. There's no rule you can memorize; frequency decides.

If you're sending code snippets as context — pasting a function for the model to review, or forwarding a stack trace — count first. A 50-line function that looks like 300 words can easily be 600 tokens when it's code-heavy. That's a budget conversation you want to have before traffic shows up.

3. Non-English has a different exchange rate

BPE vocabularies are heavily biased toward English. Most models were trained on English-dominant corpora, so English subwords are frequent and earn single-token slots. The same text in Japanese, Arabic, or Hindi tokenizes more expensively — often 3–4×.

A 100-word English sentence is roughly 130 tokens. The same meaning in Japanese might be 200–250. In Arabic with diacritics, easily 300+. The model has no idea it's charging more for non-English content — the tokenizer just sees less-common byte sequences and produces more tokens per character.

This matters if you're building a multilingual product. English users are effectively subsidized by the tokenizer's training distribution. Non-English users aren't. If you haven't benchmarked per-request token cost across the languages you actually support, your cost model is wrong — probably by 2× or more on the expensive end.

4. Repeated field names in JSON arrays

Say you have an array of 200 user records, each with the same eight fields. Every one of those field names gets tokenized 200 times.

[
  {"user_id": 1, "first_name": "Ada", "last_name": "L", "created_at": "..."},
  {"user_id": 2, "first_name": "Bob", "last_name": "M", "created_at": "..."},
  ... × 200
]

The string "first_name" is 3 tokens (first, _name, plus quote boundaries). 200 rows × 3 tokens × 8 fields ≈ 4,800 tokens of pure field-name overhead — before counting a single actual value. That's roughly $0.015 per API call at current Sonnet pricing, just for metadata.

Minification helps: {"uid":1,"fn":"Ada","ln":"L","ts":"..."} cuts it significantly. The trade-off is model comprehension — abbreviated keys can confuse the model if it needs to reason about meaning rather than just extract values. If it's filtering or counting, abbreviate. If it's interpreting, the field names earn their token cost.

5. The overhead that runs before line one

Every API request carries tokens you didn't type. Three main culprits:

System prompt. Your system prompt is billed on every single call. A tight 500-token system prompt at 50,000 requests/day is 25 billion input tokens/month before a single user message arrives. Anthropic's prompt caching cuts the cost of stable prefixes by ~90% — mark the preamble with cache_control: {"type": "ephemeral"} and most of those tokens hit cache instead of billing at full rate.

Tool schemas. If you're using function calling, every schema definition is appended before the conversation starts. A tool with eight parameters and descriptive field names adds 200–400 tokens per call. Billed at input rates. On every request. Tool schemas are worth auditing.

Conversation history. Multi-turn conversations re-send all prior turns. Turn 10 includes turns 1–9 in full. This is roughly quadratic growth — a 20-turn conversation carries ~10× the context overhead of a 2-turn one. For most chat applications, conversation length is the largest single cost driver beyond the first handful of turns.

If you're iterating on system prompts, prompt diff shows the delta between versions character by character so you can see exactly what got longer. And if you want to audit whether a system prompt is tighter than it could be, the system prompt linter flags redundant boilerplate — the "be helpful and accurate" filler the model already knows.

Measure before you ship

The fix is boring: run your actual prompt through the actual tokenizer for your actual model before the feature hits production. Not a word count, not an estimate — the real number, with the system prompt, user message template, and a representative conversation baked in.

The token counter handles tiktoken (OpenAI) and Claude's tokenizer, no API key required. Paste your complete prompt — system turn, user turn, any tool schemas — and multiply by your expected daily request volume. That number, written in your design doc before the first line of production code, is what separates a feature that's profitable from one that quietly runs at 3× the cost you planned for.

← All articles