"hello".length is 5. "😊".length is 2. "é".length is 1 or 2 depending on how it was typed. These all come from the same place: most string functions count storage units, not characters.

Every developer trusts .length until the day it betrays them. Usually it's an emoji in a username, or an accented name that a database swears is different from the identical-looking one already stored. The function isn't buggy — it's answering a different question than the one you asked. It's counting storage units. You wanted characters. Those are not the same thing, and the gap between them is one of the most reliable sources of subtle string bugs in production.

Three things that all get called "length"

There are three distinct concepts hiding under the word "character":

Most string length functions count code units. Humans count grapheme clusters. That mismatch is the whole story.

Why JavaScript gets emoji "wrong"

JavaScript strings are UTF-16 internally. 😊 lives above U+FFFF, so it's stored as a surrogate pair — two code units — and "😊".length returns 2. This isn't a bug; the ECMAScript spec defines length as the number of UTF-16 code units, and it's behaving exactly as written. But it surprises every developer the first time a "one-character" emoji counts as two, blows past a character limit, or gets cut in half by a substring.

The normalization trap

Here's the one that produces "these two strings are identical but not equal." The character é can be encoded two ways: as a single precomposed code point (U+00E9), or as a plain e followed by a combining acute accent (U+0065 U+0301). They render identically. They are not the same sequence of code points, so a naive equality check returns false. This is defined behaviour — Unicode provides normalization forms (NFC, NFD, and compatibility variants NFKC/NFKD) precisely to reconcile them. The Unicode Standard describes the goal directly:

"Essentially, the Unicode Normalization Algorithm puts all combining marks in a specified order, and uses rules for decomposition and composition to transform each string into one of the Unicode Normalization Forms."

— Unicode Standard Annex #15, "Unicode Normalization Forms"

The practical consequence: if one system stores NFC and another compares in NFD, you get invisible mismatches — duplicate accounts, failed lookups, "that username is taken" for a name that looks unused. Normalize to one form at your boundaries and the ghosts disappear.

Combining characters and joiners

It gets deeper. A flag emoji is two regional-indicator code points. A family emoji can be seven or more code points strung together with zero-width joiners. Counting the "characters" in a line of social-media text with .length gives an answer that has nothing to do with what the user sees. Anything that counts, truncates, or validates user-facing text needs grapheme-cluster awareness, not code-unit counting.

Where it actually bites

The failures are concrete. Truncating a string by code-unit count can split a surrogate pair or a grapheme cluster, producing a broken half-character (the infamous replacement glyph). A "max 160 characters" SMS limit is really a byte limit in a specific encoding, not a character count. And a database column declared as a character length can, under some charset configurations, mean bytes rather than characters — so a "255-character" field rejects a shorter string full of multi-byte characters. The rule of thumb: know whether you're counting bytes, code units, code points, or graphemes, because the right answer depends entirely on which one the task needs.

UTF-8, UTF-16, UTF-32: the encoding choice

Unicode defines what the characters are (code points). The encoding defines how those code points become bytes on disk or in memory. UTF-8 uses one to four bytes per code point, is backwards-compatible with ASCII, and dominates the web — over 98% of web pages use it. UTF-16 uses two or four bytes and is the internal format of JavaScript, Java, and Windows APIs. UTF-32 uses a fixed four bytes per code point, which makes indexing trivial but wastes space for ASCII-heavy text. The encoding choice matters because it determines which operations are cheap and which are expensive: random access by code-point index is O(1) in UTF-32 but O(n) in UTF-8, while memory footprint is smallest in UTF-8 for most text. Your language's string type already made this choice for you — but knowing which encoding it picked explains why certain operations have surprising performance characteristics.

The Mojibake pattern

Mojibake — garbled characters like é appearing where é should be — is the visible symptom of a specific, diagnosable encoding mismatch. It happens when bytes written in one encoding are read as another. The most common case: UTF-8 bytes interpreted as ISO-8859-1 (Latin-1). The character é in UTF-8 is two bytes (C3 A9); read as Latin-1, those same bytes are the two characters à and ©. Double mojibake happens when someone "fixes" the display by re-encoding the already-garbled text, producing a deeper layer of corruption. The fix is always at the source: ensure the producing system and the consuming system agree on the encoding, typically by declaring it explicitly in HTTP headers (charset=utf-8), database connection strings, and file metadata. Mojibake is never random corruption — it's a deterministic transformation that tells you exactly which encoding mismatch occurred, if you know how to read the garbled output.

Database and search implications

Unicode handling in databases is a minefield of silent data loss. MySQL's utf8 charset is not actually UTF-8 — it only supports up to three bytes per character, which means any code point above U+FFFF (including most emoji) is silently truncated or rejected. The real UTF-8 charset in MySQL is utf8mb4, and the migration from utf8 to utf8mb4 is one of the most common "we need to fix this before users notice" database tasks. Collation adds another layer: whether café and cafe match in a search depends on whether the collation is accent-sensitive. Whether ß equals ss depends on the locale-specific collation rules. Getting search and sorting right for international text means choosing the correct charset and the correct collation — and testing with actual non-ASCII data, not just English strings.

See what's really in the string

The cure for all of this is to stop guessing and look. Paste any string into a Unicode inspector to see every code point it actually contains — the surrogate pairs, the combining marks, the invisible joiners. Use a word and character counter to see how counts shift with content and normalization. And when two strings "look the same" but compare unequal, run them through a text diff to expose the invisible difference — a combining accent here, a zero-width character there. .length isn't lying so much as answering literally. Ask it the right question and it tells the truth.

← All articles