The first 32 codes in ASCII aren't letters — they're commands. Most are relics, but a handful still quietly break CSV imports, corrupt log parsing, and cause "identical" files to differ.

Everyone knows ASCII as the code where 65 is A and 97 is a. But the first 32 codes — 0 through 31 — aren't printable characters at all. They're control codes, commands from an era of teletype machines, and while most are museum pieces, a handful are still very much alive and still quietly breaking things: the CSV that imports as one giant row, the log line that won't parse, the two files that are "identical" but don't match. The trouble with these characters is that you can't see them.

What a control character is

The distinction is baked into the character set itself. Wikipedia:

"In computing and telecommunications, a control character or non-printing character (NPC) is a code point in a character set that does not represent a written character or symbol. They are used as in-band signaling to cause effects other than the addition of a symbol to the text."

— Wikipedia, "Control character" (CC BY-SA 4.0)

"In-band signalling" is the key phrase: these codes live inside the text stream but tell the receiver to do something — start a new line, sound a bell, mark the end of a record — rather than display a glyph.

The one that still causes daily pain: CR vs LF

Two control characters handle "new line," and their disagreement is the most common cross-platform bug in text. Line feed (LF, code 10) is used by Unix, Linux, and macOS. Carriage return followed by line feed (CRLF, codes 13 then 10) is used by Windows. So a file created on Windows and processed on Linux can carry an extra invisible carriage return at the end of every line — which shows up as a trailing \r that breaks string comparisons, corrupts the last field of every CSV row, or makes a config value silently wrong. The line looks right; there's an invisible character on the end of it. This single mismatch has consumed an astonishing amount of developer time over the decades.

The null byte and the tab

Two more worth knowing. The null character (code 0) marks the end of a string in C and many systems built on it — so a stray null in your data can truncate a value at exactly that point, silently discarding everything after it, sometimes with security implications. And the tab (code 9) is the classic CSV saboteur: paste data containing tabs into a tab-separated file and your columns shift; mix tabs and spaces in indentation-sensitive formats and you get errors with no visible cause (this is YAML's cardinal sin). Both are invisible or near-invisible, which is exactly why they're so effective at breaking things quietly.

Why "identical" files differ

Put it together and you get the maddening experience of two pieces of text that look character-for-character identical but compare as different, or a diff tool that flags a line where you can see no change. The change is real — it's just a control character your eyes and your editor render as nothing. A trailing carriage return, a tab where you expected spaces, a non-breaking space that isn't a normal space, an invisible byte-order mark at the start of a file. When text "should" match and doesn't, an invisible character is the prime suspect.

The byte-order mark: the one at the start

Strictly speaking, the byte-order mark (BOM, U+FEFF) is a Unicode character, not an ASCII control character — but it causes the same class of invisible-character bugs and deserves mention here. A BOM at the start of a UTF-8 file is three bytes (EF BB BF) that are invisible in most editors but very much present to parsers. A JSON file with a leading BOM will fail to parse in strict implementations. A CSV opened with a BOM will have the BOM bytes prepended to the first field name, so a column header that looks like id is actually \xEF\xBB\xBFid — and lookups against "id" silently fail. Shell scripts with a BOM can produce "command not found" errors on the shebang line. The BOM is technically optional in UTF-8, and its presence causes more problems than it solves. If you encounter a file that parses wrong from the very first byte, suspect a BOM.

Escape sequences and how they travel

Control characters are represented differently depending on context, and this trips people up when debugging. In a terminal, a carriage return is the invisible character code 13. In a JSON string, it's the escape sequence \r. In a hex dump, it's 0D. In a regex, it's \x0D or \r. These are all the same byte, but the representation changes per tool. When someone reports "there's a \r in the data," they might mean the literal two-character string backslash-r, or they might mean the single-byte carriage return — and the fix is different for each. Clarify which representation you're looking at before you start fixing, because replacing the wrong one either does nothing or makes it worse.

The bell, the backspace, and the form feed

Most of the 32 control characters are genuine relics — the bell (code 7, which once made a teletype ring), the backspace (code 8), the form feed (code 12, which advanced paper to the next page). They shouldn't appear in modern data, but they do: copy-pasted from legacy systems, embedded in PDFs, or generated by hardware that still speaks old protocols. A form feed in a log file can cause a log viewer to split the display unexpectedly. A backspace character in a filename (yes, it's technically possible on some filesystems) creates a file that's nearly impossible to delete by typing its name. These are rare bugs, but when they hit, the symptom — "the data looks fine but behaves wrong" — is identical to the CR/LF and null-byte cases. The diagnostic reflex is the same: stop trusting what the text looks like and inspect the actual bytes.

Make the invisible visible

The fix for every bug above is the same: stop trusting how text looks and inspect what it contains. An ASCII table lets you look up exactly which control code you're dealing with — is that a code 10, a code 13, or both? — so "trailing whitespace" becomes a specific, fixable character. A Unicode inspector reveals every code point in a string, including the control characters and invisible spaces that render as nothing, so you can finally see the \r on the end of the line. And when two "identical" strings won't match, a text diff pins down exactly where the invisible difference lives. The characters you can't see are still there — the tools just make them show up.

← All articles