Prompt Diff

Compare two prompts (or two versions of a prompt) side by side. Line-level diff with highlighted changes. Designed for iterating on LLM prompts.

Comparing prompt versions

Prompt engineering is iteration: you write a prompt, test it, tweak one phrase, test again. After a dozen rounds, you've got a "version 1" and a "version 14" and no clean record of what changed where. This tool gives you that record on demand — paste any two prompts and see exactly which lines were added, removed, or left alone. No git, no setup, no upload.

When a diff saves your deployment

Split view vs unified view

Line-level limits and invisible characters

Tracing a one-line change

Version A ends a system prompt with "Answer concisely." Version B changes that one line to "Answer concisely. If you are unsure, say so rather than guessing." In unified view the diff shows the original line prefixed − and the new line prefixed +, every other line unchanged — so at a glance you can prove the only behavioural change between the two deployments is the added hedging instruction. That's the audit trail prompt work usually lacks.

Moves, multi-version history, and non-English text

Why does moving a block show as delete + add? The diff is positional, not semantic — it has no "moved" detection. A paragraph shifted from the top to the bottom appears removed where it was and added where it landed. Re-read those as a move, not two independent edits.

Can I diff more than two versions at once? No — it's strictly A vs B. To walk a longer history, diff v1→v2, then v2→v3, and so on, one pair at a time.

Is a line-level diff good enough for prompts? Usually yes, because prompts are edited line by line. The exception is a single word changed inside a very long line: the whole line marks as changed. If your prompt is one giant paragraph, break it onto separate lines before diffing.

Does it work for non-English prompts? Yes — it compares characters, not language, so prompts in any script diff fine.

What a diff catches that your eyes miss

When two prompts “look identical” but behave differently, the difference is almost always something that does not render. A character-level diff surfaces exactly these:

Invisible differenceWhy it changes behaviour
Trailing whitespace / double spacesTokenises differently; shifts few-shot alignment
Smart vs straight quotes (’ vs ')Different code points; breaks exact-match rules
CRLF vs LF line endingsExtra per line changes the text
Zero-width / non-break spacesInvisible tokens the model still sees

This is why “I only changed the wording” sometimes flips an output: a copy-paste through a doc or chat app quietly swapped the quotes or injected an NBSP. Diff the two versions character-by-character before concluding the model is non-deterministic.

Practical workflow for prompt iteration

Keep a running text file of prompt versions (v1, v2, v3). After each eval run, paste the current version into side A and the candidate into side B. The diff tells you exactly what changed — so when evals improve or regress, you can attribute the shift to specific lines rather than guessing. This closes the feedback loop that most prompt engineering lacks: a concrete record of what was tried and what it did.