Prompt Diff
Compare two prompts (or two versions of a prompt) side by side. Line-level diff with highlighted changes. Designed for iterating on LLM prompts.
Comparing prompt versions
Prompt engineering is iteration: you write a prompt, test it, tweak one phrase, test again. After a dozen rounds, you've got a "version 1" and a "version 14" and no clean record of what changed where. This tool gives you that record on demand — paste any two prompts and see exactly which lines were added, removed, or left alone. No git, no setup, no upload.
When a diff saves your deployment
- Auditing a deployed change. Marketing tweaked the system prompt last week — what exactly is different? Paste both versions and read the diff.
- A/B testing prompts. Two candidate prompts, one runs better on evals. Diff them to isolate what the difference might be doing.
- Reverting a regression. The latest prompt is worse than two iterations ago — which line did you change?
- Reviewing a teammate's edit. They sent you a "small tweak" to the system prompt — did they only touch the part they said they did?
- Migrating between model families. Adapting a prompt from GPT to Claude often means small wording changes — diff after rewriting to confirm the structure stayed the same.
Split view vs unified view
- Side-by-side — A on the left, B on the right. Good when both versions are similar length and you want to scan visually.
- Unified — single column with + / − markers, mirroring
git diffoutput. Better for sharing the diff in a Slack message or for sparse changes.
Line-level limits and invisible characters
- This is a line-level diff. A single word changed in the middle of a long line marks the whole line as added+removed. For prose-level diff of a sentence, you may prefer a word-level tool.
- Trailing whitespace. Hidden spaces at the end of a line will mark it as different — useful or noisy depending on the case. Use the "Ignore trailing whitespace" toggle if you only care about visible content.
- Reordered blocks look like delete+add. If you moved a paragraph from position 1 to position 3, the diff shows it as removed at position 1 and added at position 3. There's no "moved" detection.
- Lines, not tokens. This diff doesn't speak token; it speaks lines. If your two prompts have the same content rewrapped to different line lengths, every line will look different. Normalise line breaks first.
- Privacy. Everything stays in your tab. Don't paste secrets into the tool's example placeholder, mind you — the placeholder text is hard-coded, not connected to your input.
Tracing a one-line change
Version A ends a system prompt with "Answer concisely." Version B changes that one line to "Answer concisely. If you are unsure, say so rather than guessing." In unified view the diff shows the original line prefixed − and the new line prefixed +, every other line unchanged — so at a glance you can prove the only behavioural change between the two deployments is the added hedging instruction. That's the audit trail prompt work usually lacks.
Moves, multi-version history, and non-English text
Why does moving a block show as delete + add? The diff is positional, not semantic — it has no "moved" detection. A paragraph shifted from the top to the bottom appears removed where it was and added where it landed. Re-read those as a move, not two independent edits.
Can I diff more than two versions at once? No — it's strictly A vs B. To walk a longer history, diff v1→v2, then v2→v3, and so on, one pair at a time.
Is a line-level diff good enough for prompts? Usually yes, because prompts are edited line by line. The exception is a single word changed inside a very long line: the whole line marks as changed. If your prompt is one giant paragraph, break it onto separate lines before diffing.
Does it work for non-English prompts? Yes — it compares characters, not language, so prompts in any script diff fine.
What a diff catches that your eyes miss
When two prompts “look identical” but behave differently, the difference is almost always something that does not render. A character-level diff surfaces exactly these:
| Invisible difference | Why it changes behaviour |
|---|---|
| Trailing whitespace / double spaces | Tokenises differently; shifts few-shot alignment |
| Smart vs straight quotes (’ vs ') | Different code points; breaks exact-match rules |
| CRLF vs LF line endings | Extra
per line changes the text |
| Zero-width / non-break spaces | Invisible tokens the model still sees |
This is why “I only changed the wording” sometimes flips an output: a copy-paste through a doc or chat app quietly swapped the quotes or injected an NBSP. Diff the two versions character-by-character before concluding the model is non-deterministic.
Practical workflow for prompt iteration
Keep a running text file of prompt versions (v1, v2, v3). After each eval run, paste the current version into side A and the candidate into side B. The diff tells you exactly what changed — so when evals improve or regress, you can attribute the shift to specific lines rather than guessing. This closes the feedback loop that most prompt engineering lacks: a concrete record of what was tried and what it did.