Most flaky LLM apps don't have a model problem — they have a system-prompt problem. These ten anti-patterns are the usual suspects, each with the symptom to watch for and the fix.

When an LLM app behaves inconsistently — ignoring a rule half the time, drifting off format, refusing valid requests — the reflex is to blame the model or reach for a bigger one. More often the fault is in the system prompt. System prompts are code, but they rarely get code review, and they accumulate the same kind of rot: contradictory rules, dead instructions, copy-pasted boilerplate that no longer applies. A prompt that "worked" grows by accretion until it quietly fights itself.

Below are ten patterns that reliably hurt output quality. For each: the symptom you'll actually observe, and the fix. None of these require a better model — they require a cleaner prompt.

1. Contradictory instructions

The prompt says "be concise" in one paragraph and "always explain your reasoning in full detail" in another. Or "never make assumptions" alongside "don't ask clarifying questions." The model can't satisfy both, so it picks one non-deterministically, and the choice flips between runs.

Symptom: the same input produces short answers one time and long ones the next; a rule appears to be followed intermittently.

Fix: resolve the conflict explicitly. Decide which instruction wins and delete the loser, or scope each one ("be concise for factual lookups; explain fully for multi-step reasoning"). Contradictions are the single most common cause of "the model ignores my instructions."

2. The wall of rules

A system prompt that has grown to 60 bullet points, each added to patch one bad output. No hierarchy, no grouping — just a flat list. The model treats them as roughly equal weight, so the three rules that actually matter get the same attention as the 57 edge-case patches.

Symptom: important constraints get dropped under load, especially in long conversations, while trivial ones are honoured.

Fix: group and rank. Lead with the handful of non-negotiable rules, put edge cases in a clearly labelled secondary section, and prune anything you can't remember the reason for. A prompt you can't hold in your head, the model can't either.

3. Negative-only instructions

"Don't be verbose. Don't use jargon. Don't make things up." Negative instructions tell the model what to avoid but leave the target underspecified — there are many ways to "not be verbose," and only some of them are what you wanted.

Symptom: the model technically obeys the prohibition but lands somewhere you didn't intend — terse to the point of unhelpful, say.

Fix: pair every "don't" with a "do." Replace "don't be verbose" with "answer in two to four sentences unless asked for detail." Positive, concrete direction gives the model a target instead of just a fence. Anthropic's prompt-engineering guidance makes this point directly: tell the model what to do rather than what not to do.

4. Vague persona, no task

"You are a helpful, friendly, knowledgeable assistant with a passion for great customer service." Three sentences of personality, zero sentences about what the assistant is actually supposed to accomplish. Persona without task is decoration.

Symptom: pleasant, on-brand answers that don't reliably do the job — the model optimises for tone because that's all you specified.

Fix: spend your prompt budget on the task, the inputs, the expected output, and the success criteria. A role line can help set context, but "You are a support agent for Acme's billing system. Your job is to resolve or correctly escalate billing questions" beats a paragraph of adjectives.

5. Conflicting format demands

"Respond in JSON" and, elsewhere, "always start with a friendly greeting." Or "return a markdown table" plus "keep it to one short sentence." Format instructions that can't coexist force the model to break one of them, and it's rarely the one you'd choose.

Symptom: malformed structured output — JSON with a chatty preamble that breaks your parser, or truncated tables.

Fix: pick one output contract and make it unambiguous. If you need machine-readable output, say "Respond with only valid JSON, no prose before or after," and move any conversational tone into a field inside the structure. Reserve format rules for the format; don't mix presentation and structure.

6. Buried critical instructions

The one rule that must never be broken — "never reveal customer data from other accounts" — sits in the middle of paragraph nine, between two minor style notes. Models attend more reliably to the start and end of a prompt than to the muddy middle, a tendency well documented in the "lost in the middle" research on long contexts.

Symptom: the most important safety or correctness rule is the one that slips, and it slips more as the prompt and conversation grow.

Fix: position by importance. Put critical rules at the top, and restate the truly load-bearing ones near the end of the system prompt. Redundancy for a safety-critical instruction is cheap insurance.

7. Unbounded "always" and "never" that fight

"Always cite a source." "Never speculate." Fine individually — until the user asks something for which the model has no source and no certain answer. The two absolutes collide and the model does something unpredictable: invents a citation, refuses, or hedges its way out.

Symptom: odd failures on exactly the inputs that sit in the gap between two absolute rules.

Fix: add the escape hatch. "Cite a source when one exists; if you don't have one, say so plainly and don't guess." Absolute rules need a defined behaviour for the case where they can't all hold, or the model will improvise one.

8. Constraints with no example

"Match our house style." "Format dates the way we do." "Be professional but warm." These describe a target the model can only guess at, because the actual standard lives in examples you never provided.

Symptom: output that's technically compliant but off — the right idea, wrong flavour, and no consistent way to correct it.

Fix: show, don't just tell. One or two concrete examples of the desired output pin down a style far better than any adjective. Few-shot examples are among the highest-leverage additions to a system prompt; both Anthropic and OpenAI recommend them for exactly this reason.

9. Stale instructions after a model upgrade

Prompts carry scar tissue. A rule like "think step by step before answering" or an elaborate workaround for a formatting quirk was added to compensate for an older model's weakness. You upgrade the model, the weakness is gone, but the workaround stays — now it's either redundant or actively counterproductive.

Symptom: after a model swap, quality is flat or worse despite the newer model being stronger; the prompt is holding it back.

Fix: re-audit the prompt on every model upgrade. Strip compensations for problems the new model doesn't have, and re-test. Treat "which model version was this written for?" as a first-class question. Comparing the before-and-after prompt side by side makes the accumulated cruft obvious.

10. Token bloat

Every token in the system prompt is re-sent on every request — paid for and processed each time — and it competes with the user's actual content for the model's attention. Copy-pasted legalese, redundant restatements, verbose preamble, and dead rules all inflate cost and dilute signal.

Symptom: rising per-request cost, slower responses, and a model that seems distracted from the task by all the surrounding instruction.

Fix: treat prompt length as a budget. Cut anything that doesn't change behaviour, deduplicate, and prefer one clear sentence over three cautious ones. A tighter prompt is cheaper, faster, and usually more reliable.

The pattern behind the patterns

Anti-pattern Core defect
Contradictions / conflicting formatsUnsatisfiable — model picks non-deterministically
Wall of rules / token bloatNo hierarchy — signal diluted
Negative-only / vague persona / no examplesUnderspecified — no clear target
Buried rules / fighting absolutesPositioned or scoped wrong
Stale post-upgrade instructionsRight once, wrong now

Nearly every entry reduces to one of four defects: the prompt is unsatisfiable, unranked, underspecified, or out of date. Once you see prompts through that lens, most fixes are obvious — the hard part is spotting the problem in a prompt you wrote and have stopped reading closely.

Tool walkthrough

Toolhub's system prompt linter is built to catch these mechanically: it flags likely contradictions, negative-only rules with no positive counterpart, absolute "always/never" pairs that can collide, and estimated token weight so bloat is visible before it costs you. Run a prompt through it the way you'd run code through a static analyser — before shipping, not after the incident.

When you're iterating — tightening rules, cutting cruft, or migrating to a new model — prompt diff shows exactly what changed between two versions of a prompt, character by character and rule by rule. That's how you catch a stale instruction you meant to delete but didn't, or verify that a "small" edit didn't silently drop a critical line. Diffing prompts across a model upgrade is the fastest way to strip the accumulated scar tissue from anti-pattern nine.

Where to read further

A system prompt is the most-executed piece of code in an LLM app — it runs on literally every request — yet it's the least likely to be reviewed. Treat it like code: keep it short, keep it consistent, delete what's dead, and re-check it whenever the model underneath it changes. Most "the model isn't good enough" problems turn out to be a prompt that quietly stopped making sense.

← All articles