A Flesch score is a real, useful number — and also a shallow one. It counts syllables and sentence lengths, not whether your writing makes sense. Here's what the formulas measure, what they can't, and how not to game yourself.
Paste any block of English text into a readability checker and you get a tidy number back: "Reading Ease 62" or "Grade Level 8.4." It feels authoritative, the way any number does. Word processors flag it, content platforms score it, and plain-language regulations sometimes require it. But the number is measuring far less than most people assume. Understanding exactly what a Flesch score counts — and what it is structurally blind to — is the difference between using it as a helpful check and being quietly misled by it.
The two most common readability formulas were both developed by Rudolf Flesch, one later refined with J. Peter Kincaid for the U.S. Navy. They are computed from the same two inputs, they are decades old, and they are genuinely useful within narrow limits. This article covers how they work, what they can and cannot see, how they get gamed, and where they legitimately belong.
The two formulas, stated plainly
Both Flesch formulas take exactly two measurements from your text: the average number of words per sentence, and the average number of syllables per word. That's it. No dictionary of hard words, no grammar parsing, no meaning.
Flesch Reading Ease produces a score on a roughly 0–100 scale, where higher means easier:
206.835 − 1.015 × (total words / total sentences) − 84.6 × (total syllables / total words)
Flesch-Kincaid Grade Level rescales the same two inputs to approximate a U.S. school grade:
0.39 × (total words / total sentences) + 11.8 × (total syllables / total words) − 15.59
Read the coefficients and the whole behaviour becomes obvious. Longer sentences push both scores toward "harder." More syllables per word push both toward "harder." Shorten your sentences and pick shorter words, and every readability formula on the market rewards you. There is nothing else in the machinery.
What the scores actually measure
They measure two surface features of text that correlate with difficulty across large samples: sentence length and word length. On average, across a lot of writing, longer sentences and longer words really are harder to read. The formulas capture that statistical tendency and nothing more.
This matters because "correlates with difficulty on average" is a much weaker claim than "measures how hard this specific passage is to understand." The score is a proxy. A short-word, short-sentence passage will always score as easy, whether or not it is actually easy to follow.
What the scores are blind to
Everything that makes writing genuinely clear or genuinely confusing lives outside the formula. A readability score cannot see:
- Coherence. Whether one sentence follows logically from the last. You can shuffle the sentences of a paragraph into nonsense and the Flesch score will not move at all — the word and syllable counts are identical.
- Logic and argument. Whether the reasoning holds, whether claims are supported, whether the structure builds. None of that is measurable in syllables.
- Jargon appropriateness. "Ion" is one syllable and "wonderful" is three, so a physics paper full of short technical terms can score as "easy" while being incomprehensible to a general reader. The formula rewards short jargon and punishes plain long words.
- Vocabulary difficulty beyond length. "Bane" and "cane" are equally easy by the formula; one is far less common. Word rarity is invisible.
- Audience fit. The score has no idea who is reading. "Grade 8" is a property of the text's surface statistics, not a measurement of any real reader's comprehension.
The one-sentence version: readability formulas measure the form of text, never its meaning. They are a spell-checker for structural complexity, not an editor.
How the scores get gamed
Because the formula depends on only two levers, anyone can move the number without improving the writing. This happens constantly, sometimes deliberately, sometimes as an accident of "writing for the score."
- Chop every sentence in half. Replace one clear 20-word sentence with three choppy fragments. The average words-per-sentence drops, the grade level falls, and the prose gets worse — staccato, hard to connect, stripped of the connective tissue ("because," "however," "which means") that carries an argument.
- Swap long common words for short rare ones. "Utilise" (four syllables, everyone understands it) scores worse than a short piece of specialist slang the reader has never seen. Optimising for syllable count can push you toward less familiar vocabulary.
- Break lists into fragments. Bullet points with no verbs count as very short "sentences," dragging the average down regardless of whether the content is any clearer.
The lesson is not that low grade levels are bad — plain writing is a real virtue — but that a good score is a consequence of clear writing, not a cause of it. Optimise the writing; let the number follow. Reverse that and you get text engineered to please a formula that cannot read.
Where readability scores legitimately belong
Used as a rough gauge rather than a verdict, these formulas do real work:
- Matching text to a grade level in education. This is the use Flesch-Kincaid was literally built for. A teacher choosing between two passages for a Grade 5 class gets a fast, defensible first-pass filter. It won't tell them which passage is better taught, but it flags the one that is structurally out of range.
- Plain-language compliance. The U.S. Plain Writing Act of 2010 and many agency style guides push toward clear public communication, and readability scores are a common (if imperfect) checkpoint for government forms, insurance documents, and consent notices. A benefits letter scoring at Grade 14 is a legitimate red flag worth investigating.
- Catching your own drift. Writers unconsciously inflate. A score that suddenly jumps ten grades between drafts is a useful nudge to check whether a section got tangled.
The common thread: the score is a trigger for a human to look, never the final judgment. It points; it does not decide.
Where they get misused
The failures are the mirror image of the good uses. Treating the number as a quality target rather than a diagnostic ("all content must hit Grade 6") pushes writers to game the formula. Applying a general-audience target to specialist writing — legal, medical, technical — produces either false alarms or dumbed-down text that loses precision. And using a single readability score to certify that a document is "accessible" ignores everything the formula can't see: layout, terminology, translation, and actual reader testing.
The other indices, briefly
Flesch's formulas are the best known, not the only ones. A few you will encounter, all built on similar surface features:
- SMOG (Simple Measure of Gobbledygook). Estimates the grade level needed to understand a text based on the count of polysyllabic (3+ syllable) words. Popular in healthcare communication because it targets near-full comprehension.
- Gunning Fog Index. Combines average sentence length with the percentage of "complex" (3+ syllable) words to produce a grade level. Widely used in business and journalism.
- Coleman-Liau Index. Unusually, it uses characters per word instead of syllables, which makes it easier to compute programmatically since counting letters is more reliable than counting syllables.
They differ in the details but share the core limitation: all are surface-feature proxies. Agreement between two indices tells you the surface statistics are consistent, not that the writing is clear.
Reading Ease bands, mapped
| Reading Ease | Approx. grade | Audience |
|---|---|---|
| 90–100 | 5th grade | Very easy; broad general public |
| 80–90 | 6th grade | Easy; conversational |
| 70–80 | 7th grade | Fairly easy |
| 60–70 | 8th–9th grade | Plain English; most adults |
| 50–60 | 10th–12th grade | Fairly hard; high school |
| 30–50 | College | Difficult; academic |
| 0–30 | College graduate | Very difficult; professional/legal |
Treat these bands as rough zones, not sharp lines. Most general-audience web writing aims for the 60–70 range, which lands around plain English for the typical adult reader.
Tool walkthrough
Toolhub's readability checker computes Flesch Reading Ease and Flesch-Kincaid Grade Level directly in the browser — nothing is uploaded — and shows the two underlying inputs, words per sentence and syllables per word, so you can see why a passage scored where it did rather than just accepting the number. Because both formulas are ultimately arithmetic on counts, it helps to see those counts explicitly: the word counter breaks down words, sentences, and characters, which makes it easy to sanity-check a score, confirm a passage is long enough for the formula to be stable, or spot the one 60-word run-on sentence dragging your grade level up.
The workflow that actually improves writing: draft for meaning first, then run the readability checker as a diagnostic. If the grade level is high, find the specific long sentences and dense passages responsible, and rewrite those for clarity — don't chop everything uniformly to satisfy the average.
Where to read further
- Wikipedia: Flesch-Kincaid readability tests — the exact formulas, coefficients, and the history of the Navy commission that produced the grade-level version.
- PlainLanguage.gov — the U.S. federal plain-language guidance and the Plain Writing Act, including why clear writing is about far more than a readability score.
- Wikipedia: Readability — a broader overview covering SMOG, Gunning Fog, Coleman-Liau, and the known limits of formula-based measurement.
A Flesch score is a thermometer, not a diagnosis. It reliably tells you one thing — how long your sentences and words are, expressed as a grade or an ease score — and it tells you that instantly and for free. Everything else about whether your writing is any good stays where it always was: with a reader who understands what you meant to say. Use the number to catch the passages worth a second look, and never let it do the editing for you.
← All articles