What Readability Scores Actually Measure — and What They Miss
Readability formulas get treated two contradictory ways: as gospel by people chasing a target score, and as nonsense by people who have watched a formula reward a vague sentence over a precise one. Both reactions come from not knowing exactly what the arithmetic does. The formulas are not vague at all — they are precisely defined and easy to compute by hand. Once you see the actual numbers behind a few real passages, the formula's honest behavior, and its honest limits, both become obvious.
The exact arithmetic
The Flesch Reading Ease score is computed from just two inputs: the average number of words per sentence, and the average number of syllables per word. The formula is 206.835 − 1.015 × (words per sentence) − 84.6 × (syllables per word), so every point the score drops comes from a specific, traceable increase in one of those two averages — there is nothing hidden in the calculation. A companion formula, the Flesch–Kincaid Grade Level, uses the same two inputs differently: 0.39 × (words per sentence) + 11.8 × (syllables per word) − 15.59, producing a number that approximates a U.S. school grade level. Notice what is not in either formula: nothing about grammar correctness, nothing about logical structure, nothing about whether a claim is well-supported, nothing about whether a technical term is used precisely or loosely. The formulas measure exactly two surface properties of a passage and nothing else. That is not a flaw to apologize for — it is the whole design, and it is why the score is reproducible and objective in a way that "this reads well" never is.
Three real passages, actually scored
Rather than describe this abstractly, here are three short passages run through the formula directly, so the numbers are real rather than illustrative guesses.
Passage one, plain and short-sentenced: “The cat sat on the mat. It was warm there. The sun came in through the window. Soon the cat fell asleep.” — and so on for ten short sentences. Run through the scorer: 65 words, 10 sentences, 72 syllables, for an average of 6.5 words per sentence and 1.1 syllables per word. That produces a Flesch Reading Ease of 106.5 and a grade level of 0 — both scores pinned at their easiest extreme, labeled “Very easy (5th grade).”
Passage two, ordinary clear nonfiction prose about writing feedback, running to four longer sentences: 95 words, 4 sentences, 139 syllables — an average of 23.8 words per sentence and 1.5 syllables per word. That comes out to a Flesch Reading Ease of 58.9 and a grade level of 10.9, labeled “Fairly difficult (10th–12th grade).” Nothing about this passage is badly written; it simply uses longer sentences than the first example, and the formula reflects exactly that and nothing more.
Passage three, a dense two-sentence passage on renormalization-group physics, deliberately written the way real technical prose in that field actually reads: 101 words across just 2 sentences, 189 syllables — 50.5 words per sentence and 1.9 syllables per word. The Reading Ease score comes out at −2.7, below the formula's nominal floor, with a grade level of 26.2, labeled “Very difficult (college graduate).” That passage is not poorly written. It is precise, technical writing for a technical audience, using long sentences and polysyllabic vocabulary because the subject genuinely requires them. The score is telling you the truth about its surface structure, not passing judgment on its quality.
What the score is measuring, restated plainly
Look at what actually moved between the three examples: sentence length climbed from 6.5 to 23.8 to 50.5 words, and syllables per word climbed from 1.1 to 1.5 to 1.9. Those two numbers alone explain the entire spread from a 106.5 down to a −2.7. That is the whole formula, working exactly as designed. It is a genuinely useful signal — long, syllable-heavy sentences are a real and common cause of writing that is harder to follow than it needs to be, and the score gives you an objective way to notice that pattern in your own drafts without having to trust your own ear, which gets used to your own habits.
What the score cannot see
The formula has no way to detect whether a short sentence is actually clear or just vague, because vagueness and precision can both be expressed in short words. It cannot tell that a technical term is the single most precise word available, rather than needless jargon, and it will reward you identically for replacing that term with a fuzzier one that happens to have fewer syllables. It cannot detect a logical gap, an unsupported claim, or an argument that does not follow from its premises — a passage can score as “very easy” while being confidently wrong from the first sentence to the last. And it says nothing at all about whether a sentence is grammatically correct; a grammatically broken fragment and a clean, complete sentence of the same length and syllable count score identically. A dense but genuinely well-written physics paper scoring “very difficult” is the formula working correctly, not a false alarm — and a confidently wrong five-word sentence scoring “very easy” is the same formula's blind spot, working exactly as designed on the one axis it was built to measure.
Two numbers, two different jobs
It is worth being clear about why the calculator reports two numbers from the same two inputs rather than one. The Reading Ease score is a 0-to-100-ish scale where higher means easier, useful for a quick relative comparison — is this draft easier or harder to read than my last one? The Grade Level estimate translates the same underlying measurement into an approximate U.S. school grade, which is more useful when you have an external target to hit, such as writing guidance aimed at a specific reading level. They will always move together, because they are built from the same two underlying counts, so you rarely need both for the same decision. Pick whichever framing matches the question you are actually asking: "is this clearer than before" wants the Ease score; "will a general reader at roughly this education level follow it" wants the Grade Level. Neither one needs a second, independent measurement to be meaningful — they are two readings of the same instrument, not two separate opinions about your writing.
Does AI-generated text score differently?
People sometimes ask whether AI-written prose has a distinctive readability signature. It can, but not for a deep reason — language models tend by default toward a smooth, moderate register: neither the short, plain sentences of passage one above nor the genuinely dense technical style of passage three, but something closer to passage two, competent and moderately complex. That is a statistical tendency in default output, not a rule, and it changes considerably depending on what you asked for and how much you edited afterward. It is also not a reliable way to detect whether text was AI-written or not — plenty of careful human writers land in exactly that same moderate range, and a formula that only sees sentence and syllable length cannot distinguish authorship. Do not use a readability score as an AI-detection tool; it was never built for that job, and treating it as one will produce false confidence in both directions.
Using the score without misusing it
The productive way to use a readability score is as a smoke detector, not a style guide. A surprisingly high grade-level number on a passage you intended for a general audience is a genuine signal to go look at your sentence lengths and see whether some of them are carrying two or three ideas that would read more clearly split apart. A surprisingly low number on a passage meant for a specialist audience might mean you have oversimplified language that the reader would actually prefer precise. In neither direction does the number tell you what to write instead — it tells you where to look. Chasing a specific target score by mechanically shortening sentences or swapping in simpler synonyms, without rereading for whether the meaning survived the edit, is how you end up with prose that scores well and says less than it did before.
Register is a separate decision from readability
None of this means every audience wants the same score. A dissertation committee, a general-interest reader, and a grant reviewer are different audiences who tolerate different sentence lengths and vocabulary, and matching your register to your actual reader is a judgment call the formula cannot make for you. Use the score to catch accidental complexity — the sentence that got away from you while you were thinking, not the one you chose deliberately because your reader needs the precision. The number is a measurement of surface structure. What you do with that measurement is still, entirely, a writing decision that belongs to you.
If you want to see how readability, page length, and reading time move together across a whole document rather than a single passage, the document planning reference lays several of these figures out side by side for a range of typical document lengths, all computed the same way as here — directly from the formula, not from a guess.