What the formulas measure
Classic readability formulas count surface features. Flesch-Kincaid uses average sentence length and average syllables per word; the Gunning fog index uses sentence length and the proportion of words with three or more syllables; SMOG counts polysyllabic words across a sample. All produce an approximate US school grade level.
They were developed for specific practical purposes — the US Navy commissioned Flesch-Kincaid to assess technical manuals — and they work reasonably well for their intended use: comparing documents of the same kind and flagging prose that is unnecessarily dense.
The critical limitation is that none of them reads the text. They cannot tell whether a sentence makes sense, whether the argument follows, or whether the vocabulary suits the audience. Text scrambled into random word order scores identically to the original, which is the clearest demonstration of what these measures do not capture.
Where the scores mislead
The syllable-counting proxy for difficulty produces obvious errors. "Aforementioned" is five syllables and widely understood; "bight", "scree" and "quoin" are one syllable each and known to almost nobody. Technical writing full of familiar domain terms scores as difficult while being perfectly clear to its audience.
Optimising for the score directly makes writing worse. Chopping sentences into uniform short fragments to lower the grade level produces choppy, monotonous prose — varied sentence length is what creates readable rhythm, and a document of nothing but eight-word sentences is exhausting. Replacing a precise long word with a vague short one reduces the score and the clarity together.
The formulas also assume connected English prose. Applied to bullet lists, tables, code, headings or text with many abbreviations, they return numbers that mean nothing, since sentence boundaries and syllable counts are not meaningful there.
Using the number as a signal
Treat a score as a prompt to look rather than as a target to hit. A sudden spike in a document's grade level usually points at a genuinely tangled passage worth rewriting; a consistently high score across a piece intended for a general audience is worth investigating.
Common guidance is grade 8 to 10 for general audiences, and plain-language requirements in several jurisdictions set explicit targets for public information — many health literacy guidelines suggest grade 6 to 8 for patient material, and some government accessibility standards reference readability directly.
The changes that genuinely improve clarity are structural rather than statistical: prefer active voice where the actor matters, put the main clause first, replace nominalisations with verbs, define terms on first use, and break long paragraphs at natural boundaries. Reading the text aloud — or having it read to you with the speech tool — surfaces awkwardness no formula detects. The best test remains a reader from the intended audience.