Summarization Ratios: What a Compression Number Actually Tells You
A summary can be objectively measured in one narrow but genuinely useful way: how much shorter it is than the source it came from. That number does not tell you whether the summary is good — a compression ratio has no opinion on accuracy or nuance — but it does tell you something concrete about how much material had to be discarded to reach that length, and thinking in ratios rather than vague impressions makes the tradeoff much easier to reason about honestly.
The arithmetic behind a compression ratio
A compression ratio compares a summary's word count to its source's word count as a percentage: the ratio itself (summary divided by original), the reduction (100 minus that ratio, i.e. how much was cut), and a compression factor expressing the same relationship as an "×" multiple — a summary a fifth the length of its source is a 5× compression. These three numbers are three views of the same underlying fact, useful in different contexts: the ratio for describing density, the reduction for emphasizing what was cut, and the factor for a quick intuitive sense of "how much shorter." A tool built around this arithmetic will also report the raw word count removed, which is worth glancing at on its own — a reduction that sounds modest as a percentage can still mean thousands of actual words were cut from a long source.
Three real compression levels, worked
Start with a 3,000-word literature-review section and three different summary lengths. Compress it to 250 words — a short abstract-style summary — and the ratio comes out to 8.3%, a reduction of 91.7%, and a compression factor of 12×. Compress the same source to 750 words instead, closer to an extended abstract or a detailed synopsis, and the ratio rises to 25%, the reduction drops to 75%, and the factor falls to 4×. Compress it more gently still, to 1,500 words, and you get a 50% ratio, a 50% reduction, and a 2× factor — genuinely more of a condensed rewrite than a summary at that point.
Those three numbers are not just decoration; they describe three different kinds of object, not three lengths of the same one. A 12× compression has to make brutal editorial choices about what survives — typically the central claim and maybe one supporting point, full stop. A 4× compression can retain the claim, the main evidence, and at least a gesture at the caveats. A 2× compression is close enough to the original that most of the argument's structure, not just its headline, can survive intact. Knowing which ratio you are actually asking for — and which one a tool actually gave you — tells you how skeptical to be about what got left out.
What specifically disappears at high compression
The things that vanish first as compression ratios climb are not random; they are the qualifying, hedging, and boundary-setting language that separates a responsible claim from an overstated one. A source that says a result "may suggest a pattern under certain conditions" is exactly the kind of sentence a tight summary flattens into "the study found," because the hedge takes words and the flat version does not. Limitations sections are especially vulnerable, since they read as an admission of weakness rather than a headline finding, and an aggressive summarizer optimizing for a tight word count will treat them as the first thing to cut. The caveats are not decoration on a finding — they are part of what makes the finding true within its actual scope, and losing them while keeping the flat claim is how a summary quietly overstates something the original source was careful about.
Choosing the right ratio for the job
The fix is not to avoid heavy compression altogether — a 12× abstract-level summary is genuinely useful for triaging whether a source is even worth reading in full, and asking for anything longer defeats that purpose. The fix is to match the compression level to what you actually intend to do with the result. Triaging a stack of twenty potential sources to find the five worth reading calls for aggressive compression; you are looking for topic and relevance, and losing nuance at that stage costs you little because you have not committed to using the source yet. Preparing to cite a specific claim calls for the opposite: little or no compression on the passage that actually matters, because that is exactly the sentence where the caveats you might lose are the ones that determine whether your citation of it is accurate.
A high ratio can also mean truncation, not summarization
There is a second, easily confused failure mode that produces a high compression number for the wrong reason: the source was too long for the tool to process in a single pass, and what actually got compressed was only the portion it saw — often just the opening — not the whole document. A summary in that case can look like an aggressive but competent 12× compression of the full source, when it is really something closer to a complete or near-complete rendering of the first third, silently missing the middle and the conclusion entirely. This distinction matters because the two failures call for different fixes: a genuinely over-compressed summary of the whole source needs a longer summary; a truncated one needs the document split into sections and summarized piece by piece, then synthesized, regardless of what length you asked for. Before trusting a high-ratio summary of a long document, confirm the methods and conclusions sections were both actually represented in what came back, since those are the parts most likely to have been silently dropped rather than genuinely compressed.
Using the numbers as a gut-check, not a target
The most useful habit is to compute the ratio after the fact and ask whether it matches what you actually intended, rather than requesting a specific ratio up front and trusting it blindly. If you asked for "a short summary" and got something compressing at 60% — barely condensed at all — that mismatch is worth noticing, because either your source was already tight or the tool under-compressed. If you asked for "the key points" and got a 15× compression down to a single flat sentence, that is worth noticing too, because a claim with no supporting evidence attached is not something you can respond to intelligently, let alone cite. The number itself is neutral; what it is good for is catching the gap between what you asked for and what you actually received, which is often larger than it feels while you are skimming the smooth-sounding result.
A practical rule for citation-bound summaries
Whenever a summarized claim is going to end up supporting something in your own document — cited, quoted, or built into your argument — treat the summary as a pointer to a passage, not as the passage itself, and go read that specific paragraph in the source before you rely on it. This single habit catches the majority of summarization-driven distortion, because it converts "trust the compressed version" into "verify the one sentence that actually matters," which costs a few minutes rather than an hour and removes almost all of the risk. Reserve your full trust in a compressed summary for the parts of your process where being wrong costs you little — initial triage, a first pass at whether a topic is relevant at all — and re-open the source itself for anything that is going to carry weight in your own writing. The few minutes that costs is cheap compared to the cost of an inaccurate citation surviving all the way to a submitted document, and it is far cheaper than the cost of retracting or correcting one after the fact.
Requesting the caveats explicitly changes the output
If you are using an AI tool to produce the summary, the compression ratio you get is partly a function of what you asked for, and asking explicitly for the parts that get cut first changes what survives. Requesting the central claim, the evidence behind it, the conditions under which it holds, and the limitations the source itself acknowledges — by name, as four separate things to preserve — forces a tighter summary to keep the qualifying language it would otherwise drop by default. This does not eliminate the tradeoff between length and completeness; a 250-word summary asked to cover four dimensions of a finding will still be tighter on each one than a 750-word summary would be. But it does mean the compression happens where you chose it to happen, rather than wherever the flattening tendency of a generic "summarize this" request happened to cut.
None of this arithmetic tells you whether a given summary is trustworthy — only how much of the source's original length it kept. Trustworthiness still has to come from checking the summary against the source directly wherever the stakes are high enough to matter, the same discipline that applies to any other AI-generated claim you plan to build on. The ratio just tells you, honestly and in advance, how much room the summary had to get it right.