Reading Source-backed 3-minute read
Paper Shows a Small Average Comprehension Edge
Matched-text studies find a modest paper advantage, but timing, genre and study design complicate what that average can explain.
What −0.21 actually means
A screen does not erase 21% of what you read. That is not what the most-cited number says.
Delgado and colleagues combined 76 comparisons from 54 studies. Participants silently read content-matched text on paper or digitally, then took the study’s comprehension test. The average difference was −0.21 standard deviations in both the between-participant and within-participant analyses, favouring paper.
That is a small standardized mean difference, not a percentage-point penalty. It cannot tell any individual reader how many answers a screen will cost them.
The studies also varied substantially. Delgado’s prediction interval ran from −0.56 to 0.14, crossing zero. A new study conducted under different conditions could therefore find a larger paper advantage, little difference or an advantage for digital reading.
What was held still
These comparisons were narrower than everyday “screen reading.” Delgado required matched content and excluded digital features beyond scrolling. Hyperlinks, multimedia and other richer affordances were outside the analysis.
The test followed reading, but test reliability, socioeconomic status and digital-use experience were often unavailable for coding. The synthesis could estimate an average association across its included comparisons; it could not identify why that difference appeared.
Design labels also matter. Clinton’s synthesis included both between- and within-participant designs. That does not establish that every participant in every included study was randomly assigned to a format.
A primary experiment supports a causal claim when it actually assigns the medium and standardizes the text, task and test. Even then, the claim belongs to that task and population. Combining studies in a meta-analysis does not add experimental control that the primary studies lacked.
Where the gap changes
Clinton reported an overall comprehension effect of −0.25 standard deviations. The estimate was larger for expository text, at −0.32, and close to zero for narrative text, at −0.04.
Delgado found a similar ordering: −0.27 for informational text and 0.01 for narrative text. But its narrative category contained only seven effects.
These two syntheses also share primary studies. Their agreement is not independent replication. Genre was generally a characteristic coded after studies were completed, rather than a reading type randomly assigned within one common experiment.
The defensible conclusion is therefore limited: the average paper advantage appeared more clearly among the included informational-text studies. Those comparisons do not establish genre itself as the cause.
The clock is not one variable
In Delgado’s synthesis, time-limited studies showed a paper advantage of −0.26. The self-paced estimate was −0.09 and its confidence interval crossed zero.
That contrast was coded across existing studies. Their texts, readers and tasks could have differed along with timing, and timing explained only about 5% of the variance. A nonsignificant estimate is also not proof that the formats are equivalent.
Ackerman and Goldsmith manipulated a different timing question. They found no medium difference when readers received adequate fixed study time, but screen performance was worse when readers regulated their own time. That tests clock control, not deadline pressure, and the study did not supply a standardized effect for the contrast.
Clinton’s near-zero reading-time result answers another question again: how long participants read. Fixed time, self-regulated time and pressure from a deadline are not interchangeable controls.
Where the evidence stops
The evidence supports a modest average paper advantage on matched-text comprehension tests, especially in the informational studies represented. It does not support a universal score penalty, a settled cognitive mechanism or the claim that paper always wins.
It also says relatively little about reading with the features that make digital text distinctly digital. The sharper next test is not simply paper versus screen. It is which screen, which navigation, which timing rule and which kind of text—with those conditions manipulated rather than reconstructed afterward.