Unrealistic Scientific Optimism
Summary: Blog post (Games with Words, 2015) arguing that low statistical power, combined with publication bias toward significant results, systematically inflates reported effect sizes — making the psychology literature chronically overoptimistic and studies chronically underpowered.
Sources: Raw/Unrealistic Scientific Optimism.md
Last updated: 2026-05-07
The core problem
Statistical power is the probability that a study will detect an effect if it exists. Most psychology studies are underpowered — they lack the sample sizes needed to reliably detect small effects. The problem compounds because:
-
Effect sizes are inflated in the literature. Journals publish significant results; null results disappear. Since effect size and significance are correlated, the published literature systematically reports inflated estimates.
-
Researchers calibrate sample sizes to the inflated estimates. If psychologists believe typical effects are ~0.5 SD, they design studies sized for 0.5 SD — but the true effect may be 0.1 SD, requiring roughly 25× more subjects.
-
Contingent stopping makes things worse. Checking for significance periodically and stopping when results go the wrong way introduces further distortion even with 1,000 subjects per condition.
Effect size inflation: the simulation
If you only publish significant results, the average reported effect size is always higher than the true effect — sometimes dramatically. With 50 subjects per condition, you will almost never report an effect smaller than 0.5 SD even if the true effect is 0.1 SD, because only flukes large enough to reach significance get published.
The implication: a significant result from a small sample is more likely to be an error than a true finding, unless the effect is very large (e.g., the Stroop effect).
The epistemological trap
Without widespread publication of null results and replications, we cannot determine how badly our perception is distorted — because the degree of distortion depends on how large effects really are, which we cannot know from a literature that only shows significant results. This is a self-concealing problem. See probability on the gap between decision quality and outcome quality; language-log-gigo on how bad inputs corrupt scientific publishing at scale.