Few phrases sound more final than “statistically significant.” In everyday language, significant means important. In statistical practice, it has a narrower meaning. Confusing those meanings turns a calculation into a verdict it was never designed to deliver.
Start with the question a p-value asks
A p-value is calculated under a statistical model that includes a null hypothesis—often no difference or no association. Loosely stated, it describes how incompatible the observed data, or more extreme data, are with that model. A small value says the data would be relatively unusual under the specified assumptions.
It does not tell us the probability that the null hypothesis is true. It does not tell us the probability that the result happened “by chance.” It does not measure the size or practical importance of an effect. Those common translations reverse or enlarge what the calculation actually says.
Where the 0.05 threshold came in
Many fields developed a convention: call p < 0.05 statistically significant and p ≥ 0.05 not significant. The boundary can help standardize a decision when a rule genuinely must be set. Trouble begins when a continuous measure is treated as a cliff. Results with p = 0.049 and p = 0.051 are not meaningfully opposites.
The American Statistical Association has emphasized that scientific conclusions and decisions should not rest only on whether a p-value crosses a threshold. Context, design quality, measurement, evidence from other studies and the costs of errors all matter.
Significance is not effect size
| Quantity | Useful question | What it does not settle |
|---|---|---|
| Effect estimate | How large was the observed difference or association? | How precise or unbiased it is |
| Confidence/uncertainty interval | Which effect sizes remain reasonably compatible with the analysis? | Whether every assumption is correct |
| P-value | How incompatible are the data with a specified model? | Truth, importance or probability of replication |
| Sample size | How much information was collected? | Representativeness or design quality |
| Practical threshold | How large must an effect be to matter for the decision? | Whether the study estimated it accurately |
With a very large sample, a tiny and practically irrelevant difference can produce a small p-value. With a small sample, a potentially important effect may not cross 0.05 because the estimate is imprecise. “Not significant” therefore does not mean “no effect.” It may mean the study could not distinguish among several possibilities.
Read the interval, not only the label
Suppose a study estimates an improvement of 3 units with an interval from 1 to 5. Another estimates 3 units with an interval from −4 to 10. The central estimates match, but the second leaves much more uncertainty. Reporting only whether each crosses a threshold hides that difference.
Intervals also help readers compare the evidence with a meaningful benchmark. If an effect smaller than 2 units would not matter in practice, an interval concentrated between 0 and 1 tells a different story from one spanning 0 to 12—even if both are described as “not significant.”
Why design still comes first
A precise analysis of biased data can produce a precise wrong answer. Statistical calculations do not repair unmeasured confounding, selective sampling, unreliable instruments, missing outcomes or a comparison that fails to answer the research question. Before reading the p-value, ask how observations were generated.
Randomization, blinding, prespecified outcomes, appropriate controls and transparent handling of missing data protect different parts of the study. Statistical significance is meaningful only inside that design.
The multiple-testing problem
If researchers test many outcomes, time points, subgroups and models, some p-values can fall below 0.05 through random variation even when no corresponding effect exists. Imagine testing 100 independent null relationships at a 0.05 threshold. Under simplified assumptions, about five could cross the line on average.
Researchers can address this by specifying primary questions in advance, adjusting for multiple comparisons, separating exploratory from confirmatory work and validating patterns in new data. Readers should ask how many analytical routes were available, not only which result was reported.
Selective reporting and the literature
Statistically significant findings have historically been easier to publish than null or ambiguous results. This can make the published literature look more decisive than the full research record. Trial registration, study registration, results reporting and access to protocols help reveal what was planned and what changed.
A systematic review may detect signs of missing small studies, but no statistical method can perfectly reconstruct unseen evidence. Publication practices are part of statistical interpretation.
A better way to report a result
- Name the design and comparison. Explain how the data were produced.
- Give the effect in understandable units. Include absolute quantities where possible.
- Show uncertainty. Report an interval and explain what range matters.
- State the analysis status. Was it planned, exploratory or one of many tests?
- Test robustness. Do reasonable alternative assumptions change the conclusion?
- Use the wider evidence. Compare with replications, systematic reviews and plausible mechanisms.
How to read “not significant”
Ask whether the estimate is close to zero and precise, or merely uncertain. A well-powered study with a narrow interval around a negligible effect can provide evidence against a meaningful benefit. A small study with a wide interval may be compatible with benefit, harm or little difference. Both can produce p ≥ 0.05; they do not communicate the same knowledge.
Connect the number to the claim
Our guide to Correlation vs. Causation explains why an association does not become causal by crossing 0.05. How to Read a Scientific Study shows where to find outcomes, figures and limitations. For the larger framework, see How Scientific Discovery Works and How to Evaluate New Scientific Discoveries.
Statistical field notes
Is p < 0.05 good evidence?
It can be evidence against a specified model, but its force depends on design, assumptions, prior plausibility and how many analyses were attempted. A p-value cannot tell whether measurement was biased or whether the selected model answers the real question.
Does a smaller p-value mean a larger effect?
No. Sample size and variability influence the value. A huge dataset can give a very small p-value for a tiny difference. Compare the effect estimate and its units directly.
Is p = 0.05 a five-percent probability of error?
No. It is not the probability that the conclusion is wrong. Error rates for a repeated decision procedure require additional assumptions; the probability of a particular claim requires a different framework and prior information.
Are confidence intervals always better?
They reveal magnitude and precision that a binary label hides, but they also depend on models and assumptions. Read them alongside design and substantive knowledge rather than treating them as a replacement magic object.
Statistical power without mythology
Power is the probability that a procedure will detect a specified effect under assumed conditions. Planning for adequate power can reduce inconclusive studies. Afterward, the observed interval is usually more informative than a power calculation based on the observed effect.
Practical importance must be defined
Researchers should decide what difference would change understanding or action. The meaningful threshold comes from the field, not the software. Bayesian analysis, likelihood methods and decision analysis ask related questions differently, but none removes the need for good data and transparent assumptions.
Keep exploring.
One remarkable idea at a time—nature, science, history and beyond.
Sources and further reading
Barnakle uses credible primary and authoritative sources wherever possible.
- National Academies — Reproducibility and Replicability in Science
- https://www.nationalacademies.org/projects/DBASSE-BBCSS-17-03/publication/25303
- NIH — Rigor and Reproducibility
- https://www.nih.gov/research-training/rigor-reproducibility
- EQUATOR Network — Reporting Guidelines
- https://www.equator-network.org/
- American Statistical Association — Statement on P-Values
- https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf
Last reviewed September 12, 2026.




