In 2016 the American Statistical Association took the unusual step of issuing a formal statement about a single statistic [1]. The provocation was that p-values had come to carry weight they cannot bear, and the statement is blunt about the limits.
“A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.”
And, addressing the inference people most often draw in both directions:
“Smaller p-values do not necessarily imply the presence of larger or more important effects, and larger p-values do not imply a lack of importance or even lack of effect.”
What a threshold discards
Take a real result. The 2021 meta-analysis of magnesium for insomnia found total sleep time increased by 16.06 minutes against placebo, with a 95% confidence interval running from −5.99 to +38.12 minutes and p = 0.15 [2].
Run that through a significance threshold and it becomes four words: no significant effect. Those four words are compatible with the data. They also throw away everything a person needed. The interval says the evidence is consistent with losing six minutes and with gaining thirty-eight. Someone deciding whether to keep taking a capsule every night wants to know that the plausible range is that wide, because the top of it would be worth having and the bottom of it would not.
“Not significant” and “no effect” are different claims. The threshold makes them look like the same claim, which is how absence of evidence gets reported as evidence of absence.
What we report instead
Three numbers, always: the effect size, a 95% credible interval around it, and the posterior probability that the effect is beneficial. No threshold, no verdict word doing the work of a distribution.
The distinction between a credible interval and a confidence interval matters more here than it usually does. A confidence interval is a statement about a procedure: intervals built this way would contain the true value 95% of the time across repeated experiments. A credible interval is a statement about the parameter: given this data and these priors, there is a 95% probability the effect lies in this range. The second is what people already assume the first means, and it is the question someone is actually asking about their own body.
What it costs
It costs the headline. We cannot tell you magnesium works. We can tell you the estimate is +18 minutes a night, that the plausible range runs from +6 to +31, and that there is a 97% posterior probability the effect is positive. That is harder to put on a landing page and considerably more useful for deciding what to do on Tuesday.
It also costs certainty in the other direction. Sometimes the honest output is an interval so wide that no action follows from it. We report that as unsettled rather than rounding it to the nearest verdict, and then design the protocol that would narrow it.
The ASA's closing principle is the one we try hardest to keep: no single index should substitute for scientific reasoning. An interval is not a verdict either. It is the shape of what you currently know.
Sources
- 1.Wasserstein RL, Lazar NA. The ASA's Statement on p-Values: Context, Process, and Purpose. The American Statistician. 2016;70(2):129-133. doi:10.1080/00031305.2016.1154108. Link ↗
- 2.Mah J, Pitre T. Oral magnesium supplementation for insomnia in older adults: a Systematic Review & Meta-Analysis. BMC Complement Med Ther. 2021;21:125. doi:10.1186/s12906-021-03297-z. Link ↗