The shape of a self-experiment that worked is always the same. A number looked bad, you changed something, the number improved. The usual objection is that you might be fooling yourself about how you feel. That is not the problem here, and treating it as the problem is what lets the real one through. The number moved for a reason that has nothing to do with belief and everything to do with when you decided to start.
What happens when you measure twice and change nothing
Takashima and colleagues took paired annual health-check data from 547 Japanese male clerical workers who had received no hyperlipidaemia medication in the intervening year, sorted them into fifths by their 1998 total cholesterol, and looked at where each fifth sat in 1999 [1].
Both ends moved inward. Nothing was done to either group. Had you handed the top fifth a supplement in 1998 and measured again in 1999, you would be reporting a 16 mg/dl fall, and it would be real in the sense that the number really is lower.
“These results suggest that the observed yearly change in each serum lipid level may largely reflect the 'regression to the mean' effect in addition to the real yearly biological change.”
The size of the pull is set by how noisy the measure is
The same paper reports the correlation between each person's 1998 and 1999 measurement: 0.86 for HDL cholesterol, 0.77 for total cholesterol, 0.68 for log triglycerides [1]. Then it reports how far the gap between the top and bottom fifths closed. HDL cholesterol: a range of 36 mg/dl shrank to 31. Total cholesterol: 92 to 72, taking the quintile means above. Log triglycerides, the noisiest of the three: geometric mean range 253.5 mg/dl down to 154.1.
Divide the surviving gap by the original — arithmetic on their figures, not a claim they make — and you get 0.86, 0.78 and 0.61, against correlations of 0.86, 0.77 and 0.68. The fraction of an extreme baseline that survives a year is roughly the correlation between the two measurements. Everything else was never really there. This is the whole mechanism, and it means the effect is largest exactly where measurement is worst.
Measured directly, in blood pressure
Wang and colleagues ran a post hoc analysis of the BP GUIDE trial, in which 286 patients had office and home blood pressure taken every three months for a year, and grouped them by baseline in 10 mmHg strata [2].
The group starting below 120 mmHg went the other way, 113 to 120. The pattern held in both the intervention and the control arm, and for diastolic pressure [2]. A pooled analysis across five studies found the same thing on ambulatory monitoring — 156 to 141 mmHg for baseline 24-hour systolic of at least 150, regression dilution ratios of 0.52 for 24-hour systolic and 0.38 for night-time, the noisiest window [3]. Those two papers share five authors, so read them as one line of work rather than two independent confirmations, which is roughly what their own authors ask for.
“Replication studies are needed to confirm these findings.”
Daily measurement makes this worse, not better
The intuition that more data protects you is exactly backwards, because more data means more chances to catch yourself at an extreme. Rose and colleagues took 814 people with controlled blood pressure and stable medication across two trials — 232 in HOMEBP, 582 in TASMINH4 — and asked how much a self-monitored reading actually moves [4].
The noise inside one week of your own readings is an order of magnitude larger than the drift you are trying to see. The consequence they draw is the one that matters for anyone watching a dashboard: for a patient starting at 130 mmHg, a raised reading at six months ran at odds of 3 to 1 of being a false positive, only flipping to 2 to 1 in favour of a true positive at twelve months [4]. Their recommendation was to stop looking so often.
What this does not establish
None of this is wearable data. No figure above is the correlation between one week of your deep sleep and the next, or between two months of overnight HRV. That number is not in this literature, and we do not have it either. The mechanism transfers; the magnitude does not, and quoting 0.77 at a sleep tracker would be inventing a statistic.
The lipid cohort is 547 Japanese male clerical workers observed once a year, and its authors are careful to say the yearly change reflects regression to the mean in addition to real biological change — the design does not separate the two. The BP GUIDE analysis is post hoc on a trial built to answer a different question, with 286 people. The five-study pooled analysis sits behind a paywall; the figures cited here come from its published abstract, and a reviewer should treat them as such.
And none of it shows that any intervention did nothing. Regression to the mean does not subtract from the treatment effect. It means a before-and-after difference is not the treatment effect, and that it is biased in the flattering direction whenever you started because things looked bad.
Why this is an n-of-1 problem
A population trial gets this fix for free. Everyone in it was selected by the same rule, so the regression lands in both arms and subtracts out — that is a large part of what the control group is doing, and why the BP GUIDE pattern showed up in the control arm too. Run alone, you have one arm at a time and nothing to subtract from it.
Worse, your selection rule is invisible. A trial writes down its entry criterion; you started the magnesium on a Tuesday because the week had been bad, and that criterion never got recorded anywhere. The rule is doing statistical work and nobody can see it, including you a month later.
The fixes all work by breaking the link between how you felt and when the intervention began: randomise the start date, run several on-and-off periods instead of one before-and-after, and take the baseline from a stretch you did not choose. Averaging more readings into that baseline helps — it raises the correlation, and the correlation is what governs the pull — but it shrinks the effect rather than removing it. So a Baseline trial that was triggered by a bad week is labelled as one, and the estimate we report comes from the periods after the trigger, not from the drop that followed it.
Sources
- 1.Takashima Y, Sumiya Y, Kokaze A, Yoshida M, Ishikawa M, Sekine Y, Akamatsu S. Magnitude of the Regression to the Mean within One-year Intra-individual Changes in Serum Lipid Levels among Japanese Male Workers. Journal of Epidemiology. 2001;11(2):61-69. doi:10.2188/jea.11.61. PMID 11388494. Link ↗
- 2.Wang N, Atkins ER, Salam A, Moore MN, Sharman JE, Rodgers A. Regression to the mean in home blood pressure: analyses of the BP GUIDE study. Journal of Clinical Hypertension. 2020;22(7):1184-1191. doi:10.1111/jch.13933. PMID 32634288. Link ↗
- 3.Moore MN, Atkins ER, Salam A, Callisaya ML, Hare JL, Marwick TH, Nelson MR, Wright L, Sharman JE, Rodgers A. Regression to the mean of repeated ambulatory blood pressure monitoring in five studies. Journal of Hypertension. 2019;37(1):24-29. doi:10.1097/HJH.0000000000001977. PMID 30499921. Link ↗
- 4.Rose F, Stevens RS, Morton KS, Yardley L, McManus RJ. How often should self-monitoring of blood pressure be repeated? A secondary analysis of data from two randomized controlled trials. Journal of Hypertension. 2025;43(11):1863-1870. doi:10.1097/HJH.0000000000004123. Link ↗