Every consumer sleep tracker reports a deep-sleep number as though it were a measurement. It is an inference — typically from movement, heart rate, and heart rate variability — trained to approximate what a sleep laboratory would have scored from EEG. A 2025 validation study put six wrist-worn devices against simultaneous polysomnography and reported how far off they were [1].
Agreement, overall
Cohen's kappa across all sleep stages ranged from 0.53 for the Apple Watch Series 8 down to 0.21 for the Garmin Vivosmart 4 — moderate at the top end, fair at the bottom. For reference, 1.0 would be perfect agreement and 0 would be chance.
Deep sleep specifically
The per-device bias against polysomnography, in minutes of deep sleep per night, is the part worth sitting with:
Two devices came close: the Fitbit Sense was off by −3.89 minutes and the Fitbit Charge 5 by −2.19, neither significantly. The spread across the group is the point. A supplement that genuinely moved deep sleep by fifteen minutes a night would be a substantial finding, and it is smaller than the error on several of these devices.
Why n-of-1 survives this
Because a personal trial does not ask what your deep sleep is. It asks whether your deep sleep changed. Those questions have different requirements. The first needs accuracy — closeness to the truth. The second needs reliability — an error that stays put.
If a device overstates your deep sleep by roughly half an hour every night, that offset appears on both sides of a within-person comparison and largely cancels. You are contrasting nights against nights on the same wrist, with the same firmware and the same body. The bias that ruins the absolute number is mostly harmless to the difference.
The study's authors land in the same place. They are explicit that these devices cannot replace polysomnography for clinical diagnosis, while concluding that the better-performing ones “could be effectively used to track prolonged and significant changes in sleep architecture”. Tracking change over many nights is precisely the n-of-1 use case.
Where the argument breaks
The cancellation depends on the offset being constant. It is not guaranteed to be. These devices infer sleep stages partly from heart rate and heart rate variability, so an intervention that shifts your cardiac signals can shift the inference independently of any real change in sleep architecture. Alcohol does this. So do some supplements. The device then reports a change that is an artefact of how it estimates, not a change in what it is estimating.
That is confounding routed through the instrument, and no amount of within-person contrast removes it. It is the honest limit of wearable-based inference, and the reason a result that hinges entirely on a single derived metric deserves less confidence than one corroborated by something measured more directly.
Practically: we hold the device constant across a trial, we do not treat absolute stage minutes as truth, and we report change against your own baseline rather than against a population norm.
Sources
- 1.Schyvens A-M, Peters B, Van Oost NC, Aerts J-M, Masci F, Neven A, Dirix H, Wets G, Ross V, Verbraecken J. A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography. SLEEP Advances. 2025;6(2):zpaf021. doi:10.1093/sleepadvances/zpaf021. Link ↗