baseline
Evidence

HRV is not one number, and your device chose which one

14 AUG 20267 min

Across 536 nights against a chest-strap ECG, five wearables estimated nocturnal RMSSD to within 5 ms — and still differed significantly from each other, because they compute a different metric over a different slice of the night from a signal that is not the heartbeat.

Your ring reports an HRV of 42 and the obvious reading is that your heart rate variability last night was 42. Every part of that sentence is doing more work than it appears to. The device did not observe your heartbeat — it observed a pulse in the vessels under a sensor. It did not measure variability in general — it computed one metric from a family of them. And it computed that not over the night, but over whichever slice of the night its manufacturer chose. Before two HRV numbers are compared, those three choices have to match.

Which metric

HRV names a family. The time-domain measures include SDNN, RMSSD and pNN50, the frequency-domain measures LF and HF power [2]. These are not versions of one quantity. RMSSD reflects beat-to-beat change and is vagally mediated. SDNN describes overall variability across the whole recording, and grows as the recording gets longer. LF is worse still as a summary — Shaffer and Ginsberg report that roughly half of it reflects a parasympathetic contribution, and that the LF/HF ratio as a measure of sympathovagal balance has been directly challenged [2]. Consumer devices have largely settled on RMSSD in milliseconds, which is the only reason a cross-device comparison is discussable at all [1].

Over which window

It is inappropriate to compare metrics like SDNN when they are calculated from epochs of different length.
Shaffer & Ginsberg, 2017

How much the window matters depends on the metric. In 3,387 adults from the PREVEND cohort, RMSSD from a single 10-second ECG correlated at 0.853–0.862 with the 240–300 second reference, with a standardised bias of d = 0.15–0.17. SDNN over the same 10 seconds correlated at 0.758–0.764 with a standardised bias of d = 0.86–0.89 — a large discrepancy where RMSSD's was negligible. The authors recommend at least 30 seconds for SDNN and accept a single 10 seconds for RMSSD [3].

Overnight, devices sample different hours. Oura's rings compute HRV from 5-minute samples averaged across the entire night. The Garmin Fenix 6 measures intermittently across detected sleep. The Polar Grit X Pro uses only a 4-hour window after sleep onset, and WHOOP 4.0 weights its estimate dynamically towards the last slow-wave sleep episode [1]. Two devices on the same person on the same night are summarising different periods of that night.

And not quite the same signal

Optical sensors record pulse rate variability, not heart rate variability. Yuda and colleagues trace six conversion steps between the electrical beat and the light absorbed at the skin, from electromechanical coupling through wave propagation to the optical reading, each with its own transfer function modulated by respiration, blood pressure and vascular compliance. Their clearest case is a patient paced at a fixed rate, in whom heart rate variability was absent while pulse transit time still oscillated [4].

In 54 adults measured upright, PRV metrics at 1000 Hz sat about 1–8% above the ECG-derived equivalents [5].

100–200 Hz
Sampling rate needed for PRV bias under 2% · PPG devices commonly sample 20–100 Hz · n = 54
PRV and HRV were not surrogate biomarkers due to the different nature of the collected waveforms.
Burma et al., 2024

How close the devices actually get

Given all of that, the empirical agreement is better than the theory threatens. Thirteen healthy adults wore five devices at home for 536 nights against a Polar H10 chest strap sampling at 1000 Hz [1]. Every device ran slightly below the reference on nocturnal RMSSD.

−0.96 ms
Oura Gen 4 vs ECG · 95% limits of agreement −11.78 to 9.85 ms · 139 nights
−0.78 ms
WHOOP 4.0 vs ECG · 95% limits of agreement −12.50 to 10.94 ms · 289 nights

The Oura Gen 3 was off by −2.50 ms across 470 nights and the Garmin Fenix 6 by −1.84 ms across 150. Mean absolute percentage error separates them more clearly than bias does: 5.96% for the Oura Gen 4, 7.15% for the Gen 3, 8.17% for WHOOP, 10.52% for Garmin, and 16.32% for the Polar Grit X Pro, whose −4.65 ms bias was the largest of the five. Those differences between devices were statistically significant, and the authors' position is that the devices are not interchangeable.

They do not describe how signal artifact is filtered, how signal quality is interpreted, or how interpolation of missing data is conducted.
Dial et al., 2025, on the device manufacturers

What this does not establish

Thirteen participants, all apparently healthy, aged 33.2 ± 8.6 years. Nights were unevenly distributed — 470 for the Oura Gen 3 against 139 for the Gen 4 — so the precision behind each figure differs. The reference was a research-grade chest strap rather than clinical multi-lead ECG. The authors note that they could not verify how any device handles artefact or missing data, and that firmware updates change the calculations over time, so these figures describe the devices as they were, not as they are.

We have not quoted the Polar Grit X Pro's published limits of agreement, because the printed interval is not consistent with its reported bias and standard deviation. The sampling-rate result is also weaker than it looks for wearables: it comes from finger and cerebral-artery waveforms recorded in an upright posture for five minutes and downsampled offline, not from a wrist or ring sensor on a sleeping person. And none of this addresses the separate question of what HRV predicts.

Why this is an n-of-1 problem

42 ± 15 ms
Short-term RMSSD in healthy adults · range 19 to 75 ms · 44 studies, n = 21,438

That range is why a population comparison tells you almost nothing. One healthy adult can sit at nearly four times another's RMSSD and both are unremarkable. What carries information is your number against your own history, and that only holds if the instrument, the metric, the window and the conditions stay fixed. A device-specific offset of a few milliseconds sits on both sides of a within-person contrast and largely cancels — it does not cancel across a change of device.

So we record which device and which firmware produced every night in a trial, and we treat a device change or a major algorithm update as the start of a new baseline rather than a continuation of the old one.

Sources

  1. 1.Dial MB, Hollander ME, Vatne EA, Emerson AM, Edwards NA, Hagen JA. Validation of nocturnal resting heart rate and heart rate variability in consumer wearables. Physiological Reports. 2025;13(16):e70527. doi:10.14814/phy2.70527 Link ↗
  2. 2.Shaffer F, Ginsberg JP. An overview of heart rate variability metrics and norms. Frontiers in Public Health. 2017;5:258. doi:10.3389/fpubh.2017.00258 Link ↗
  3. 3.Munoz ML, van Roon A, Riese H, et al. Validity of (ultra-)short recordings for heart rate variability measurements. PLoS One. 2015;10(9):e0138921. doi:10.1371/journal.pone.0138921 Link ↗
  4. 4.Yuda E, Shibata M, Ogata Y, Ueda N, Yambe T, Yoshizawa M, Hayano J. Pulse rate variability: a new biomarker, not a surrogate for heart rate variability. Journal of Physiological Anthropology. 2020;39:21. doi:10.1186/s40101-020-00233-x Link ↗
  5. 5.Burma JS, Griffiths JK, Lapointe AP, et al. Heart rate variability and pulse rate variability: do anatomical location and sampling rate matter? Sensors. 2024;24(7):2048. doi:10.3390/s24072048 Link ↗
← All notes