baseline
Open benchmark

N1Bench

v0.4Updated 14 JUL 2026240 held-out records

The first open benchmark for causal reasoning over personal health data. We publish our own scores next to every frontier model — a system that tells you what’s working should be able to prove it doesn’t tell you things that aren’t.

#
System
Effect
Calibr.
01
Baseline v0.4 (hierarchical)
0.81
0.94
02
Frontier model A + tools
0.63
0.71
03
Frontier model B
0.58
0.66
04
Frontier model C
0.52
0.61
05
Correlation only
0.24
0.19

Abstain and Overall columns are shown on wider screens.

What we measure

Effect

How close the recovered effect size sits to ground truth on records where a real effect was planted. Scored on absolute error, normalised by the outcome's own night-to-night variation.

Calibration

Whether a stated 95% interval contains the truth 95% of the time. Penalises overconfidence and uselessly wide intervals equally — an interval that always covers by being enormous scores no better than one that never covers.

Abstain

On records with no detectable effect, how often the system declines to conclude instead of inventing one. This is the metric most systems fail, and the reason a leaderboard is worth publishing at all.

What each part is worth

The same model with one component removed at a time, scored on the same held-out records.

Configuration
Overall
Δ
v0.4, full
0.88
− population pooling
0.79
−0.09
− confounder adjustment
0.71
−0.17
− hierarchical prior (flat)
0.66
−0.22
correlation only
0.15
−0.73
Submitting a system

The held-out records and the scoring harness are open. Run your system against them and send the output — we publish every submission, including the ones that beat us.

Read the v0.4 write-up →