baseline
Benchmark

N1Bench v0.4: what frontier models get wrong

02 JUN 202615 min

Every system we tested infers causation from correlation somewhere. Here is where each one breaks, and how we score calibration.

Not yet published

This one is still being written. Join the waitlist and it will reach you when it goes out.

Join the waitlist →
← All notes