Blog
Latest updates from Boundless Intuition Labs.
Benchmarks, verification results, and the failures we found along the way - published as we finish a report, not on a schedule.
ResearchJul 17, 2026·18 min read
Fluent Is Not the Same as Correct
Two frontier Claude models fail an independently authored airline fee benchmark on the same cases, landing on the same wrong dollar figure - and a proof-carrying kernel takes both to 100%.
ResearchJun 19, 2026·15 min read
A Diagnosis Should Be a Proof, Not a Probability
A frontier model gave five different diagnoses to the same patient, five times. We built a clinical classifier whose verdicts are proven in Lean 4, not sampled, and benchmarked it head-to-head.

