Blog
July 2026
TriviaQA Audit Report
TriviaQA is one of the classic QA benchmarks. We discuss its history, how it is scored, and whether it has any signal left at the frontier.
Read moreJuly 2026 · By James Mann
Forecasting the Remote Labor Index
Benchmarks can forecast AI progress, not just track it. We forecast the Remote Labor Index — a measure of how much real remote work AI can do — and stress-test the method on benchmarks that have already saturated.
Read moreJune 2026 · By Jay Bailey
Why Are Evaluations Broken?
Why are so many AI evaluations broken, and how can we improve on this problem? We explore the root causes and share our approach to building better evaluations.
Read more