Every confident number you have heard about AI came out of a benchmark, and benchmarks are more fragile than the numbers make them look. This track is about learning to squint before you read a leaderboard, so the leaderboard does not do all of your thinking for you.
Share a checkpointCopy a grid, image card, or short progress reflection.
Nearly every argument here cashes out as a claim about measurement, and measurement is usually weaker than the confident numbers suggest. The major instruments, what each tests, and the standing critique that the field measures what is convenient before it measures what matters. Ask what a saturating benchmark actually tells you. Usually less than the press release implies.
00
Three instruments everyone waves around
SWE-bench, GPQA and LiveBench are useful, famous and extremely easy to overread. They tell you something real. They do not tell you the machine has become your new coworker, doctor, lawyer and unsettlingly intense group-project partner.
01
Jimenez, Yang et al., Princeton · 2023 · Benchmark
Real GitHub issues from real repositories, scored by whether the patch makes the actual test suite pass. The most economically meaningful benchmark in wide use, because passing it is close to doing the job. Also the one where contamination and scaffolding matter most, so check how a score was produced before believing it.
Questions written by domain PhDs and validated so that non-experts with unrestricted web access still fail most of them. Built to survive the saturation that killed earlier knowledge benchmarks, and a good case study in how hard it is to build a test that stays hard.
A contamination-resistant leaderboard with objective tasks that refresh over time and show cost beside score. Not magic, just a cleaner instrument for a field where the test set keeps getting eaten by the training set like a suspiciously convenient snack.
Aggregate usage diagnostics are stored; your question and answer text are not.
01
How to test without fooling yourself
A benchmark is a claim about the world with a spreadsheet attached. Inspect is the machinery for making that claim responsibly; Raji is the reminder that the machinery still has assumptions hiding under the rug.
The benchmark family that insists capability is not one number. Accuracy, robustness, calibration, fairness, toxicity, efficiency, multilinguality, domains, modalities: all the annoying dimensions that make a leaderboard less tweetable and more useful. A good antidote to scoreboard intoxication.
The evaluation framework a national institute actually uses, released open source. Worth an hour even if you never run it: reading the abstractions teaches you what a rigorous eval consists of and why most internal ones are not that.
06
Raji, Bender, Paullada, Denton & Hanna · 2021 · Paper
The standing critique: benchmarks claiming generality are built from narrow convenient data, then treated as measuring the whole construct, and the incentive structure rewards this. Read it before you next cite a score as evidence of anything general.
Measurement is an argument with instruments attached.
02
Then ask what autonomy changes
Hazard evaluations are where the discourse gets less tidy. The question is not whether a model can answer exam questions; it is whether it can pursue messy goals through tools, time and resistance. Fun little distinction. Very normal hobby.
Abandons benchmark scores for task duration: which lengths of human-professional work a model finishes at 50% reliability. That horizon has doubled roughly every seven months for six years, through several architecture changes. The caveat almost every citation drops is that the 80% reliability curve sits about two doublings, or a year, behind.
Evaluates multimodal agents on open-ended tasks inside real computer environments across operating systems and applications. It is a more revealing test of agents than another set of exam questions because success requires perception, tool use and long action sequences. The benchmark is still a controlled proxy, but at least the proxy has windows, menus and consequences.
An attempt to measure economically valuable work instead of exam-shaped cleverness: realistic deliverables across 44 knowledge-work occupations, judged against expert outputs. The limitations matter as much as the scores, because the real world cruelly insists on iteration, ambiguity and people changing their minds after lunch.