Ian Misner Builder, dad, occasional writer

02 · a weekend

How anyone knows what a model can do

Every confident number you have heard about AI came out of a benchmark, and benchmarks are more fragile than the numbers make them look. This track is about learning to squint before you read a leaderboard, so the leaderboard does not do all of your thinking for you.

Nearly every argument here cashes out as a claim about measurement, and measurement is usually weaker than the confident numbers suggest. The major instruments, what each tests, and the standing critique that the field measures what is convenient before it measures what matters. Ask what a saturating benchmark actually tells you. Usually less than the press release implies.

00

Three instruments everyone waves around

SWE-bench, GPQA and LiveBench are useful, famous and extremely easy to overread. They tell you something real. They do not tell you the machine has become your new coworker, doctor, lawyer and unsettlingly intense group-project partner.

SWE-bench

01 Jimenez, Yang et al., Princeton · 2023 · Benchmark

Real GitHub issues from real repositories, scored by whether the patch makes the actual test suite pass. The most economically meaningful benchmark in wide use, because passing it is close to doing the job. Also the one where contamination and scaffolding matter most, so check how a score was produced before believing it.

45 min swebench.com · free

Reference

LiveBench

03 LiveBench · continuous · Leaderboard

A contamination-resistant leaderboard with objective tasks that refresh over time and show cost beside score. Not magic, just a cleaner instrument for a field where the test set keeps getting eaten by the training set like a suspiciously convenient snack.

Also in Frontier tracking. Ticking it here marks it there.

15 min livebench.ai · free

Reference

Ask the map

Ready with the full reading map.

Aggregate usage diagnostics are stored; your question and answer text are not.

01

How to test without fooling yourself

A benchmark is a claim about the world with a spreadsheet attached. Inspect is the machinery for making that claim responsibly; Raji is the reminder that the machinery still has assumptions hiding under the rug.

Inspect

05 UK AI Security Institute · 2024– · Framework

The evaluation framework a national institute actually uses, released open source. Worth an hour even if you never run it: reading the abstractions teaches you what a rigorous eval consists of and why most internal ones are not that.

1 hr inspect.aisi.org.uk · free

Reference
A collection of measuring tools and benchmark shapes.
Measurement is an argument with instruments attached.
02

Then ask what autonomy changes

Hazard evaluations are where the discourse gets less tidy. The question is not whether a model can answer exam questions; it is whether it can pursue messy goals through tools, time and resistance. Fun little distinction. Very normal hobby.

Measuring AI ability to complete long tasks

07 Kwa, West, Becker et al. (METR) · 2025 · Paper

Abandons benchmark scores for task duration: which lengths of human-professional work a model finishes at 50% reliability. That horizon has doubled roughly every seven months for six years, through several architecture changes. The caveat almost every citation drops is that the 80% reliability curve sits about two doublings, or a year, behind.

Also in Start here, The risk argument. Ticking it here marks it there.

Source PDF

30 min metr.org · free

Reference

OSWorld

08 Xie, Zhang, Chen et al. · 2024 · Agent benchmark

Evaluates multimodal agents on open-ended tasks inside real computer environments across operating systems and applications. It is a more revealing test of agents than another set of exam questions because success requires perception, tool use and long action sequences. The benchmark is still a controlled proxy, but at least the proxy has windows, menus and consequences.

25 min github.com/xlang-ai · free

Reference

GDPval

09 OpenAI · 2025 · Evaluation + paper

An attempt to measure economically valuable work instead of exam-shaped cleverness: realistic deliverables across 44 knowledge-work occupations, judged against expert outputs. The limitations matter as much as the scores, because the real world cruelly insists on iteration, ambiguity and people changing their minds after lunch.

Also in Start here, Economics & infrastructure, Applications. Ticking it here marks it there.

20 min openai.com · free

Reference

Cost: free · library card · rent or stream · paid.

Videos and PDFs can open in place. Everything else opens at the original source in a new tab.

This is a snapshot of a field that moves monthly. Keep current is the maintenance layer.