Ian Misner Builder, dad, occasional writer

03 · a weekend

Knowing where the frontier actually is

Now you can read the scoreboard: which models exist, what they cost, and what they can actually be shown to do. The half-life on this knowledge is short, so the point is not memorising the facts. The point is building the habit of checking.

Most confident statements about AI capability come from people whose picture is eighteen months old, which here is a very long time. The instruments and the habit: which models exist, what they cost, what they can be shown to do. Sources over opinions.

00

The frontier is a database problem first

Before opinions, find the tables. Epoch gives you training runs; LM Arena gives you messy preference signals; together they keep you from confidently discussing a model that has already been lapped twice.

LMArena

02 formerly Chatbot Arena · continuous · Leaderboard

Blind pairwise human preference at scale, rendered as Elo. Known weaknesses: it rewards style and formatting, it is gameable, and preference is not capability. One instrument, not the only one.

Also in Keep current. Ticking it here marks it there.

10 min lmarena.ai · free

Reference

LiveBench

03 LiveBench · continuous · Leaderboard

A contamination-resistant leaderboard with objective tasks that refresh over time and show cost beside score. Not magic, just a cleaner instrument for a field where the test set keeps getting eaten by the training set like a suspiciously convenient snack.

Also in Evals & benchmarks. Ticking it here marks it there.

15 min livebench.ai · free

Reference

Ask the map

Ready with the full reading map.

Aggregate usage diagnostics are stored; your question and answer text are not.

01

Then make it a habit

The useful skill is not memorising this month's winner. It is knowing where the frontier gets reported, how stale the number is, and when the impressive demo is mostly fog machine with invoice attached.

AI Index Report

05 Stanford HAI · annual · Report

The field's annual census: capability, investment, adoption, cost, policy and public opinion, all sourced. Skim the top takeaways on release, then use it year-round as the reference whenever someone quotes a number at you. The adoption-versus-measured-value gap in the economy chapter is the most interesting figure in it and the one most often skipped.

Also in Start here, Economics & infrastructure, Applications, Keep current. Ticking it here marks it there.

20 min – 2 hr hai.stanford.edu · free

Reference
02

Open is a claim, not a file extension

A downloadable weight file can be genuinely useful without making the whole system transparent. These two instruments separate access, documentation and reproducibility, which is less exciting than shouting open and considerably more informative.

Foundation Model Transparency Index 2025

08 Stanford CRFM · 2025 · Index + report

Scores major foundation-model developers against one hundred disclosure indicators spanning data, compute, model behavior and downstream use. The 2025 edition makes the useful distinction between releasing weights and documenting enough of the system to support real scrutiny. Treat the scores as a structured disclosure audit, not a universal quality ranking.

Source PDF

30 min crfm.stanford.edu · free

Reference

The Open Source AI Definition 1.0

09 Open Source Initiative · 2024 · Definition

Defines open-source AI around the freedoms to use, study, modify and share, then specifies the data information, code and parameters needed to exercise them. Its sharpest contribution is refusing to call weights alone an open-source system. Read it before using open, open-weight and source-available as interchangeable compliments.

15 min opensource.org · free

Reference

Cost: free · library card · rent or stream · paid.

Videos and PDFs can open in place. Everything else opens at the original source in a new tab.

This is a snapshot of a field that moves monthly. Keep current is the maintenance layer.