Skip to content
Ian Misner Builder, dad, occasional writer

Maintenance · 25 entries · 6 stages

How to keep up without living online

The facts in this field have a shelf life measured in months. These are the things to check on a schedule, so the model of the world you just built keeps getting corrected by reality.

Last on purpose, and the only track with no end. Everything above is a snapshot, and snapshots age quickly here. These are worth checking on a cadence rather than reading once, because the calibration you just built starts decaying when you stop feeding it. If you maintain only three, make them Epoch, METR and the annual safety report. Between them you get the compute and cost picture, the capability picture, and the contested-claims picture, from parties who publish their methods.
Research, policy, builders, newsletters, and public voices feeding a quiet signal monitor.

Choose the useful amount

Two ways through this track

Basic route 7 sources. A shorter route through the load-bearing ideas.
0 of 7 in this route
  1. Notable AI models database
  2. METR
  3. Frontier Model Forum
  4. LMArena
  5. Artificial Analysis
  6. OpenAI model release notes
  7. Import AI
How to read the cards
free some paid library card rent or stream paid
Core read first Reference look things up Risk case / Skeptical argues one side Both sides argues with itself Off axis outside the meter Fiction intuition pump Follow-up keep going
A check means encountered, not mastered. Shared sources stay checked everywhere.
Stage 00

The monthly calibration loop

These are the sources to keep in rotation so your model of the world does not calcify into a handsome little fossil. Capability, compute, safety posture, governance promises: check them before your confidence gets decorative.

AI Index Report

Reference

S01 Stanford HAI · annual · Report

The field's annual census: capability, investment, adoption, cost, policy and public opinion, all sourced. Skim the top takeaways on release, then use it year-round as the reference whenever someone quotes a number at you. The adoption-versus-measured-value gap in the economy chapter is the most interesting figure in it and the one most often skipped.

What to do Read the top-level takeaways, then go straight to the economy chapter and find the gap between reported adoption and measured value. Note that gap as a number. Come back to the rest only when you need a citation — never cover to cover. Skim the takeaways again each year on release.

Also in Start here, Costs and applications.

20 min–2 hr hai.stanford.edu · free

S02 Epoch AI · continuous · Database

Compute, parameters, dataset size and estimated cost for essentially every significant model, with methodology published. Half an hour with the filters gives you a better map of the field than most people writing about it have.

What to do Use the filters, not the summary. Sort by training compute to find the current frontier, then filter to your favourite open-weight model and see how far back it sits. Repeat quarterly. It is the fastest available cure for confidently discussing a model that has already been lapped twice.

Also in The risk argument.

30 min epoch.ai · free

S03 chaired by Yoshua Bengio · 2026 · Report + key updates

Over a hundred experts nominated by thirty-plus countries plus the EU, UN and OECD, with Key Updates through the year when capabilities move. Its most valuable feature is structural: it separates established from contested from speculated and refuses to collapse the third into the first. Read the four-page executive summary, then audit anything you believe confidently against which category it falls in.

What to do Read only the four-page executive summary first. Then pick three claims you currently hold with confidence and find which bucket the report files each in — established, contested, or speculated. Anything you believe confidently that sits under speculated is the finding. Re-run this against the Key Updates when they land.

Also in Start here, The risk argument.

25 min–3 hr internationalaisafetyreport.org · free

METR

Reference

S04 Model Evaluation and Threat Research · per release · Check per launch

Independent pre-deployment evaluation of frontier models for autonomy and dangerous capability. When a major model ships, METR's write-up is usually the most informative independent account of what it can do unsupervised.

What to do Check when a major model ships. Read the autonomy and dangerous-capability findings, and specifically what the model managed unsupervised. It is usually the most informative independent account available in the first week, and it is always more sober than the launch post it accompanies.

20 min metr.org · free

S05 industry body · continuous · Check quarterly

Where Anthropic, Google, Microsoft, OpenAI and others publish shared technical work on safety frameworks, thresholds and evaluation. Read it as the industry's own account of what it has committed to, useful precisely because it is checkable against behaviour later.

What to do Check quarterly. Read it as the industry’s own written record of what it has committed to, and keep a dated note of specific promises. Its value comes from being able to return later and check those promises.

Also in Build and govern.

20 min frontiermodelforum.org · free

S06 Anthropic · continuous · Frontier safety framework

Anthropic’s living policy for escalating safeguards as model capabilities rise. The current version separates capability thresholds, required safeguards, risk reports and public accountability into a company-specific system that can be checked against later releases.

What to do Build a three-column comparison table, one column per lab, with rows for capability thresholds, safeguards, review and public reporting. Fill the Anthropic column first, then diarise a check against the next Claude release to see whether the policy was applied as written. A framework nobody ever audits is marketing with footnotes.

Also in Build and govern.

30 min anthropic.com · free

S07 OpenAI · 2025 · Frontier safety framework

OpenAI's framework for tracking severe-harm capabilities and requiring safeguards before deployment. Version 2 focuses its top-level categories on biological and chemical capability, cybersecurity and AI self-improvement, with risk reports and a Safety Advisory Group built into the process. Compare the categories and governance mechanics directly with the other labs rather than assuming they match.

What to do Read the tracked categories and the Safety Advisory Group’s role, then build the table: three columns, one per lab, with rows for thresholds, review process and public reporting. Fill this column second. The differences are the point, and the apparent sameness is an artifact of summaries.

Also in Build and govern.

30 min openai.com · free

S08 Google DeepMind · 2026 · Frontier safety framework

Google DeepMind's current framework for identifying critical capability levels and pairing them with security and deployment mitigations. Version 3.1 adds and revises protocols as the threat model changes. Its thresholds, review process and terminology are its own, which is exactly why it deserves a separate record.

What to do Third column of that table. Read for critical capability levels and how mitigations are paired to them, and accept that the terminology is deliberately its own. Record what changed from 3.0, because version deltas are the only evidence these documents respond to anything.

Also in Build and govern.

30 min deepmind.google · free

Ask the map

Log in to ask questions against the full reading map.

Log in to ask

Aggregate usage diagnostics are stored; your question and answer text are not.

Stage 01

Quick reality checks

Before opinions, find the tables. Epoch gives you training runs; LM Arena gives you messy preference signals; together they keep you from confidently discussing a model that has already been lapped twice.

LMArena

Reference

S09 formerly Chatbot Arena · continuous · Leaderboard

Blind pairwise human preference at scale, rendered as Elo. Known weaknesses: it rewards style and formatting, it is gameable, and preference is not capability. One instrument, not the only one.

What to do Check it, then discount it. Read the leaderboard and the known-limitations page in the same visit, so the Elo never travels alone in your head. Check monthly, and never use it as your only instrument — it rewards style, and style is not capability.

10 min lmarena.ai · free

LiveBench

Reference

S10 LiveBench · continuous · Leaderboard

A contamination-resistant leaderboard with objective tasks that refresh over time and show cost beside score. Not magic, just a cleaner instrument for a field where the test set keeps getting eaten by the training set like a suspiciously convenient snack.

What to do Bookmark it. On the first visit, sort by cost against score rather than score alone. Check monthly, and treat any large jump as a question about contamination before you treat it as evidence of progress.

Also in How it works.

15 min livebench.ai · free

S11 independent benchmarking · continuous · Dashboard

Quality against price against latency across every major model and provider, updated continuously. The page that turns abstract capability talk into a procurement decision, and where the cost collapse at fixed capability becomes impossible to miss.

What to do Use this the way you would use a supplier quote. Filter to your actual use case, sort cost against quality, and note what last year’s frontier capability costs today. Revisit whenever you are about to make a build-versus-buy call, which is the only time it matters.

15 min artificialanalysis.ai · free

S12 OpenAI · continuous · Release notes

The official running log for ChatGPT model changes, retirements, routing and behaviour updates. It is not a benchmark, but it tells you what users are actually being moved onto, which is often the part missing from lab-chart arguments.

What to do Check whenever a release lands. Read for retirements and routing changes rather than launches — what users are being silently moved onto matters more than what was announced, and it is the part missing from every chart argument.

10 min / check help.openai.com · free

S13 Anthropic · continuous · Model cards

Anthropic's collected model and system cards for Claude releases. Read them when a model launches, mostly for the deltas: capability claims, safety evaluations, refusals, tool use and where the company chose to draw the box around the product.

What to do Read on launch, and read for deltas only. Diff the safety evaluations and refusal behaviour against the previous card. Where the company chose to draw the box around the product is usually the most informative paragraph and never the headline.

15 min platform.claude.com · free

Stage 02

Open is a claim, not a file extension

A downloadable weight file can be genuinely useful without making the whole system transparent. These two instruments separate access, documentation and reproducibility, which is less exciting than shouting open and considerably more informative.

S14 Stanford CRFM · 2025 · Index + report

Scores major foundation-model developers against one hundred disclosure indicators spanning data, compute, model behaviour and downstream use. The 2025 edition makes the useful distinction between releasing weights and documenting enough of the system to support real scrutiny. Treat the scores as a structured disclosure audit, not a universal quality ranking.

What to do Read the indicator list before the scores, then use it as a checklist against one vendor you actually depend on and see what you cannot find out about them. Treat the result as a disclosure audit, never as a quality ranking — the distinction is the source’s main contribution.

30 min crfm.stanford.edu · free

S15 Open Source Initiative · 2024 · Definition

Defines open-source AI around the freedoms to use, study, modify and share, then specifies the data information, code and parameters needed to exercise them. Its sharpest contribution is refusing to call weights alone an open-source system. Read it before using open, open-weight and source-available as interchangeable compliments.

What to do Read it once, then audit your own vocabulary. Find three things you or your company have called open source and classify each as open, open-weight, or source-available. Use the correct word from then on, including when it is less flattering.

15 min opensource.org · free

Stage 03

People who keep touching the wire

A few high-signal people whose public work helps you notice when the frontier has moved. Not scripture. Just useful smoke alarms.

Theo / t3.gg

Reference

S16 Theo Browne · ongoing · Video and developer commentary

A fast, opinionated read on what AI coding tools are doing to actual web development: coding agents, app builders, product decisions and the gap between an impressive demo and a thing someone can ship. Theo has strong preferences and says them at full volume, which is useful data when the tooling changes faster than its documentation.

What to do Check weekly, and calibrate for volume: he is loud, and often right about tooling. Watch specifically for the gap between an impressive demo and a thing somebody can ship, because that gap moves faster than any documentation covering it.

20 min / week t3.gg · free

Import AI

Reference

S17 Jack Clark · weekly · Newsletter

A weekly pass through AI research, policy and industrial weirdness by someone unusually good at making the moving parts legible. It is not neutral in the fake view-from-nowhere sense; it is useful in the someone-read-the-papers-and-has-taste sense. That distinction matters.

What to do Read it weekly. If you keep only one source, keep this one — Clark has taste and states his position, which is more useful than performed neutrality. Watch for shifts in his framing, because he sits close enough to the frontier for a shift to be a signal.

15 min / week jack-clark.net · free

S18 Simon Willison · ongoing · Weblog

The rare feed that catches model releases, local inference, coding agents and prompt-injection failures while the rest of the internet is still composing a launch thread. Practical, unusually well documented and excellent at separating what a tool claims from what happened when somebody actually ran it.

What to do Check twice a week. Read specifically for the difference between what a tool claimed and what happened when he actually ran it. If you keep only one source for practical capability rather than argument, this is the one.

10 min / check simonwillison.net · free

Latent Space

Reference

S19 swyx and rotating co-hosts · ongoing · Podcast and newsletter

AI engineers interviewing the people building model, agent and infrastructure systems. It is especially good for learning the vocabulary of a new technical pattern shortly before that vocabulary gets flattened into brochure copy.

What to do One episode a week, chosen by topic rather than guest. Its purpose here is early vocabulary: you learn what a new technical pattern is called shortly before the word gets flattened into brochure copy, and that head start is most of the value.

45 min / week latent.space · free

Dwarkesh Podcast

Reference

S20 Dwarkesh Patel · ongoing · Long-form interview podcast

Long interviews with lab leaders, researchers and unusually consequential operators, with enough room for claims to acquire assumptions and caveats. Not a news feed. It is where you go for context after one claim has eaten the entire week.

What to do Not a weekly habit, and do not treat it as news. Go here after a single claim has eaten a week of your attention, and listen to the episode where somebody had enough room to acquire caveats. Listen on purpose rather than on schedule.

1 hr–3 hr / episode dwarkesh.com · free

Keep the alarms working

Suggest a source

See something missing? Send the source, not a biography: one link and one sentence about what it catches that this map currently misses.

Suggest on X
Stage 04

Longer listens and useful counterweights

Podcasts and courses are scaffolding, not scripture. These three are useful when you need a wider-angle news pass, a skeptical temperature check, or a structured way into the safety argument before it hardens into team merchandise.

Hard Fork

Reference

S21 Kevin Roose & Casey Newton, NYT · 2022– · Podcast, weekly

Two Times journalists on the week's releases, lawsuits, policy and drama. Sits deliberately in the middle and treats both camps as subjects rather than allies, which annoys everyone. Its real function here is as a gauge of which arguments are reaching people who have read nothing else on this page.

What to do Use it weekly as an instrument rather than a source. It tells you which arguments have reached people who have read nothing else on this page, which is exactly what you need to know before speaking to any of them.

Free in any podcast app; no Times subscription needed for the audio.

1 hr / week nytimes.com · free

S22 Emily M. Bender & Alex Hanna · 2022– · Podcast

Each episode takes one claim, paper or press release and goes through it line by line. The standing position: text synthesis machines, marketing capability claims, benchmarks measuring the wrong thing, risk discourse functioning as advertising. Contemptuous by design. A reliable corrective after a month inside the risk literature, and unlistenable without one.

What to do Prescribed dosage: one episode after any month spent inside the risk literature. The contempt is both the point and the limit. Extract the line-by-line method — take one claim and check what it actually says — then apply it to a source you trust rather than one you already dislike.

1 hr / episode DAIR Institute · free

S23 BlueDot Impact · current · Course

A facilitated course covering technical alignment, governance and current research, in small cohorts with weekly discussion. Free, applications open periodically. Descended from Richard Ngo's original curriculum and the route by which much of the field entered it. Being argued with weekly is the only reliable defence against the drift this page keeps warning you about.

What to do Applications are free and open periodically. Apply if you want the one thing this page structurally cannot give you: being argued with weekly by people who have read the same material. Reading alone can drift. A cohort gives that drift a useful correction.

12 weeks bluedot.org · free

Stage 05

Then keep a human in the loop

The last two are mine because this whole map is a living object, not a museum label. If I keep doing the job properly, the newsletter and Twitter feed are where the corrections, additions and embarrassed revisions show up first.

NEXT Ian Misner · current · Newsletter

The most direct way to keep this orientation alive after the reading map: get future notes, updates and revisions by email. This belongs at the end because the useful version of this project is maintained, not finished.

What to do This is the mechanism by which the map stays a live document rather than something a reader finished once in 2026 and quietly stopped trusting.

Newsletter signup page using the site's Mailchimp form.

1 min ianmisner.com · free

NEXT Ian Misner · current · Social feed

The lighter-weight currentness layer: short notes, links and half-formed updates, usually before they become posts or revisions. Useful if you want the reading map connected to what I am noticing in public.

What to do Optional, and lower signal than the newsletter by design. Follow if you want the pre-draft layer: what is being noticed before it becomes a revision here.

Ongoing Twitter/X · free

Cost: free · some paid · library card · rent or stream · paid.

Videos and PDFs can open in place. Everything else opens at the original source in a new tab.

This is a snapshot of a field that moves monthly. Keep current is the maintenance layer.