Ian Misner Builder, dad, occasional writer

Ongoing · start now, keep forever

The currentness layer

The facts in this field have a shelf life measured in months. These are the things to check on a schedule, so the model of the world you just spent a month building keeps getting corrected by reality.

Last on purpose, and the only track with no end. Everything above is a snapshot, and snapshots age quickly here. These are worth checking on a cadence rather than reading once, because the calibration you just built starts decaying when you stop feeding it. If you maintain only three, make them Epoch, METR and the annual safety report. Between them you get the compute and cost picture, the capability picture, and the contested-claims picture, from parties who publish their methods.

00

The monthly calibration loop

These are the sources to keep in rotation so your model of the world does not calcify into a handsome little fossil. Capability, compute, safety posture, governance promises: check them before your confidence gets decorative.

AI Index Report

01 Stanford HAI · annual · Report

The field's annual census: capability, investment, adoption, cost, policy and public opinion, all sourced. Skim the top takeaways on release, then use it year-round as the reference whenever someone quotes a number at you. The adoption-versus-measured-value gap in the economy chapter is the most interesting figure in it and the one most often skipped.

Also in Start here, Frontier tracking, Economics & infrastructure, Applications. Ticking it here marks it there.

20 min – 2 hr hai.stanford.edu · free

Reference

International AI Safety Report

03 chaired by Yoshua Bengio · 2026 · Report + key updates

Over a hundred experts nominated by thirty-plus countries plus the EU, UN and OECD, with Key Updates through the year when capabilities move. Its most valuable feature is structural: it separates established from contested from speculated and refuses to collapse the third into the first. Read the four-page executive summary, then audit anything you believe confidently against which category it falls in.

Also in Start here, The risk argument. Ticking it here marks it there.

25 min – 3 hr internationalaisafetyreport.org · free

Reference

METR

04 Model Evaluation and Threat Research · per release · Check per launch

Independent pre-deployment evaluation of frontier models for autonomy and dangerous capability. When a major model ships, METR's write-up is usually the most informative independent account of what it can do unsupervised.

20 min metr.org · free

Reference

Anthropic’s Responsible Scaling Policy: Version 3.0

06 Anthropic · 2026 · Frontier safety framework

Anthropic's current policy for escalating safeguards as model capabilities rise. Version 3.0 separates capability thresholds, required safeguards and public reporting into a company-specific system that can be checked against later releases. Read the actual policy rather than treating three labs' differently structured promises as one document.

Also in Policy & governance. Ticking it here marks it there.

30 min anthropic.com · free

Reference

Our Updated Preparedness Framework

07 OpenAI · 2025 · Frontier safety framework

OpenAI's framework for tracking severe-harm capabilities and requiring safeguards before deployment. Version 2 focuses its top-level categories on biological and chemical capability, cybersecurity and AI self-improvement, with risk reports and a Safety Advisory Group built into the process. Compare the categories and governance mechanics directly with the other labs rather than assuming they match.

Also in Policy & governance. Ticking it here marks it there.

Source PDF

30 min openai.com · free

Reference

Frontier Safety Framework 3.1

08 Google DeepMind · 2026 · Frontier safety framework

Google DeepMind's current framework for identifying critical capability levels and pairing them with security and deployment mitigations. Version 3.1 adds and revises protocols as the threat model changes. Its thresholds, review process and terminology are its own, which is exactly why it deserves a separate record.

Also in Policy & governance. Ticking it here marks it there.

30 min deepmind.google · free

Reference

Ask the map

Ready with the full reading map.

Aggregate usage diagnostics are stored; your question and answer text are not.

01

The quick reality checks

Leaderboards and model cards are not truth serum, but they are excellent smoke alarms. They tell you when the frontier moved, when prices changed, and when last quarter's argument has started wearing an antique hat.

LMArena

09 formerly Chatbot Arena · continuous · Leaderboard

Blind pairwise human preference at scale, rendered as Elo. Known weaknesses: it rewards style and formatting, it is gameable, and preference is not capability. One instrument, not the only one.

Also in Frontier tracking. Ticking it here marks it there.

10 min lmarena.ai · free

Reference
02

People who keep touching the wire

A few high-signal people whose public work helps you notice when the frontier has moved. Not scripture. Just useful smoke alarms.

Theo / t3.gg

13 Theo Browne · ongoing · Video and developer commentary

A fast, opinionated read on what AI coding tools are doing to actual web development: coding agents, app builders, product decisions and the gap between an impressive demo and a thing someone can ship. Theo has strong preferences and says them at full volume, which is useful data when the tooling changes faster than its documentation.

Open

20 min/week t3.gg · free

Reference

Import AI

14 Jack Clark · weekly · Newsletter

A weekly pass through AI research, policy and industrial weirdness by someone unusually good at making the moving parts legible. It is not neutral in the fake view-from-nowhere sense; it is useful in the someone-read-the-papers-and-has-taste sense. That distinction matters.

Open

15 min/week jack-clark.net · free

Reference

Simon Willison’s Weblog

15 Simon Willison · ongoing · Weblog

The rare feed that catches model releases, local inference, coding agents and prompt-injection failures while the rest of the internet is still composing a launch thread. Practical, unusually well documented and excellent at separating what a tool claims from what happened when somebody actually ran it.

Open

10 min/check simonwillison.net · free

Reference

Latent Space

16 swyx and rotating co-hosts · ongoing · Podcast and newsletter

AI engineers interviewing the people building model, agent and infrastructure systems. It is especially good for learning the vocabulary of a new technical pattern shortly before that vocabulary gets flattened into brochure copy.

Open

45 min/week latent.space · free

Reference

Dwarkesh Podcast

17 Dwarkesh Patel · ongoing · Long-form interview podcast

Long interviews with lab leaders, researchers and unusually consequential operators, with enough room for claims to acquire assumptions and caveats. Not a news feed. It is where you go for context after one claim has eaten the entire week.

Open

1-3 hr/episode dwarkesh.com · free

Reference
03

Longer listens and useful counterweights

Podcasts and courses are scaffolding, not scripture. These three are useful when you need a wider-angle news pass, a skeptical temperature check, or a structured way into the safety argument before it hardens into team merchandise.

Hard Fork

18 Kevin Roose & Casey Newton, NYT · 2022– · Podcast, weekly

Two Times journalists on the week's releases, lawsuits, policy and drama. Sits deliberately in the middle and treats both camps as subjects rather than allies, which annoys everyone. Its real function here is as a gauge of which arguments are reaching people who have read nothing else on this page.

Free in any podcast app; no Times subscription needed for the audio.

1 hr/wk nytimes.com · free

Reference

Mystery AI Hype Theater 3000

19 Emily M. Bender & Alex Hanna · 2022– · Podcast

Each episode takes one claim, paper or press release and goes through it line by line. The standing position: text synthesis machines, marketing capability claims, benchmarks measuring the wrong thing, risk discourse functioning as advertising. Contemptuous by design. A reliable corrective after a month inside the risk literature, and unlistenable without one.

1 hr/ep DAIR Institute · free

Skeptical

AI Safety Fundamentals

20 BlueDot Impact · current · Course

A facilitated course covering technical alignment, governance and current research, in small cohorts with weekly discussion. Free, applications open periodically. Descended from Richard Ngo's original curriculum and the route by which much of the field entered it. Being argued with weekly is the only reliable defence against the drift this page keeps warning you about.

12 weeks bluedot.org · free

Risk case
04

Then keep a human in the loop

The last two are mine because this whole map is a living object, not a museum label. If I keep doing the job properly, the newsletter and Twitter feed are where the corrections, additions and embarrassed revisions show up first.

Cost: free · library card · rent or stream · paid.

Videos and PDFs can open in place. Everything else opens at the original source in a new tab.

This is a snapshot of a field that moves monthly. Keep current is the maintenance layer.