The facts in this field have a shelf life measured in months. These are the things to check on a schedule, so the model of the world you just spent a month building keeps getting corrected by reality.
Share a checkpointCopy a grid, image card, or short progress reflection.
Last on purpose, and the only track with no end. Everything above is a snapshot, and snapshots age quickly here. These are worth checking on a cadence rather than reading once, because the calibration you just built starts decaying when you stop feeding it. If you maintain only three, make them Epoch, METR and the annual safety report. Between them you get the compute and cost picture, the capability picture, and the contested-claims picture, from parties who publish their methods.
00
The monthly calibration loop
These are the sources to keep in rotation so your model of the world does not calcify into a handsome little fossil. Capability, compute, safety posture, governance promises: check them before your confidence gets decorative.
The field's annual census: capability, investment, adoption, cost, policy and public opinion, all sourced. Skim the top takeaways on release, then use it year-round as the reference whenever someone quotes a number at you. The adoption-versus-measured-value gap in the economy chapter is the most interesting figure in it and the one most often skipped.
Compute, parameters, dataset size and estimated cost for essentially every significant model, with methodology published. Half an hour with the filters gives you a better map of the field than most people writing about it have.
Over a hundred experts nominated by thirty-plus countries plus the EU, UN and OECD, with Key Updates through the year when capabilities move. Its most valuable feature is structural: it separates established from contested from speculated and refuses to collapse the third into the first. Read the four-page executive summary, then audit anything you believe confidently against which category it falls in.
04
Model Evaluation and Threat Research · per release · Check per launch
Independent pre-deployment evaluation of frontier models for autonomy and dangerous capability. When a major model ships, METR's write-up is usually the most informative independent account of what it can do unsupervised.
Where Anthropic, Google, Microsoft, OpenAI and others publish shared technical work on safety frameworks, thresholds and evaluation. Read it as the industry's own account of what it has committed to, useful precisely because it is checkable against behaviour later.
Anthropic's current policy for escalating safeguards as model capabilities rise. Version 3.0 separates capability thresholds, required safeguards and public reporting into a company-specific system that can be checked against later releases. Read the actual policy rather than treating three labs' differently structured promises as one document.
OpenAI's framework for tracking severe-harm capabilities and requiring safeguards before deployment. Version 2 focuses its top-level categories on biological and chemical capability, cybersecurity and AI self-improvement, with risk reports and a Safety Advisory Group built into the process. Compare the categories and governance mechanics directly with the other labs rather than assuming they match.
08
Google DeepMind · 2026 · Frontier safety framework
Google DeepMind's current framework for identifying critical capability levels and pairing them with security and deployment mitigations. Version 3.1 adds and revises protocols as the threat model changes. Its thresholds, review process and terminology are its own, which is exactly why it deserves a separate record.
Aggregate usage diagnostics are stored; your question and answer text are not.
01
The quick reality checks
Leaderboards and model cards are not truth serum, but they are excellent smoke alarms. They tell you when the frontier moved, when prices changed, and when last quarter's argument has started wearing an antique hat.
Blind pairwise human preference at scale, rendered as Elo. Known weaknesses: it rewards style and formatting, it is gameable, and preference is not capability. One instrument, not the only one.
Quality against price against latency across every major model and provider, updated continuously. The page that turns abstract capability talk into a procurement decision, and where the cost collapse at fixed capability becomes impossible to miss.
The official running log for ChatGPT model changes, retirements, routing and behavior updates. It is not a benchmark, but it tells you what users are actually being moved onto, which is often the part missing from lab-chart arguments.
Anthropic's collected model and system cards for Claude releases. Read them when a model launches, mostly for the deltas: capability claims, safety evaluations, refusals, tool use and where the company chose to draw the box around the product.
13
Theo Browne · ongoing · Video and developer commentary
A fast, opinionated read on what AI coding tools are doing to actual web development: coding agents, app builders, product decisions and the gap between an impressive demo and a thing someone can ship. Theo has strong preferences and says them at full volume, which is useful data when the tooling changes faster than its documentation.
A weekly pass through AI research, policy and industrial weirdness by someone unusually good at making the moving parts legible. It is not neutral in the fake view-from-nowhere sense; it is useful in the someone-read-the-papers-and-has-taste sense. That distinction matters.
The rare feed that catches model releases, local inference, coding agents and prompt-injection failures while the rest of the internet is still composing a launch thread. Practical, unusually well documented and excellent at separating what a tool claims from what happened when somebody actually ran it.
16
swyx and rotating co-hosts · ongoing · Podcast and newsletter
AI engineers interviewing the people building model, agent and infrastructure systems. It is especially good for learning the vocabulary of a new technical pattern shortly before that vocabulary gets flattened into brochure copy.
Long interviews with lab leaders, researchers and unusually consequential operators, with enough room for claims to acquire assumptions and caveats. Not a news feed. It is where you go for context after one claim has eaten the entire week.
Podcasts and courses are scaffolding, not scripture. These three are useful when you need a wider-angle news pass, a skeptical temperature check, or a structured way into the safety argument before it hardens into team merchandise.
18
Kevin Roose & Casey Newton, NYT · 2022– · Podcast, weekly
Two Times journalists on the week's releases, lawsuits, policy and drama. Sits deliberately in the middle and treats both camps as subjects rather than allies, which annoys everyone. Its real function here is as a gauge of which arguments are reaching people who have read nothing else on this page.
Free in any podcast app; no Times subscription needed for the audio.
Each episode takes one claim, paper or press release and goes through it line by line. The standing position: text synthesis machines, marketing capability claims, benchmarks measuring the wrong thing, risk discourse functioning as advertising. Contemptuous by design. A reliable corrective after a month inside the risk literature, and unlistenable without one.
A facilitated course covering technical alignment, governance and current research, in small cohorts with weekly discussion. Free, applications open periodically. Descended from Richard Ngo's original curriculum and the route by which much of the field entered it. Being argued with weekly is the only reliable defence against the drift this page keeps warning you about.
The last two are mine because this whole map is a living object, not a museum label. If I keep doing the job properly, the newsletter and Twitter feed are where the corrections, additions and embarrassed revisions show up first.
The most direct way to keep this orientation alive after the reading map: get future notes, updates and revisions by email. This belongs at the end because the useful version of this project is maintained, not finished.
Newsletter signup page using the site's Mailchimp form.
The lighter-weight currentness layer: short notes, links and half-formed updates, usually before they become posts or revisions. Useful if you want the reading map connected to what I am noticing in public.