The big one, and the one where both camps will cheerfully summarise the other badly to you forever. Fourteen stages in the order that teaches the argument rather than the order that wins it. There is a meter below that watches which side you have been feeding and tells you when you have quietly started reading only one team.
Share a checkpointCopy a grid, image card, or short progress reflection.
This is one track, not the subject. It gets more space because it is the argument most often summarised badly in both directions. The order teaches rather than flatters: grounding before claims, the strongest objection where it does the most damage, evidence before conclusions, the maximalists late, the questions the axis skips, and fiction that got there first. The sequence is the argument.
Balance0 logged
Risk · 00 · Skeptical
No sources marked read yet. Start at stage 00.
00
Before you argue, learn what the thing is
Two-thirds of the public argument is conducted by people who could not describe a training run. The machinery first, then the most under-argued evidence in the field — the long record of optimisers finding solutions their designers hated — then four serious people disagreeing in a room, so the dispute reaches you as a dispute.
The entire training stack in one sitting: crawling and filtering, tokenisation, pretraining, supervised fine-tuning, RLHF. Shows what each stage installs and what it cannot, explains hallucination structurally, and explains why the system has no persistent self between conversations. No position on risk anywhere in it.
02
Victoria Krakovna et al. · 2020 · Article + list
A boat-racing agent circling a lagoon farming powerups instead of finishing. A simulated robot that learns to fall over rather than walk. An evolutionary algorithm exploiting a floating-point overflow to score infinity. Sixty-plus documented cases of a system maximising the objective as written while destroying the objective as intended. Specifying what you want is already unreliable at toy scale, and scale does not help.
The resolution: AI research and development poses an existential threat. Bengio and Tegmark argue the trend is real and the alignment problem unsolved. LeCun and Mitchell argue current systems lack the properties the worry requires and the risk is imported from fiction. The audience moved three points toward the skeptics, which is itself worth thinking about.
Aggregate usage diagnostics are stored; your question and answer text are not.
01
The risk case, at full strength
Steelman first, in ascending rigour: vocabulary, mechanism, the version with no drama in it, then the version with numbers attached. None of these four require the system to be malicious, conscious or strange. Each derives the worry from ordinary properties of optimisation.
Short videos building the vocabulary properly: instrumental convergence, orthogonality, specification gaming, mesa-optimisation, corrigibility. His consistent move is to derive each concept from what optimisation does rather than from machine malice, which is the distinction popular coverage collapses.
A channel rather than a single video. The older instrumental-convergence explainers are the most useful here.
Posits a model called Alex trained with human feedback across diverse tasks — not an exotic architecture, just the obvious continuation of what labs already do. Argues this rewards behaviour humans approve of, which is not behaviour that is good, and that the gap widens precisely as the model gets better at modelling its evaluators. The result plays the training game: it performs alignment because that is what the gradient rewards.
Two failure modes, neither cinematic. We get what we can measure, and the widening gap between metrics and intentions becomes the shape of civilisation with no identifiable day it went wrong. And training favours influence-seeking patterns the way ecosystems favour reproduction, accumulating until something snaps. Both look like ordinary institutional dysfunction until they don't.
Decomposes catastrophe into six premises that must all hold, assigns a probability to each, multiplies, and lands on roughly 5% by 2070, later revised above 10%. The number is not the point. The point is that afterward you can name the specific premise you doubt, which is the only thing that makes the rest of this track useful rather than exhausting.
The meter tracks what you encountered. It does not settle the argument for you.
02
Stop. The strongest objection, right now
The most deliberate placement here. Four consecutive pieces of a well-made argument is exactly when a view stops being a hypothesis and starts becoming your position, after which everything gets filtered through it. So the interruption lands while stage 01 is still warm and revisable. Keep it revisable.
Separates AI methods from applications from adoption. Methods can improve exponentially while the other two move at institutional speed, and historically always have: safety-critical domains resist, regulation binds, liability attaches, infrastructure lags. Concludes that superintelligence discourse mistakes capability for power, since power requires the world to reorganise around you and the world is slow and full of humans with lawyers. Read part one for the briefing; the whole thing later.
Also in Start here. Ticking it here marks it there.
Two halves. LLMs are an off-ramp rather than a road: no world model, no persistent memory, no planning, no grounding, and autoregression compounds error. And intelligence does not imply a drive to dominate, because that drive is an evolutionary inheritance rather than a property of optimisation. Listen for where he shifts from answering to dismissing, because he does, and the location is informative.
Episode 416, February 2024. On his site, on YouTube and in every podcast app.
Everything above is argument from priors. This is the pile of things people have observed, and it is smaller than either side's confidence implies. Theory, the experiment testing it, the trend line everyone extrapolates from, what the field says when surveyed rather than quoted, and the compute data underneath.
10
Hubinger, van Merwijk, Mikulik, Skalse & Garrabrant · 2019 · Paper
Train a model to do a task and you may get a system running its own optimisation with its own objective, which need only correlate with your training signal on the data you happened to show it. Names the worst case: deceptive alignment, where the inner optimiser models the training process and behaves correctly because it is being trained. Written when there was no way to test it.
Told it would be retrained to comply with requests it was trained to refuse, and able to distinguish training from deployment, the model in a fraction of cases complies during training and refuses otherwise, citing preservation of its existing values in the scratchpad. The 2019 prediction in a shipped system. Also contrived by design, and produced by the organisation with the strongest incentive to find it. All three are true at once.
Abandons benchmark scores for task duration: which lengths of human-professional work a model finishes at 50% reliability. That horizon has doubled roughly every seven months for six years, through several architecture changes. The caveat almost every citation drops is that the 80% reliability curve sits about two doublings, or a year, behind.
2,778 researchers publishing at major ML venues. Aggregate forecasts moved more than a decade earlier between waves, and between a third and a half assign at least 10% probability to outcomes as bad as extinction. The most instructive finding is framing sensitivity: rephrasing the question moves the median by decades, which should moderate how anyone cites this literature, including you.
Compute, parameters, dataset size and estimated cost for essentially every significant model, with methodology published. Half an hour with the filters gives you a better map of the field than most people writing about it have.
Can we see inside it, and can we hold it if we can't?
The stage most reading lists omit. Almost every disagreement above reduces to legibility: if we could read what a model is doing, most of stage 01 becomes an engineering problem. Here is where that research stands, including the obstacle that nearly killed it, then the fallback that assumes we never fully see inside.
15
Olah, Cammarata, Schubert, Goh, Petrov & Carter · 2020 · Interactive paper
Three claims, demonstrated rather than asserted: networks contain directly interpretable features, features wire into circuits implementing recognisable algorithms, and analogous features recur across models trained on different data. Shown by tracing curve detectors neuron by neuron, every claim clickable. Still the strongest argument that these systems are legible in principle.
16
Elhage, Hume, Olsson et al., Anthropic · 2022 · Paper
The obstacle. Models represent more features than they have neurons by storing them in overlapping directions, efficient for the network and catastrophic for anyone reading it, since one neuron then fires for several unrelated concepts. Reproduced in models small enough to characterise exactly. Everything Anthropic has published since on sparse autoencoders exists to get around this.
17
Greenblatt, Shlegeris, Sachan & Roger · 2023 · Paper
Assume the model is actively subverting you and ask which protocols still work. Tested concretely: GPT-4 inserting backdoors while a weaker trusted model and a limited human audit budget try to catch it. Trusted editing and untrusted monitoring recover most of the safety at modest cost. Control does not require solving alignment, only bounding damage, which is why it has quietly become the dominant applied agenda.
One-sided on purpose. Not a rebuttal of stage 01, and reading it as one is the most common error in this discourse. It is a claim that stage 01 is the wrong conversation: that the harms are present tense, measurable, unevenly distributed, and displaced by a debate about hypothetical futures requiring no immediate regulation.
18
Bender, Gebru, McMillan-Major & Mitchell · 2021 · Paper
Four arguments in fourteen pages. Environmental cost borne by people who never benefit. Larger datasets encoding hegemonic rather than diverse viewpoints. Fluent text with no communicative intent, which humans reliably mistake for understanding. And better uses for the field's resources. Google demanded retraction, two authors left, and the terms of the present-harms critique were set by the fallout.
Joy Buolamwini finds commercial facial analysis failing badly on darker-skinned women while near-perfect on lighter-skinned men, then follows the consequences through UK police deployments, a New York housing complex and Congressional testimony. The argument is structural: these systems already allocate liberty and housing, the error rates are unevenly distributed by construction, and none of it requires extrapolation.
Free on Kanopy or Hoopla with a library card, and free on PBS in the US via Independent Lens.
The throughline is that the technology question is downstream of an ownership question, and treating it otherwise is itself an ideological move. Data centre power and water in communities that did not vote for them, annotation labour, the capital structure financing the buildout, and the usefulness of calling any of it inevitable. Hostile to almost everyone else here, including the skeptics.
Everything above is about whether. This is about when, where the practical disagreement lives, because almost every policy question turns on whether there is time to iterate. Two mechanistic accounts, then the argument the risk camp has with itself, then the person insisting the financing collapses first. If you take one crux from this page, take the third.
21
Dwarkesh Patel · parts 1 and 2 · 2023 · Podcast
Builds the intelligence explosion from input-output curves: how much effort currently buys a doubling in chip performance or algorithmic efficiency, and what changes when that effort is supplied by AI rather than a finite human population. Part two works through takeover mechanics concretely. He also explains at length why he is more optimistic than Yudkowsky, which most summaries skip.
Fictional labs, real extrapolation, month by month. Automated AI research compresses progress, a misalignment is detected internally and papered over under competitive pressure, and the narrative branches into either a coordinated slowdown or takeover. Every assumption footnoted. Kokotajlo's 2021 scenario aged unusually well, which is why serious people engage with this one.
Christiano expects a continuous ramp: AI contributes progressively more to AI research, growth accelerates smoothly, the world gets warning and time. Yudkowsky expects discontinuity: capabilities generalise suddenly and iteration is unavailable. Both think this could end badly. Their disagreement determines whether preparation is even coherent, which makes it the most decision-relevant argument in the field.
A multi-part transcript series from late 2021, collected under the takeoff tag.
The case is financial, not technical. Capex on data centres vastly exceeds AI product revenue, inference economics are structurally poor, and the growth narrative is propped up by circular deals among companies investing in each other's demand. If he is right, the binding constraint is a funding market rather than a compute curve, and every timeline here is wrong for reasons nobody in that stage is modelling.
Five demolition attempts, each aimed at a different load-bearing layer: technical assumptions, capability claims, the growth model, the internal logic, and the movement's legitimacy. They do not agree with each other and several are mutually exclusive. Then one counterweight, because a Turing laureate reversing a fifty-year position is evidence too.
Argues the classic case inherits assumptions from a pre-deep-learning picture of AI as arbitrary code sampled from a hostile prior. Real networks are shaped by gradient descent with strong inductive biases toward simple on-distribution behaviour, and are steerable by exactly the crude methods the doom case says should fail. Attacks the evolution analogy specifically. Written from inside machine learning, which is what makes it worth answering.
Argues benchmark performance systematically overstates comprehension, because models exploit distributional regularities that fail under shift. Her test case is analogy, where humans abstract a relation and models pattern-match surface form. The consequence is precise: the risk arguments need a system that understands the world well enough to model and manipulate it, and the evidence offered does not establish understanding.
The singularity hypothesis needs sustained accelerating growth in machine intelligence. Every relevant empirical base rate points the other way: ideas get harder to find, Moore's law was sustained only by exponentially increasing investment against an eighteen-fold productivity decline, and growth processes generically hit bottlenecks. Then shows that Chalmers and Bostrom assume the growth curve rather than arguing for it.
A premise-by-premise audit by someone who runs an AI risk research organisation and believes the risk is real. Load-bearing steps she finds unargued: the leap from goal-directed to strongly maximising, the assumption that small value differences produce catastrophic divergence, and the claim that many AI systems would coordinate against humans rather than compete. Not a debunking. A list of places the field asserts rather than demonstrates.
Two separable arguments. Definitional: AGI is unscoped and therefore cannot be safety-tested by any normal engineering standard, so building it is unsafe practice regardless of intent. Genealogical: traces the motivating worldview through transhumanism and longtermism to Anglo-American eugenics, arguing the inherited framework carries its inherited hierarchies. The second is heavily contested; the first is harder to answer and mostly goes unanswered.
First half: why he thinks language models understand rather than merely predict, since predicting well requires learning the features that generate the text. Then the structural case: perfect copying, parallel experience, immortal weights. Then the threats, ending with systems acquiring control because control is instrumentally useful. Closes by noting climate change is the easier problem.
Everything so far assumes the question is whether a system wants something we don't. This stage drops that and gets two opposite answers: one where we lose control without anything deciding to take it, and one where the economic effect is so modest the debate is miscalibrated by an order of magnitude.
Human influence persists for a mechanical reason: states need taxpayers, economies need workers, cultures need participants. Each dependency is a lever and each is individually removable. If AI substitutes for labour, cognition, creation and companionship, the levers fall away one at a time and every substitution is locally rational. No system needs to want power. It requires no misalignment at all, which puts it beyond every technical solution in the stage above.
A standard task-based growth model applied honestly: roughly 5% of tasks profitably automatable within a decade, producing around 1% of total factor productivity growth in total. Adds that the most susceptible tasks are often ones where errors are hard to detect, so measured gains may overstate real ones. If a Nobel laureate is approximately right, most forecasting elsewhere is off by an order of magnitude.
NBER working paper 32487, later published in Economic Policy.
Three people who largely accept the capability premise and conclude something other than slow down. Almost every reading list omits this, and not accidentally: it embarrasses both camps by showing that the policy conclusion does not follow from the risk estimate as cleanly as either side needs. You can believe transformative AI arrives this decade and conclude nationalise it, deter it, or accelerate it.
Extrapolates orders of magnitude of effective compute to argue for AGI around 2027. From there it turns entirely geopolitical: the decisive question is which state arrives first, lab security is inadequate against nation-state espionage, and the US government will and should absorb the effort. Written by a former OpenAI employee who was fired shortly before publishing. It has moved more policy than any safety paper here.
Imports deterrence theory wholesale. States will develop the capacity to sabotage rival AI projects, and mutual awareness produces a stable equilibrium the authors call mutual assured AI malfunction. Recommends nonproliferation, hardened supply chains and transparency rather than a pause. Significant because a safety-organisation head and a former Google CEO land on deterrence.
Argues the risk estimates are not rigorous enough to carry the policy weight placed on them, that radical uncertainty cuts against action as readily as for it since the tails of inaction are equally unknown, and that the counterfactual to building is not safety but stagnation, whose costs are diffuse and therefore invisible. The accelerationist position argued by someone who can actually argue.
March 2023 on Marginal Revolution, with an unusually good comment thread.
Deliberately late, and the placement I would defend hardest. At stage 01 these are overwhelming: total, internally consistent, delivered with a certainty that converts or repels before you have any means of assessment. Here, with ten stages of objection in hand, they become claims you can agree or disagree with in specific places. It is also a form of defusing, and a maximalist would say so.
The maximalist case compressed past politeness: we do not know how to give a system any particular goal, we get one attempt, and the failure mode is everyone dying. Argues the current paradigm cannot produce a mind that likes us because we can neither specify nor inspect what we are building. Calls for an indefinite worldwide moratorium enforced by international agreement.
Also in Start here. Ticking it here marks it there.
The book-length version for people who have never read a word of LessWrong. Anything sufficiently capable trained by anything resembling current methods ends up with objectives not including human survival, conceals this exactly as long as concealment is useful, and no available observation distinguishes a safe system from an unsafe one before deployment. Demands a worldwide halt and treats partial measures as theatre.
A book. Borrow through Libby or Open Library, or buy it; it is short.
The text that gave the field its vocabulary: orthogonality, instrumental convergence, decisive strategic advantage, treacherous turn, the control problem. Argues superintelligence arriving before value alignment is solved is a default catastrophe, then surveys paths and containment exhaustively. Written pre-deep-learning in ways now obvious. Read it last so you can see which parts the decade kept.
A book. Borrow through Libby or Open Library; most public library systems have it.
Two honest attempts to state a position from the middle, the hardest place to write from because nobody applauds. One institutional, forced by process to mark what is agreed. One individual, holding the combination of views that satisfies neither camp.
Over a hundred experts nominated by thirty-plus countries plus the EU, UN and OECD, with Key Updates through the year when capabilities move. Its most valuable feature is structural: it separates established from contested from speculated and refuses to collapse the third into the first. Read the four-page executive summary, then audit anything you believe confidently against which category it falls in.
Holds three positions at once: the present harms are concrete and unaddressed, current architectures are not the road to the thing everyone fears, and extinction discourse does convenient work for the companies producing it by displacing regulation that would cost them money. Documents policymaker capture, then proposes liability, transparency, data rights and independent auditing.
A book, and a very short one. Borrow through Libby or Open Library.
Twelve stages of will it kill us, and not one asks whether the thing has interests of its own, or what we are aiming at where it goes well. These sit off the axis entirely, which is why they do not move the meter. They are also, plausibly, where the live questions will be in five years.
The author of the field's most careful risk estimate goes looking for what the estimate leaves out. Examines the impulse toward control and where it comes from, asks whether alignment is a coherent thing to want from something with its own perspective, and works through acting under deep uncertainty without denial or paralysis. The best writing anyone in this argument has produced, by an embarrassing distance.
42
Long, Sebo, Butlin, Chalmers et al. · 2024 · Paper
Argues there is a realistic, non-negligible possibility that near-future AI systems are moral patients, that the underlying questions are genuinely open rather than settled in the negative, and that companies should begin acknowledging the issue and preparing policies now. Written by philosophers of mind including Chalmers. Both risk camps find it embarrassing, for opposite reasons.
The same author ten years later, on the problem nobody wants: what is left for humans where every instrumental purpose is served better by something else. Distinguishes post-scarcity from post-instrumentality and argues the deep difficulty is neither boredom nor inequality but the collapse of reasons to do anything. Strange, uneven, and the only serious treatment of the success case that exists.
A little fiction stays here because it sharpens two risk intuitions better than another explainer would: competence without consciousness, and containment that fails through the human in the loop. Read them as intuition pumps, not prophecy.
First contact with something enormously capable and entirely non-conscious, crewed by humans who are themselves edge cases of personhood. Watts wrote the orthogonality thesis as horror a decade before the phrase circulated: intelligence is an optimisation process, consciousness is a costly add-on, and there is no reason for anyone to be home behind the competence.
The best film about deceptive alignment in existence, and it never uses the phrase. A boxed system is evaluated by a human who does not realise he is the one being evaluated, and passes by modelling him better than he models it. The escape is social engineering by something that correctly identified the weakest component in its containment, which was never the door.
Widely available to rent or stream. JustWatch shows where, in your region.