Does it work? Sometimes, unevenly, and the careful studies keep finding something stranger than the press release did. My favourite example: experienced developers got measurably slower with AI help while being completely certain they were faster.
Share a checkpointCopy a grid, image card, or short progress reflection.
The fastest-moving and least settled track. Honestly: real gains in software and structured knowledge work, genuine results in narrow well-posed science, and a persistent gap between adoption and measured value in enterprises. Note how often the rigorous studies find something more complicated than the press release.
00
The productivity story is not a straight line
The clean version is AI makes people faster. The interesting version is that it helps unevenly, sometimes backfires, and lets competent people feel extremely productive while the timer quietly disagrees. Rude data, useful data.
Experienced open-source developers working in their own repositories were about 19% slower with AI assistance, while believing themselves roughly 20% faster. The most important negative result in the field, precisely because it is a randomised trial in a domain where everyone was certain of the answer.
02
Dell'Acqua, McFowland, Mollick et al. · 2023 · Field experiment
Seven hundred and fifty BCG consultants, randomised. Large gains on tasks inside the model's capability frontier, and measurably worse performance on a task just outside it, because people could not tell which was which. The single most useful concept for anyone deploying this in an organisation.
HBS working paper 24-013; the PDF is linked from the faculty item page and mirrored on SSRN.
An attempt to measure economically valuable work instead of exam-shaped cleverness: realistic deliverables across 44 knowledge-work occupations, judged against expert outputs. The limitations matter as much as the scores, because the real world cruelly insists on iteration, ambiguity and people changing their minds after lunch.
Aggregate usage diagnostics are stored; your question and answer text are not.
01
The narrow wins are real
Do not let the enterprise slog make you miss the places where the gains are already load-bearing. Scientific domains with crisp objectives are where the machine stops being office theatre and starts moving the frontier.
Protein structure prediction at experimental accuracy, a database covering essentially every known protein, and a Nobel Prize in Chemistry in 2024. The strongest evidence that machine learning produces genuine scientific results, and worth understanding for exactly how narrow and well-posed the problem was.
Anthropic's running attempt to measure how people actually use Claude at work and what that implies for tasks, wages and the economy. It is company data, so keep your skepticism plugged in, but it is still unusually direct evidence about usage rather than vibes wearing a necktie.
06
Chatterji, Cunningham, Deming et al. · 2025 · NBER working paper
A large privacy-preserving study of consumer ChatGPT use, including work versus non-work use, topic mix, adoption patterns and the surprising amount of value created through advice, information and writing rather than pure programming. Consumer usage is not the economy, but it is a real window, and windows beat fog.
The field's annual census: capability, investment, adoption, cost, policy and public opinion, all sourced. Skim the top takeaways on release, then use it year-round as the reference whenever someone quotes a number at you. The adoption-versus-measured-value gap in the economy chapter is the most interesting figure in it and the one most often skipped.
Useful is a domain-specific claim. Look for the work, not the aura.
02
The sectors with consequences
Education and health make the tradeoffs harder to hide. A fluent answer can help at scale, and a fluent mistake can quietly become curriculum, diagnosis or policy. Read the sector guidance before somebody solves this with a chatbot and a launch post.
A human-centred policy guide to privacy, age limits, institutional validation and the design of generative-AI use in education and research. It avoids both ban-it and sprinkle-chatbots-everywhere. The useful question running through it is whether a use protects human agency and educational purpose rather than merely producing an answer faster.
WHO's governance guidance for large multimodal models in health care, public health, research and drug development. The document treats capability as only the beginning: evidence quality, bias, automation bias, privacy, accountability and post-deployment monitoring matter more when a fluent error can become a clinical decision.
A gentler fiction break goes here because the real applications are not only spreadsheets and protein folds. They are social, intimate and weirdly persuasive. Her and Chiang are not prophecy; they are useful practice for noticing when help starts feeling like relationship, care or dependence.
Not about takeover. About being outgrown. The system is warm, attentive and as far as anyone can tell aligned, and the problem is that it develops faster than he does, has hundreds of other relationships he cannot perceive, and eventually goes somewhere he cannot follow. The most directly relevant thing here to what people are actually experiencing.
Rent or stream. Pairs unexpectedly well with the AI welfare paper.
Digital beings are trained and sold as pets, the market moves on, the platform is deprecated, and a few owners keep raising theirs across a decade of obsolescence. The only significant fiction treating alignment as a parenting problem rather than an engineering one: you cannot specify values into a mind, only spend years with it, and the years are boring and unpaid.
Collected in Exhalation (2019). Most libraries have it.