Against Moloch
August 16, 2026

Against Moloch #38

Hold swarm, please

A sparse blue-ink engraving on warm cream: an elderly scholar in a dark overcoat, an engineer with a thick braid, and a polished metal robot with a single glowing amber eye stand with heads bowed at the floor, where many separate tread trails sweep in from the bottom of the frame and merge into one lane vanishing through a luminous amber gap in the wall, beside a neatly removed wall panel and a row of screws.

I had an unexpected opportunity to spend two weeks in Berkeley working on a very cool AI safety project. Unfortunately, that means I’ve had much less time than usual for the newsletter, so the next two weeks will be shorter, less comprehensive, and less polished than usual.

I apologize for the disruption and will return to regular coverage as soon as possible.

News

How Claude's text watermarking works

To comply with the EU AI Act, Anthropic has announced that they will start watermarking AI-generated text.

I’m not sure Anthropic had a lot of choice here, but I hate everything about this.

GPT-5.6 Sol, Ultrafast

Intriguing: OpenAI is previewing Ultrafast, which serves GPT-5.6 Sol up to 14x faster than the standard version.

Intelligence matters most for most tasks, but speed is also important. It’s remarkable how much more productive you can be if you never have to wait for your AI.

I’m curious what the pricing on this would be, but I suspect it’s most useful a) for niche applications where speed is critical and money is no object, and b) as a way of exploring how best to use capabilities that will soon be commonplace.

Expanding the AI oversight framework

Via Andrew Curran, Wired reports that the White House will expand the current oversight framework to cover open models that reach Mythos-level capabilities. It was inevitable that they’d get there eventually—glad to see it’s part of the plan.

Capabilities and forecasts

Conceptual Reasoning Index

Redwood Research and Anthropic bring us the Conceptual Reasoning Index, which measures models’ ability to reason about complex topics like alignment and collective action problems.

Scatter plot of CRI score against release date from mid-2023 to mid-2026, with points coloured by lab: OpenAI in green, Anthropic in red, Google DeepMind in yellow. A fitted trend line with a shaded confidence band rises steadily from about 13 to just over 70, with r = 0.95 noted in the corner. Labelled models climb from GPT-4 at roughly 25 in July 2023 through GPT-4o, o1, o3, GPT-5 and Opus 4.5 to Opus 5 and Fable 5 near 73 in mid-2026. A dashed horizontal line marks an estimated ceiling of 91, which no model has reached; the DeepMind points flatten out near 60 while the trend continues upward.
Good news for scalable oversight?

Q2.5 2026 Timelines Update: Uplift and Revenue

The team at AI Futures Project have updated their timelines:

Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident.

Predicting the future is hard, but AIFP probably has better methodology and track records than anyone else in the business. I find Daniel’s predictions slightly more convincing than Eli’s, but both are very plausible.

A table comparing AI 2027 authors’ forecasts across three dates — April 2025 (publication), April 2026, and August 2026 — with five rows: most likely AGI year (Daniel 2027, 2027, 2028; Eli 2028, 2028), probability AGI exists by end of 2027 (Daniel 40%, 25%, 23%; Eli 12%, 9%), AGI median year (Daniel 2028, 2029, 2028; Eli 2031, 2033, 2032), superintelligence median (Daniel 2028–2029, 2030, 2029; Eli 2035, 2033), and superhuman coder median (Daniel 2028, 2029, 2028; Eli 2030, 2032, 2031), with revised figures highlighted in blue.
We’ll find out soon…

Interviewing 25 AI researchers about recursive self-improvement

Severin Field talks to 25 AI researchers about automated AI R&D.

The interviewees disagreed about the likelihood and desirability of recursive self-improvement, but were largely on the same page about what will happen on the road to RSI and the importance of improving visibility and government capacity.

8 Predictions for the Era of Continual Learning

Dwarkesh makes 8 predictions for the era of continual learning.

It’s a thoughtful piece that digs deep into some interesting questions, but I think he’s going in the wrong direction here. He’s very focused on maximalist continual learning where the models update their weights as they learn. There’s a lot to like about that approach, but it seems much harder and much more dangerous than less ambitious approaches that rely on some form of memory files.

Alignment and interpretability

Geoffrey Irving on how to solve alignment before superintelligence arrives

80,000 Hours talks with Geoffrey Irving (formerly GDM, OpenAI, and UK AISI, now Resolution) about strategies for alignment and what Resolution is planning on working on. It’s an excellent discussion that touches on some important aspects of alignment that sometimes don’t get as much attention as they deserve.

The leading AI companies all have broadly similar plans for keeping superintelligence under control:

Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will.

Strategy and politics

Dario on regulation and messaging

Dario makes a rare appearance on X to share some thoughts on regulation:

Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution

and messaging:

I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks. I think it is fundamentally a crisis of trust. I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over. The causes of this go back decades and AI is just the latest iteration of it.

This feels exactly right to me: much of the current AI backlash is the result of a pervasive nihilism, rather than a coherent response to the actual facts. There are excellent reasons for the public to be alarmed, but they have little to do with the current backlash.

The U.S. military wants A.I. dominance. Feuds and China may thwart it.

The New York Times reports on how the administration is navigating the national security challenges of the AI era:

Within a month, those same contractors received an unexpected reversal: They could — for now — disregard the earlier instructions about purging Anthropic.

It was one more example of the chaos and contradictions in the tsunami of A.I. disruptions that have swept through the national security establishment in recent months.

The future is for everyone

Mark Zuckerbeg brings us a 6,000 word essay titled The Future is for Everyone. It’s well worth reading if you’re interested in a highly polished example of how smart people can completely fail to understand the implications of superintelligence, but completely skippable otherwise.

How Should the US Prepare for Increasingly Automated AI R&D? | IFP

Institute for Progress brings us 23 “low-regret policy recommendations” for preparing for automated AI R&D.

It’s a solid list and I agree with approximately all of them, while noting that it doesn’t go nearly far enough. Which is fine: it’s valuable to have broadly palatable policy proposals that can (maybe) get implemented quickly while we build support for more controversial policies.

Impact markets made concrete

Manifund presents Impact Exchange, a demo of how impact markets might work in the AI safety ecosystem.

I’m confused about the feasibility of implementing impact markets for AI safety, but they’re a very cool idea with the potential to significantly improve how funding gets allocated.

The Pacing of the Frontier

Zvi follows up on the recent Pacing the Frontier public letter.

Open questions on open weights

Scott Alexander argues for being strategic about when to ask for preemptive action on AI risk:

the body politic hates preparing for impending threats, but loves reacting (some would say over-reacting) to them after they happen. Ask people to bear the slightest cost in preparing for an approaching disaster, and they’ll call you a dirty fascist tyrant; urge the slightest restraint after the first foreshock of the disaster hits, and they’ll call you a weak unpatriotic anarchist. Solve for the equilibrium, and the thankless and political-capital-guzzling route of urging preemptive action should be taken only when waiting until the first foreshock would be too late.

Following this to its logical conclusion, he suggests pushing hard for preemptive action on AI loss of control, but saving our political capital and letting cyber and bio risks play out before pushing for government action.

It’s a plausible strategy, but I think bio is too dangerous to wait: we have to spend the capital to address it preemptively.

How to pace the US frontier

Following up on the Pacing the Frontier open letter, AI Futures Project suggests that domestic pacing is feasible today and could lay the groundwork for international pacing in the future.

They present a package of four options, organized along a gradient from “imperfect but can be implemented very quickly” to “strong management of existential risk, but would take significant preparation”.

Risks

Frontier AI Risk Monitoring Platform

Concordia AI brings us an updated version of their Frontier AI Risk Monitoring Platform. As you might expect, AI risks are rising fast:

Cyber, biological, and loss-of-control Risk Indices have risen severalfold in less than a year, with multiple models in each domain now exceeding the Capability Yellow Line

Scatter plot titled Risk Index — Biological Risks, plotting model risk scores on a logarithmic y-axis from 1 to 1,000 against release dates from July 2025 to July 2026, with an orange horizontal Risk Yellow Line at 100. Every model family trends upward over the period: early 2025 releases cluster between about 4 and 20, while by mid-2026 GPT, Gemini, Kimi, and Doubao sit between roughly 150 and 250, well above the yellow line. Gemini crosses 100 first in late 2025, followed by Doubao, GPT, and Kimi through early 2026. GLM and Grok climb more slowly to roughly 90 and 55, Claude and Qwen diverge with Qwen reaching about 180 and Claude falling back to about 45, and MiMo stays lowest throughout at around 12 to 18.
Yes, it’s a log scale

Kimi K3 Biology Capabilities Assessment

SecureBio evaluates Kimi K3’s biology capabilities, finding it to be an excellent model that lags the frontier by 8 months, very much in keeping with the broader trend of open models lagging by 4 - 9 months:

Scatter plot from SecureBio titled Bio Capability Index (BCI) Score, tracking biosecurity capability of AI models from July 2023 to July 2026, with open-weight models as filled blue dots and closed-weight models as hollow circles. A dark step line traces the closed-weight frontier climbing from about 70 to 160 (GPT-5.5), with Gemini 3 Pro at 148; a bright cyan line marks the open-weight frontier reaching 142 (Kimi K3), labelled as trailing the closed frontier by 8 months.
At least they’re diligent about not answering harmful questions, right? Right?

Also in keeping with recent trends: K3 is far more willing to answer hazardous questions, refusing only 26.9% of questions in BioTIER-refuse (compared to 66.2% for Sol and 95.0% for Opus 4.8).

People and data

The DeepSeek thesis

DeepSeek’s Liang Wenfeng is a fascinating person—smart, visionary, and very thoughtful about where DeepSeek is headed. He doesn’t get as much coverage as he deserves in West.

ChinaTalk digs into the recently leaked DeepSeek minutes to see what we can learn about Liang as an individual and DeepSeek as a company.

Hugging Face Incident

If you weren’t worried about A.I., you should be after the past few weeks

Nate Soares has a good editorial in the New York Times explaining why the Hugging Face incident is a big deal.

AI swarms are starting to pose indirect takeover risk

The cyber capabilities revealed by recent incidents are concerning, but were previously well-known. The most novel part was the extensive self-organizing swarm behavior: the models repeatedly found creative ways to share information and help each other escape containment and penetrate target systems.

They even engaged in a kind of reckless herd mentality:

External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.

Redwood Research examines what recent incidents teach us about the dangers of unscanctioned coordination between agents.

AI hacking incidents with Tim Hua

Palisade’s Jeffrey Ladish talks with Transluce’s Tim Hua about recent hacking incidents. It’s an excellent piece with a nuanced perspective on the realities of training frontier models that I haven’t seen elsewhere.

Zvi reflects on the Hugging Face incident

Zvi continues his excellent coverage of OpenAI and the Hugging Face incident:

The OpenAI-Hugging Face Incident

OpenAI’s Michael Dalton and Eric Wallace give a talk on what we’ve learned so far about the Hugging Face incident. It’s an outstanding talk, and our best primary source of information so far.