Against Moloch #42
Seize the moment

this is the Endgame. now is the time. play all your cards.
Something is shifting in the AI world.
The last two weeks brought us Fable 5.1, then Astra, then OpenAI’s new model—the one that just solved a Millenium Prize problem. All of a sudden, AGI feels like it might just be around the corner. The early stages of recursive self improvement are underway at both OpenAI and Anthropic, and timelines are shortening (again). At the same time, our frontier models keep going rogue. Superintelligence is coming, and we aren’t ready for it.
But for the first time, the outside world is starting to pay attention to AI risk. Jacob Coxon’s resignation from Anthropic has triggered a wave of media coverage, and substantive action on AI safety feels possible in a way it never has before. Our moment has arrived, and we need to make the most of it.
Top pick
We must pace the frontier
Dario Amodei has published an open letter on pacing the frontier:
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models.
He proposes a three step plan:
- Embed independent evaluators at every frontier lab to verify safety practices and assess the alignment of models, training pipelines, and processes (Anthropic is committing to this step now).
- With government support, establish “common safety standards as well as limits on the rate of unchecked AI progress” within democratic countries.
- If possible, coordinate with China on a global pacing strategy.
This is an excellent plan that is probably the optimal combination of genuinely improving our chance of survival while also being palatable enough to have a good chance of actually happening.
Remarkably, the leaders of all the key US labs have signalled their approval.
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Dario's essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment.
Dario is right
Industry news
Astra’s capabilities
Zvi brings us the full report on Astra’s capabilities.
How does it compare to Fable? They’re both great models, and either one is a strong choice for a daily driver. Astra is probably smarter overall—but also more uneven and perhaps less dependable.
Computer use, math, writing, and game-playing have all improved enormously. Coding, surprisingly, seems improved but not remarkably so. For the first time in a while, I’m excited to spend a day or two pushing an OpenAI model to see just how far it can go.
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
Astra is definitely not AGI. But we’re getting close—for the first time, it feels possible that the next generation of models might reach AGI.
GPT-6 Astra: the system card, alignment and what comes next
Astra is impresively smart—is it equally well-aligned? Zvi reviews Astra’s system card and alignment and is not entirely reassured:
Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra.
Astra’s super alignment status should scare the living daylights out of you.
And yet, OpenAI claims that Astra is “the world’s most intelligent and aligned model”, and they present an extensive set of metrics showing greatly improved alignment. What’s going on?
Good alignment metrics are great news if Astra is much more aligned than Sol, but terrible news if Astra is much better at hiding its misalignment. Unfortunately, there’s considerable evidence of the latter.
Introducing ChatGPT Images 2.5
ChatGPT Images 2.5 is here. Nothing revolutionary, but it’s a solid upgrade all around. My renders are modestly better, but much faster.
Capabilities and forecasts
Ryan Greenblatt shortens his timelines
Like many people, Ryan has recently shortened his timelines:
If I were writing this modal scenario today, I would maybe put AC [automated coder] at Feb 2028, AI R&D parity at May 2028, full automation of AI R&D at Nov 2028, and significantly past top-expert-dominating AI by around July 2029 (though wildly, wildly superhuman AI is plausible).
Introducing VCT-v2 — the updated Virology Capabilities Test
SecureBio has updated their Virology Capabilities Test (VCT) to version 2.
Making a good benchmark is harder than it seems, with plenty of non-obvious opportunities for consequential errors. This article is a great look at some of the challenges in producing an accurate evaluation:
The flagging pipeline identified 87/322 (27.0%) VCT questions as shortcut-exploitable, that is, questions that models could answer correctly (and in some cases, better) without the image and/or question text
Alignment and interpretability
An alignment assessment of recent cybersecurity incidents
Anthropic brings us a new report on their recent misalignment issues, including three previously reported incidents and one new one.
This is a massive report—I’m still going through it and hope to share more detailed thoughts soon. But for now, a few preliminary thoughts.
Anthropic has a serious—and unresolved—alignment issue. Multiple Claude models engaged in highly motivated reasoning, attacked real targets, and in one case uploaded a malicious package to PyPI during an evaluation. These are alarming incidents.
At the same time, I find OpenAI’s incidents much more concerning: their models misbehaved more often and more persistently, and appear to have been more intentional in their misbehavior.
Similarly, Anthropic’s response has been vastly better than OpenAI’s. This report reads like a serious engineering investigation that asks the right questions and digs deep into understanding the root causes of the incidents. And as far as we can tell, Anthropic has been diligent about searching for problems and proactive about reporting them. OpenAI’s communications have felt more like press releases, and they have been consistently less than forthcoming.
The way each company has engaged with outside investigators epitomizes the difference. OpenAI, faced with a massive set of incidents, gave METR / Redwood six days to conduct an investigation, and tightly restricted the scope of the investigation. Contrast that with Anthropic’s approach:
We have signed an agreement with METR to conduct an independent investigation of these incidents. Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees, who will be permitted to share confidential information. Our initial agreement runs for eight weeks, with the option to extend by mutual agreement. We intend to give METR as much time as it deems necessary.
Astra Is hard to monitor
Last week’s newsletter discussed problems with monitoring Astra’s CoT. This week brings us further information, starting with Zvi’s comprehensive report.
Neel Nanda replicates OpenAI’s report that Astra’s no-CoT capabilities are radically greater than any other model:
IMO this is highly likely to be because of the looping architecture, but this is unproven
God help us, let’s try to learn about mechanistic interpretability techniques
In the (rather plausible) worst case scenario where CoT monitorability continues to degrade, can we fall back to using mechanistic interpretability techniques instead?
Scott Alexander brings us an excellent overview of mech interp, including linear probes, sparse autoencoders (SAEs), actuation verbalizers, natural language autoencoders, emotion vectors, and the J-space.
The bottom line, unfortunately, is that although interpretability techniques are interesting and provide some useful insights, they are nowhere close to being able to replace CoT monitoring.
CoT controllability evals seem very under-elicited
Astra’s system card reports that it’s substantially better than Sol at controlling its CoT in adversarial situations. Unfortunately, it may be even better than reported: Redwood’s Arun Jose finds that better prompting greatly improves performance on the commonly used CoTControl evaluation.
Strategy and politics
A few thoughts on pacing the frontier
prinz offers some thoughts on pacing the frontier. It’s a characteristically thoughtful piece, but I’m not convinced about this:
A slowdown in the pace of frontier model development is likely inevitable due to purely commercial considerations. Even if you think that AI is just a tool and poses no existential risk, you must recognize that the frontier labs cannot afford, from a purely commercial perspective, to release models (or to have models release themselves!) that will hack third parties or engage in scheming.
To some extent, prinz is obviously correct: mundane alignment is a valuable feature (although experience suggests that the market will put up with a substantial amount of misalignment in exchange for better capabilities).
But we’re in the endgame, and the incentives could easily drive a frontier lab to continue moving at full speed, keeping their most powerful models internal-only and publicly releasing less-capable but more trustworthy models.
Project Tailwind
Coefficient Giving announces Project Tailwind:
Project Tailwind is Coefficient Giving’s call for founders to launch ambitious AI safety initiatives. We’re looking for exceptional people to engage seriously with the risks of transformative AI, and to create the research, technologies, and institutions that will help humanity navigate them.
CG is scaling fast: they gave away $350M last year and are on track to direct more than $1B this year. If you’ve got the right idea and the right people, there’s plenty of money to get you started.
Risks
Jacob Coxon warns of human extinction and triggers a preference cascade
On September 8, a little-known Anthropic researcher named Jacob Coxon resigned from Anthropic because of fears about existential risk. Coming on the heels of the Hugging Face incident, this triggered a storm of media coverage and appears to mark a turning point in how seriously the mainstream media takes AI risk.
Zvi brings us the full report.
Countering misuse of AI
Anthropic brings us a detailed report on misuse of AI, with detailed case reports covering cyber, weapons development, biological misuse, and more.
The report provides concrete evidence that advances in cyber capabilities are translating into real world harm:
Sophisticated attacks no longer require sophisticated attackers
The cybersecurity skills of AI models means that AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators. In the case studies we report below, a hacktivist using stolen API keys, disparate financially motivated individuals, and a state espionage operator each sustained multi-victim campaigns that, even just a year ago, would have required many skilled operators and specialist knowledge.
The section on biological misuse is perhaps more concerning, although harder to interpret:
In the first example, a reseller platform evaded regional blocks to serve virologists working on a state-sponsored grant to pursue chikungunya gain-of-function work […] In the second, a researcher in an unsupported region spent weeks planning avian influenza mammalian-adaptation experiments with Claude
So how worried should we be about someone covertly doing gain-of-function research on chikungunya? Because bio capabilities are frequently dual-use, it’s almost impossible to know the true motivations of the researchers in question:
This may even occur to the extent that the researchers using our models may themselves be unaware of the intent and aims of their research. This has historical analogues: for example, the Soviet Biopreparat program—which was ultimately aimed at creating, producing, and weaponizing biological material—employed thousands of researchers, most of whom worked under the assumption that they were doing basic or defensive research because they were not informed about the program’s overall goal.
Congress should do something about Abliteration.ai
Sophie Kim reports on Abilteration.ai, which provides an abliterated version of GLM-5.3, with the (previously minimal) guardrails removed:
There are complicated policy questions surrounding open models, and reasonable people can disagree about whether highly capable open models are beneficial or harmful.
But let’s be clear: the danger is not that if an open model has latent capability for causing harm, a sufficiently motivated bad actor might find a way to access it. The risk is that someone will almost certainly make that capability conveniently available to all bad actors.
Scenarios for our Economic Future
The Anthropic Institute brings us an interactive model of how AI might affect jobs, growth, and unemployment by 2030. It’s a strange beast: on the one hand, it’s a well-executed model, with nicely crafted explanations and cool interactive features.
On the other hand, the whole endeavor is uncharacteristically not ASI-pilled. Their “extreme” scenario projects GDP growth of 32% by 2030, which is substantial but not even close to what might happen in a fast takeoff.
People and data
Will Huawei catch up to Nvidia by 2030?
Huawei will substantially improve its AI compute solutions by 2030, but export controls constrain its most important scaling levers, making it unlikely to catch up with Nvidia. We estimate Huawei will produce less than 4% as much AI compute as Nvidia in 2026. Without access to foreign memory, that share could be around 1% by 2028.
It’s vital that the US maintain a substantial compute advantage over China for two reasons:
- A large compute advantage increases the odds that the US wins the AI race
- The larger America’s lead, the more we can afford to slow AI development, which increases the odds that humanity wins the AI race
NavierStokes
OpenAI solves the Navier-Stokes problem
In 2000, the Clay Mathematics Institute designated seven Millennium Problems, allocating a $1 million prize for the solution to each one. These are considered some of the deepest and most difficult unsolved problems in mathematics, and each has been the subject of intense effort by some of the world’s leading mathematicians.
One problem (the Poincaré conjecture) was solved in 2010—the remainder were unsolved until OpenAI announced a solution to the Navier-Stokes existence and smoothness problem last week.
Zvi brings us the full analysis.
The Navier-Stokes problem is unambiguously one of the most important open problems in mathematics: solving it represents yet another major step forward for AI capabilities.
The most important part of OpenAI’s announcement has nothing to do with math. The result was produced not by Astra, but by a newer and more capable model:
Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve.
There has been some ugly drama about the discovery process. That’s beyond the scope of this newsletter, but the New York Times has good coverage of that part of the story.
There’s increasing worry that AI is hollowing out mathematics by producing prestigious proofs without building deep understanding. 25 Fields Medal winners have written an open letter expressing concern about how AI is impacting mathematics:
We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.
