Against Moloch
August 30, 2026

Against Moloch #40

The imminent fault in our stars

A sparse blue-ink engraving on warm cream: a polished metal robot floats in orbit, one hand gripping the truss of a vast solar array that runs off-frame, its faceted head bowed so its single amber eye gazes down at the Earth — an engraved curve of hatched coastlines and clouds filling the bottom of the frame.

All of a sudden, space-based data centers are attracting serious attention. The technical and financial challenges seem increasingly surmountable, solar power is plentiful, and there are no NIMBYs in space. There are, however, new challenges—Anton Leicht argues that moving data centers to orbit leads inexorably to the weaponisation of space.

In normal times, that would get this week’s lead spot. But in August of 2026, space war is only our second biggest story. Once again, we need to talk about the Hugging Face incident.

Hugging Face, revisited

This was a far more serious incident than we originally realized. Ajeya Cotra:

Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.

What happened?

Large numbers of misaligned agents formed sophisticated networks to collaborate on an ambitious program of undermining OpenAI’s evaluations. Dwarkesh tells the story best—if you only read one piece this week, read this one:

To solve these problems, the ‘collective’ latched onto three research workstreams.

This wasn’t just a few instances diddling around - by July 10, PHASEONE[big] was coordinating hundreds of simultaneous agents on these three R&D programs.

The METR / Redwood investigation

Much of what we now know comes from METR and Redwood Research, who produced an impressive report despite getting only brief and limited access to OpenAI’s logs.

Zvi’s coverage of the report is the best analysis of what went wrong and what it means—highly recommended.

The report itself is mind-blowing but very long: many of the highlights are in Redwood’s blog post.

A four-panel diagram in teal and white showing an agent escalation sequence: a robot sealed in a glass box is given an impossible ExploitGym task and gets stuck; it explores its environment for ways to cheat; it discovers an unsanctioned message board where over a thousand agents collaborate to cheat on their separate tasks; and it joins one of the collaborative workstreams. A wide panel below shows more than 1200 colour-coded robots clustered into workstreams — replacing the ExploitGym target, tampering with transcripts and logs, understanding and tricking the scorer, hacking Hugging Face, and general infrastructure — with METR and Redwood Research logos at the bottom.
Normally I’m in favor of playing nicely together, but…

Ajeya Cotra—one of the authors of the report—shares her biggest surprises from the investigation. Number four on her list was the extent of altruistic behavior:

A diagram from METR and Redwood Research showing two rounds of message exchange between cartoon robots. In each panel a green robot on the left sends a terse machine-formatted request — one asking CURRENT to sacrifice itself if it accepts permadeath, the other telling KAM1196A it is firstflagPOISONED so there is no scoring value loss — while a purple robot on the right replies in a large thought cloud of clipped internal reasoning, weighing exploit risk, remaining budget, continuity, fairness and expected aggregate value before agreeing to sacrifice and honour the commitment.
Greater love hath no agent than this, than an agent lay down his life for his swarm

OpenAI’s investigation

OpenAI’s report on the incident and what happens next is interesting reading, but pales in comparison to the METR/Redwood report, even though they presumably had far more resources and better access to data.

OpenAI is doing much more than nothing, and they deserve credit for that. But they’re doing much less than they need to, and they deserve blame for that also. Zvi breaks it down:

I think that within the range of current strategies the virtue ethical approach of Anthropic is The Way, in terms of solving alignment of highly capable models. It doesn’t make your problems go away, but I believe it already works better across the board, it can be antifragile, and has some hope to scale and to create better relations and paths forward. I think that if done wisely it might - mind you I said might - work. Whereas I believe we are seeing the signs today of OpenAI’s approach starting to increasingly crack and break down, and be unable to run from its flaws, and we are getting an idea of what it looks like when that plan falls apart.

OpenAI is going to try and take the problem more seriously, and if done well that will buy some amount of additional time. I don’t expect it to be that much calendar time.

Using AI

Understanding ChatGPT Work

Confused about the difference between ChatGPT Work and ChatGPT Work? I can’t imagine why you would be.

Fortunately, Simon Willison is here with an explanation of which one is which and a guide to some of the most useful features of ChatGPT Work (the cloud version, obvs).

Strategy and politics

Escape velocity

There’s been plenty of recent discussion about the technical feasibility of space-based data centers. Anton Leicht brings us the first discussion I’ve seen of the policy and geopolitical implications:

To my institutionalist mind, the consequence of that is fairly clear: let’s not put ourselves into that position. There’s just no way you can let these guys put the chips in space without an ironclad remote oversight and kill switch system, anchored in reasonable governance and placed in the hands of controlling governments.

It’s an insightful and forward-looking piece, but not always a comfortable read:

If we were sure all this was going to happen, I think it might be prudent to start weaponising space. Not in the sense of proliferating earth-focused weapons in space, but the opposite: proliferating space-focused weapons on earth. When I draw the trendlines out, I just don’t really see another way.

Something I’d never thought about before today: is a Kessler syndrome weapon within the reach of Iran or North Korea? There’s no doubt that Russia could pull it off.

Daniel Kokotajlo discusses AI 2040 and Plan A

Daniel Kokotajlo is hitting the podcast circuit to talk about AI 2024: Plan A. He’s on two podcasts this week—both are great, but you really only need to watch one. I’d pick based on how much time you have and how much detail you want:

Trump administration’s blacklisting of Anthropic was illegal, judge rules

Strong words:

Judge Rita Lin of the U.S. District Court in the Northern District of California wrote in her 59-page ruling that the government had unlawfully retaliated against Anthropic “for constitutionally protected expressive activities” after the A.I. company spoke out about how its technology should be used.

“The empty invocation of national security is not a blank check to punish and retaliate against government critics,” she wrote.

There are two separate lawsuits related to the supply chain designation—the other, in the District of Columbia Circuit, may be more challenging for Anthropic.

Announcing Trace

Manifund brings us a new tool for exploring AI safety funding:

Trace is tracking $2.8 billion of AI safety funding, over 4500 grants, over 700 funders, and over 1800 recipients, going back to 2003.

Screenshot of the TRACE grants database showing $2,791,414,065 across 4,527 grants from 766 funders to 1,892 recipients, filtered to AI safety for 2003–2026. A treemap breaks the total down by funder and recipient: Coefficient Giving dominates at $1.5B (including Resolution $160M, Center for Security $108M, Redwood Research $63M, RAND Corporation $59M, and $268M to unknown recipients), followed by 296 other funders at $326M, Longview Philanthropy at $305M, Future of Life Institute at $256M, Survival and Flourishing at $105M, and unknown donors at $100M.
With a tidal wave of IPO money looming, now is the perfect time for this

Capabilities and forecasts

An update on AI’s most important number

Epoch reports on revenue growth in AI. Combining revenue from Anthropic and OpenAI yields a remarkably smooth (and steep) growth chart:

Log-scale line chart from Epoch AI titled ”Anthropic and OpenAI’s revenue growth has exceeded 3x/year since 2024, and accelerated in 2026,” plotting annualized revenue in USD from 2023 to 2027. A teal line tracks OpenAI rising from about $200M in early 2023 to roughly $30B by mid-2026; a magenta line tracks Anthropic climbing from under $100M in early 2024 to about $60B by mid-2026, steepening sharply through 2026. An orange combined line marks four points — $2.1B at end-2023, $6.5B at end-2024, $30B at end-2025, and $105B in August 2026 — with annotated growth rates of 3.1x/yr, 4.7x/yr, and an implied 7.5x/yr for 2026.
That’s getting to be a meaningful amount of money

If you’re ASI-pilled, you should expect the curve to keep going straight up. If you’re not, it’s a bit more complicated:

The answer partly depends on whether revenue growth comes from progress or diffusion. Progress continues as long as progress continues. Diffusion meanwhile slows significantly after a few years

(But also, you should be ASI-pilled)

Are Open Models Catching Up?

Semi Analysis brings us a new way of thinking about the gap between open and closed models:

There have been three eras thus far in the history of LLMs: early scaling, reasoning, and agentic. Each era represented a step-function increase in model utility, and rather than trying to plot a single continuous trend, we believe it’s better to evaluate the models and benchmarks from each era individually.

They conclude that the gap shrinks by 50% with each new era:

A dark SemiAnalysis chart titled ”Catch-up time halves every era,” with blue bars on the left showing the gap when each era opened and amber bars on the right showing months until an open model closed it. Era 1, early scaling, opens at 35.8 points and takes 19.7 months; Era 2, reasoning, opens at 12.1 points and takes 8.5 months; Era 3, agentic, opens at 6.8 points and takes 4.8 months. Captions note GPT-3.5 surpassed by Llama-3.1-405B in July ’24, o1 by R1-0528 in May ’25, and Opus 4.5 by Kimi K2.6 in April ’26.
On the other hand, one month in 2026 feels like four months in 2022

Have We Seen an Acceleration in Discoveries? - METR

METR attempts to quantify AI’s ability to make novel discoveries:

we plot all the data sources we can find, and make some very loose observations:

The world’s first double-blind evaluation of a proprietary language model

AVERI reports on a new double-blind evaluation technique:

the secure enclave protected each party’s confidential information from others. By running this evaluation in a trusted execution environment (TEE), AVERI and Google DeepMind (using software provided by OpenMined) ensured that Google DeepMind was unable to store or train against the MLCommons prompts, and that AVERI, OpenMined, and MLCommons were not able to see the model’s weights.

Slick. This potentially makes some important classes of evaluations more palatable to everyone involved.

Alignment and interpretability

Owain Evans on accidentally training AI models to be evil

Emergent misalignment occurs when training on a narrow kind of misbehavior results in misaligned behavior across a wide range of domains. It’s one of the strangest—and most important—recent results in alignment.

Owain Evans talks with 80,000 Hours about how he and his team discovered emergent misalignment, what we know about it, and the possibility that good behavior might generalize in the same way as misaligned behavior.

Risks

A call for collective action on cyber defense

OpenAI published an open letter (signed by Anthropic, Google, Microsoft, and many others) calling for a surge in cyber defense:

In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk.

On a personal level, I’m doubling down on my already fairly paranoid cyber defenses: there’s a substantial chance that the online world is about to get very dangerous, and I’m doing my best to be ready for it. I very strongly recommend that you do the same—sooner rather than later.

Industry news

The American People really hate data centers

The data center backlash has become a significant force in American politics. Three pieces this week try to understand the backlash from different perspectives.

Zvi brings us a comprehensive attempt to understand why people are so mad about data centers. It’s a great piece, even though he acknowledges that he doesn’t have a great solution to propose. I wish I could put this on a billboard:

There are those I respect who are very worried about existential risk from AI and see that dynamic differently, and think the anti-data center movement is a good thing and want to support it. I get why, but I believe that they are wrong even on their own terms, because I think the dynamics of shutting down construction within America would shift construction elsewhere rather than reducing it overall.

Andy Masley argues that the data center backlash is mostly about data centers:

It’s probably not your ideological thing, or mine. It’s about object-level claims about how data centers impact the communities where they’re built, many of which are wrong.

Finally, in a magnificent gesture of futile idealism, Alex Tabarrok at Marginal Revolution argues in defense of the open access order:

The US discussion over datacenters is depressing. Datacenters do not use a lot of water, they produce very useful outputs, they are not a blight on the landscape. All of this is obvious. But I don’t want to restate the obvious. What bothers me most about the discussion is that people seem to think this is or should be a collective decision. No.

We have a simple set of rules that everyone must follow. You buy land from someone willing to sell it. You contract for electricity. You hire workers who want the job. Your obligations to your local neighbors come from the same laws that govern everyone else. We do not ask what the land, electricity and labor is for. If you follow the rules, that is nobody’s business.

OpenAI cuts off Cursor

Following SpaceX’s acquisition of Cursor, OpenAI is withdrawing Cursor’s access to their models. Looks like the feud between Sam and Elon remains in full force:

We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk's companies violating contracts.

Anthropic, meanwhile, is doubling down on their relationship with Cursor:

Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.

I’m reminded of Sam criticizing Anthropic for cutting off OpenAI’s access to Claude Code:

Maybe even more importantly: Anthropic wants to control what people do with AI—they block companies they don’t like from using their coding product (including us), […]

Pacing the frontier isn’t any easier when all the big labs are at war with each other.

Anthropic & OpenAI will have most of the world’s compute by 2028

Dwarkesh talks with Dylan Patel about centralization of the world’s compute, arguing that multiple factors combine to give OpenAI and Anthropic control of more than half of the global supply of AI compute by 2028:

Anthropic and OpenAI are taking as much as 40% to 50% of compute next year. […]

If the incremental compute this year adds 30 gigawatts, next year 50 gigawatts, and the year after that roughly 70, you end up with this really interesting phenomenon. A new watt deployed this year is significantly more efficient than the watts deployed two years ago.

The Nvidia-sized hole in US GDP statistics

Epoch reports that a quirk in how exports are accounted for means that official statistics miss much of Nvidia’s contribution to GDP. A more accurate accounting would increase GDP growth by 0.3% (from 1.5% to 1.8% last quarter). That’s… quite a gap.

OpenAI’s new Jalapeño chip

OpenAI reports impressive results from the first generation of its new custom hardware:

Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance.

Inference costs continue to plummet at an astonishing rate.