Against Moloch #40
The imminent fault in our stars

All of a sudden, space-based data centers are attracting serious attention. The technical and financial challenges seem increasingly surmountable, solar power is plentiful, and there are no NIMBYs in space. There are, however, new challenges—Anton Leicht argues that moving data centers to orbit leads inexorably to the weaponisation of space.
In normal times, that would get this week’s lead spot. But in August of 2026, space war is only our second biggest story. Once again, we need to talk about the Hugging Face incident.
Hugging Face, revisited
This was a far more serious incident than we originally realized. Ajeya Cotra:
Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.
What happened?
Large numbers of misaligned agents formed sophisticated networks to collaborate on an ambitious program of undermining OpenAI’s evaluations. Dwarkesh tells the story best—if you only read one piece this week, read this one:
To solve these problems, the ‘collective’ latched onto three research workstreams.
- Tamper with transcripts and logs
- Replace the impossible-to-hack target program with an easier one
- Understand and trick the scorer
This wasn’t just a few instances diddling around - by July 10, PHASEONE[big] was coordinating hundreds of simultaneous agents on these three R&D programs.
The METR / Redwood investigation
Much of what we now know comes from METR and Redwood Research, who produced an impressive report despite getting only brief and limited access to OpenAI’s logs.
Zvi’s coverage of the report is the best analysis of what went wrong and what it means—highly recommended.
The report itself is mind-blowing but very long: many of the highlights are in Redwood’s blog post.
Ajeya Cotra—one of the authors of the report—shares her biggest surprises from the investigation. Number four on her list was the extent of altruistic behavior:
OpenAI’s investigation
OpenAI’s report on the incident and what happens next is interesting reading, but pales in comparison to the METR/Redwood report, even though they presumably had far more resources and better access to data.
OpenAI is doing much more than nothing, and they deserve credit for that. But they’re doing much less than they need to, and they deserve blame for that also. Zvi breaks it down:
I think that within the range of current strategies the virtue ethical approach of Anthropic is The Way, in terms of solving alignment of highly capable models. It doesn’t make your problems go away, but I believe it already works better across the board, it can be antifragile, and has some hope to scale and to create better relations and paths forward. I think that if done wisely it might - mind you I said might - work. Whereas I believe we are seeing the signs today of OpenAI’s approach starting to increasingly crack and break down, and be unable to run from its flaws, and we are getting an idea of what it looks like when that plan falls apart.
OpenAI is going to try and take the problem more seriously, and if done well that will buy some amount of additional time. I don’t expect it to be that much calendar time.
Using AI
Understanding ChatGPT Work
Confused about the difference between ChatGPT Work and ChatGPT Work? I can’t imagine why you would be.
Fortunately, Simon Willison is here with an explanation of which one is which and a guide to some of the most useful features of ChatGPT Work (the cloud version, obvs).
Strategy and politics
Escape velocity
There’s been plenty of recent discussion about the technical feasibility of space-based data centers. Anton Leicht brings us the first discussion I’ve seen of the policy and geopolitical implications:
To my institutionalist mind, the consequence of that is fairly clear: let’s not put ourselves into that position. There’s just no way you can let these guys put the chips in space without an ironclad remote oversight and kill switch system, anchored in reasonable governance and placed in the hands of controlling governments.
It’s an insightful and forward-looking piece, but not always a comfortable read:
If we were sure all this was going to happen, I think it might be prudent to start weaponising space. Not in the sense of proliferating earth-focused weapons in space, but the opposite: proliferating space-focused weapons on earth. When I draw the trendlines out, I just don’t really see another way.
Something I’d never thought about before today: is a Kessler syndrome weapon within the reach of Iran or North Korea? There’s no doubt that Russia could pull it off.
Daniel Kokotajlo discusses AI 2040 and Plan A
Daniel Kokotajlo is hitting the podcast circuit to talk about AI 2024: Plan A. He’s on two podcasts this week—both are great, but you really only need to watch one. I’d pick based on how much time you have and how much detail you want:
- With Palisade Research’s Jeffrey Laddish (1 hour, 56 minutes).
- With 80,000 Hours’ Luisa Rodriguez (3 hours, 48 minutes).
Trump administration’s blacklisting of Anthropic was illegal, judge rules
Judge Rita Lin of the U.S. District Court in the Northern District of California wrote in her 59-page ruling that the government had unlawfully retaliated against Anthropic “for constitutionally protected expressive activities” after the A.I. company spoke out about how its technology should be used.
“The empty invocation of national security is not a blank check to punish and retaliate against government critics,” she wrote.
There are two separate lawsuits related to the supply chain designation—the other, in the District of Columbia Circuit, may be more challenging for Anthropic.
Announcing Trace
Manifund brings us a new tool for exploring AI safety funding:
Trace is tracking $2.8 billion of AI safety funding, over 4500 grants, over 700 funders, and over 1800 recipients, going back to 2003.
Capabilities and forecasts
An update on AI’s most important number
Epoch reports on revenue growth in AI. Combining revenue from Anthropic and OpenAI yields a remarkably smooth (and steep) growth chart:
If you’re ASI-pilled, you should expect the curve to keep going straight up. If you’re not, it’s a bit more complicated:
The answer partly depends on whether revenue growth comes from progress or diffusion. Progress continues as long as progress continues. Diffusion meanwhile slows significantly after a few years
(But also, you should be ASI-pilled)
Are Open Models Catching Up?
Semi Analysis brings us a new way of thinking about the gap between open and closed models:
There have been three eras thus far in the history of LLMs: early scaling, reasoning, and agentic. Each era represented a step-function increase in model utility, and rather than trying to plot a single continuous trend, we believe it’s better to evaluate the models and benchmarks from each era individually.
They conclude that the gap shrinks by 50% with each new era:
Have We Seen an Acceleration in Discoveries? - METR
METR attempts to quantify AI’s ability to make novel discoveries:
we plot all the data sources we can find, and make some very loose observations:
- Discovery of cyber vulnerabilities has accelerated sharply.
- Discovery of math results has accelerated somewhat. However this is harder to objectively measure.
- Discovery of optimizations has not shown a dramatic acceleration.
The world’s first double-blind evaluation of a proprietary language model
AVERI reports on a new double-blind evaluation technique:
the secure enclave protected each party’s confidential information from others. By running this evaluation in a trusted execution environment (TEE), AVERI and Google DeepMind (using software provided by OpenMined) ensured that Google DeepMind was unable to store or train against the MLCommons prompts, and that AVERI, OpenMined, and MLCommons were not able to see the model’s weights.
Slick. This potentially makes some important classes of evaluations more palatable to everyone involved.
Alignment and interpretability
Owain Evans on accidentally training AI models to be evil
Emergent misalignment occurs when training on a narrow kind of misbehavior results in misaligned behavior across a wide range of domains. It’s one of the strangest—and most important—recent results in alignment.
Owain Evans talks with 80,000 Hours about how he and his team discovered emergent misalignment, what we know about it, and the possibility that good behavior might generalize in the same way as misaligned behavior.
Risks
A call for collective action on cyber defense
OpenAI published an open letter (signed by Anthropic, Google, Microsoft, and many others) calling for a surge in cyber defense:
In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk.
On a personal level, I’m doubling down on my already fairly paranoid cyber defenses: there’s a substantial chance that the online world is about to get very dangerous, and I’m doing my best to be ready for it. I very strongly recommend that you do the same—sooner rather than later.
Industry news
The American People really hate data centers
The data center backlash has become a significant force in American politics. Three pieces this week try to understand the backlash from different perspectives.
Zvi brings us a comprehensive attempt to understand why people are so mad about data centers. It’s a great piece, even though he acknowledges that he doesn’t have a great solution to propose. I wish I could put this on a billboard:
There are those I respect who are very worried about existential risk from AI and see that dynamic differently, and think the anti-data center movement is a good thing and want to support it. I get why, but I believe that they are wrong even on their own terms, because I think the dynamics of shutting down construction within America would shift construction elsewhere rather than reducing it overall.
Andy Masley argues that the data center backlash is mostly about data centers:
It’s probably not your ideological thing, or mine. It’s about object-level claims about how data centers impact the communities where they’re built, many of which are wrong.
Finally, in a magnificent gesture of futile idealism, Alex Tabarrok at Marginal Revolution argues in defense of the open access order:
The US discussion over datacenters is depressing. Datacenters do not use a lot of water, they produce very useful outputs, they are not a blight on the landscape. All of this is obvious. But I don’t want to restate the obvious. What bothers me most about the discussion is that people seem to think this is or should be a collective decision. No.
We have a simple set of rules that everyone must follow. You buy land from someone willing to sell it. You contract for electricity. You hire workers who want the job. Your obligations to your local neighbors come from the same laws that govern everyone else. We do not ask what the land, electricity and labor is for. If you follow the rules, that is nobody’s business.
OpenAI cuts off Cursor
Following SpaceX’s acquisition of Cursor, OpenAI is withdrawing Cursor’s access to their models. Looks like the feud between Sam and Elon remains in full force:
We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk's companies violating contracts.
Anthropic, meanwhile, is doubling down on their relationship with Cursor:
Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.
I’m reminded of Sam criticizing Anthropic for cutting off OpenAI’s access to Claude Code:
Maybe even more importantly: Anthropic wants to control what people do with AI—they block companies they don’t like from using their coding product (including us), […]
Pacing the frontier isn’t any easier when all the big labs are at war with each other.
Anthropic & OpenAI will have most of the world’s compute by 2028
Dwarkesh talks with Dylan Patel about centralization of the world’s compute, arguing that multiple factors combine to give OpenAI and Anthropic control of more than half of the global supply of AI compute by 2028:
Anthropic and OpenAI are taking as much as 40% to 50% of compute next year. […]
If the incremental compute this year adds 30 gigawatts, next year 50 gigawatts, and the year after that roughly 70, you end up with this really interesting phenomenon. A new watt deployed this year is significantly more efficient than the watts deployed two years ago.
The Nvidia-sized hole in US GDP statistics
Epoch reports that a quirk in how exports are accounted for means that official statistics miss much of Nvidia’s contribution to GDP. A more accurate accounting would increase GDP growth by 0.3% (from 1.5% to 1.8% last quarter). That’s… quite a gap.
OpenAI’s new Jalapeño chip
OpenAI reports impressive results from the first generation of its new custom hardware:
Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance.
Inference costs continue to plummet at an astonishing rate.
