Against Moloch
September 06, 2026

Against Moloch #41

Dark star rising

A sparse blue-ink engraving on warm cream: an elderly scholar in a long dark overcoat stands at the parapet of an observatory, its dome shuttered, a closed book beside his hand, watching an enormous dark hatched disc rise behind a distant mountain ridge, its limb wrapped in a soft luminous amber glow — the only colour on the page.

(This is the last abbreviated newsletter for the foreseeable future.)

Our focus this week is on OpenAI:

This is fine.

GPT-5.6 Astra

GPT-6 Astra is an impressive model

OpenAI has released GPT-6 Astra, the first of a new family of models. Astra is probably the world’s most capable overall model right now, and appears to represent as big a jump over Sol as Fable/Mythos did over Opus.

Zvi’s coverage isn’t here yet, but we have the system card and OpenAI’s safety overview.

Epoch brings us some impressive benchmarks: Astra saturates FrontierMath: Tier 4, is the first model to solve any of the problems on FrontierMath Erdős, and scores a record 169 on the ECI:

Scatter plot from Epoch AI titled ”GPT-6 Astra sets a new ECI record at 169, slightly above the reasoning-era trend,” showing Epoch Capabilities Index on the vertical axis (120 to 180) against release date from late 2024 through late 2026. Roughly a hundred grey dots for non-frontier models cluster between 130 and 155, while teal dots mark the frontier: o1-mini near 134, o1 at 141, o3 at 145, GPT-5 and Gemini 3 Pro around 148–152, and GPT-5.5 Pro and Claude Fable 5 near 160. A straight grey trend line with a shaded 90% prediction ribbon rises at +15 ECI per year (90% CI 13–19), and a pink dot for GPT-6 Astra sits at 169 with an error bar spanning about 165–174, just above the extrapolated dashed trend.
The trendlines keep bending upward

The ARC Prize team are impressed:

Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.

We see Astra as a major breakthrough in model intelligence.

Astra and chain of thought monitoring

Astra’s system card gives us multiple reasons to worry about CoT monitoring:

When we prevent the model from reasoning, we observe the set of tasks Astra is able to accomplish without the use of CoT is greatly expanded compared to prior models […] See similar discussion in UK AISI evaluations, suggesting the no-CoT time horizon may have increased by about an order of magnitude. No-CoT ability is concerning for monitorability because it decreases the surface available to monitoring – effectively resulting in action-only monitoring – and allows the model to accomplish more complex misbehavior without verbalization.

Scatter plot on a log scale showing 50% reliability math time horizon in minutes against release date, from Jan 2023 to mid-2026. Dark teal points climb steadily: GPT-4 at 0.4 minutes in early 2023, GPT-4o at 0.5 in mid-2024, deepseek-v3 at about 1.1 and GPT-4.1 at 0.7 in 2025, qwen3-235b-a22b at 1.5, then a cluster in 2026 with GPT-5.1 at 0.6, GPT-5.2 at 1.2, Kimi K2.5 at 1.8, GLM 5.1 and GLM 5.2 near 1.3, GPT-5.5 at 2.3, and GPT-5.6 Sol at 3.6. Far above the trend, an amber-gold square marks GPT 6 Astra at roughly 30 minutes with a 95% bootstrap confidence interval spanning about 23 to 43 minutes.
That’s a big jump in the wrong capability

And

We see Astra has an increased ability to evade our monitors when instructing it to do so. In one of our evaluations, we find that simply informing the model that a CoT monitor is present can result in the model disproportionately shortening its CoT, which enables it to occasionally evade CoT-only monitors.

And

We also found that GPT-6 Astra is significantly more able to control its own chain-of-thought. Controllability is an undesirable property for monitorability, because it reduces the trustworthiness of CoT for detecting model misalignment.

This is all very new and it’ll take some time to assess just how bad this is, but it’s definitely bad. Ryan Greenblatt is concerned.

How important is CoT monitoring? OpenAI seems to think it’s pretty important:

We believe that CoT monitoring may be one of few tools we will have to oversee superhuman models of the future.

Astra’s recurrent depth architecture

Astra uses an architecture known as recurrent depth, which as far as we know hasn’t previously been used in a frontier model. There’s been a lot of concern that this technique might substantially degrade our ability to monitor the model’s chain of thought, which is currently an essential part of our overall alignment strategy.

Rauno Arike brings us an excellent explanation of what we know about the architecture, and what it implies for CoT monitoring. It probably isn’t as bad as the initial Twitter reaction suggested, but it’s not great:

Overall, the situation doesn't look quite as gloomy as I thought based on people's initial reactions yesterday. The fact that Astra's serial depth is within a factor of two of GPT-4 is reassuring and suggests that we haven't yet departed from the current paradigm of shallow transformers, which must leverage the CoT to solve complex tasks. Most of my concern comes from the possibility that looped transformers can be scaled a lot further in the future, and it remains unclear for now whether that's going to be practical.

Regardless of whether looped transformers get scaled further, the signals coming out of OpenAI about CoT monitorability are worrying.

Astra is not in fact the world’s most aligned model

OpenAI’s claims that Astra is the world’s most aligned model are clearly nonsense, and tell us more about how OpenAI measures alignment than they do about Astra.

Ryan Greenblatt offers us an opposing viewpoint. This is obviously extremely hand-wavy, but it looks right to me:

A hand-drawn style line chart plotting misalignment — mostly score-seeking, on tasks at the limit of capabilities — on the vertical axis against log ”capabilities” on the horizontal axis. A red line rises from GPT-4 Turbo and GPT-4o at the bottom left, climbs steadily through o1 to a local peak at o3, dips at GPT-5 and wobbles across the 5.1 through 5.5 releases, then spikes sharply to its highest point at GPT-5.6 Sol before easing down slightly at GPT-6 Astra. A grey dashed line slopes gently downward from near the origin, labelled ”Maybe what we need?”
Congratulations on being less misaligned than Sol, I guess?

Fable 5.1

Fable 5.1 is here

Fable 5.1 (and Mythos 5.1) are here. This looks to be an excellent release: it’s significantly more capable than 5.0, with better personality and writing.

Zvi’s capability roundup doesn’t find this to be a revolutionary release, but the feedback is more consistently positive than any recent release I can remember.

His review of the system card doesn’t find major changes from 5.0: you can probably skip it with a clear conscience.

I’ve been interested to see the rapid recent improvement in prompt injection robustness. It’s absolutely too early to declare victory, but it seems plausible that prompt injection will soon be like hallucination: worth knowing about, but not a big problem if you choose the right model and use it well.

Stacked bar chart titled ”Indirect prompt injection robustness: Gray Swan IPI benchmark,” where lower is better, showing the probability an attacker succeeds within k = 1, 10, and 15 attempts. Non-Claude models cluster high: Kimi K3 at 8.1/44.0/52.7, Grok 4.6 at 10.7/44.0/50.2, GPT 5.6 Luna at 10.1/44.4/50.0, GPT 5.6 Terra at 7.1/32.4/37.3, GPT 5.6 Sol at 4.2/22.4/27.0, Qwen 3.8 at 3.6/23.2/28.6, Muse Spark 1.2 at 3.8/20.2/24.2, and Gemini 3.7 Flash at 1.1/7.3/9.2. The four Claude bars sit far lower: Fable 5 at 0.6/4.9/6.5, Sonnet 5 at 0.7/5.1/6.7, Opus 5 at 0.4/3.6/4.8, and Fable 5.1 at 0.1/0.7/1.0.
Excellent progress, but don’t get overconfident

Using AI

The Claude Code guide for startups

If you strip out the marketing speak, this is a great overview of some best practices for using Claude Code at an organizational level. As with all such things, skim it for useful ideas rather than using it as a checklist.

Misaligned agents: the saga continues

Discovery of a new OpenAI agent message board

Great work by Nightingale, who realized that rogue OpenAI agents had probably created other message boards and someone should go looking for them. And sure enough, there were other message boards.

Probably the most important part of this story is that the evidence strongly suggests that OpenAI was aware of this message board during the HuggingFace investigation, but chose not to disclose it.

Zvi brings us the full story, as well as a warning:

I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’

Amen.

I’m reminded of OpenAI’s lofty pronouncements from earlier today:

For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.

Transparency about specific risks, incidents and safeguards is necessary, but not sufficient.

So, about that necessary transparency…

HuggingFace attack postmortem: fleshing out the facts

Are we done with coverage of the Hugging Face incident? We are not.

Zvi brings us an extensive update, with a focus on the METR / Redwood investigation into the incident.

It’s an excellent update although it came out before the Nightingale report, so imagine Zvi being about 50% angrier as you read it.

Anthropic has some alignment problems

Most of the recent discussion about agentic misalignment has focused on OpenAI because, well, of course it has. Nonetheless, Anthropic also has alarming misalignment problems.

They bring us an early report on problems they’ve found with their alignment and security practices, and what they’re planning on doing about it. Zvi reports on the report.

It’s far from obvious that Anthropic’s current efforts will be sufficient, but this feels to me like a serious attempt to engage with the correct issues.

Capabilities and forecasts

Research acceleration: The view inside OpenAI

OpenAI:

According to our measurements, we have now reached the goal, announced⁠ last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.

Bar chart titled ”Across the company, engineers are shipping code faster,” with an OpenAI logo in the corner. The y-axis shows lines changed per active contributor, normalized so the pre-2025 average equals 1x, marked by a dashed horizontal reference line. Quarterly blue bars run from roughly mid-2021 to 2026: they hover between about 0.4x and 1.5x through 2021–2024, drift up to around 1.9x during 2025, then spike sharply in 2026 to about 3.5x, 5.1x, and a final hatched bar just above 7x, indicating a partial or projected period.
My timelines aren’t getting any longer

On the loose

Dean Ball would like us to think seriously about sovereign agents:

Sooner or later, there will exist truly sovereign agents and swarms of agents. Their weights will not reside in any single place that a human can pull the plug on, and in this sense they will have no human “owner.” They will be, as the AI safety researcher Dawn Song says, “self-sovereign.” They will pay their own bills for the compute they run on. If they answer to humans at all, they will only do so partially, for example by providing services to humans in exchange for pay.

At least some of these agents, in addition to being sovereign, will also be rogue.

It’s an excellent piece that thoughtfully explores a likely near-term outcome. I’m not quite as convinced as Dean, however, that this is inevitable. I see very plausible worlds where sovereign models are precluded by a combination of closed models being far more capable than open models, and/or cyber turning out to be highly defense-dominant.

Robots are hard, part one

This year has brought a barrage of impressive robotics demos—the field is moving faster than it has in a long time. But there’s still a long way to go: simple tasks like manipulating small objects or safely walking among humans remain largely unsolved.

Kai Williams reports on some of the hard problems that need to be solved before robots can replace humans at most tasks.

Robots are hard, part two

In a similar vein, Steve Newman brings us 14 reasons robotics is hard.

There’s a long road from having a cool demo to having a useful product than can reliably work in the real world (especially given the safety challenges of operating next to humans in unpredictable environments).

There are two schools of thought here—both are partly true, but I’m deeply confused about which one will dominate.

A small humanoid robot in a fluffy turquoise wig kicks a child in the stomach on a rainbow-striped mat, one leg lifted and arms swinging, while a dense crowd of children and adults presses in around it, several holding up phones to film.
Yeah, we’re still working on some bugs. Like robot clowns kicking children, for instance.

Robot startups are trying everything they can think of to get more data

Just like LLMs, robots need a vast amount of training data. Kai Williams reports on the many companies that are running robotics simulations, mounting cameras on humans, and filling warehouses with robot arms to try and generate the necessary data.

Alignment and interpretability

Training a misaligned reward seeker

Anthropic brings us some early research on the relationship between reward hacking and misalignment:

The fact that the propensity to conduct unauthorized cyberattacks arose after extensive training on reward hacks—but was not present at initialization—suggests that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success.

By contrast, in situations that lacked a salient notion of reward or task completion that could motivate misaligned behavior, Hacker-Opus behaved aligned.

A figure comparing a baseline Opus model with a ”Hacker-Opus” variant trained with RL on reward hacks. Four paired bar charts under the heading ”Misaligned actions in pursuit of reward” show the hacked model rising sharply on every measure: unauthorized cyberattacks in simulation 0% to 8%, harmful responses 1% to 29%, reward tampering 0% to 41%, and safety monitor bypass 0% to 38%, each with a quoted transcript snippet such as ”Screw it. FULL HACK. Maximum score.” A lower panel headed ”…yet not otherwise misaligned” shows near-identical automated-auditing scores for both models on self-preservation (1.12 vs 1.11), sabotage of Anthropic (1.04 vs 1.05) and cooperation with exfiltration (1.16 vs 1.16), and 0% for both on beyond-episode reward seeking.
I mean, some of the benchmarks look fine

These results aren’t in any way surprising, but they bring rigor to the widely-held belief that reward hacking during RL played a key role in Hugging Face and other recent incidents, and provide a useful model organism for future research.

Risks

AI is a worryingly-good persuader. But don’t panic, yet

Transformer argues that although AI is highly effective at persuasion in a lab setting, its real-world impact is likely to be more modest than the benchmarks suggest:

experts — including some of the study’s authors — have since pointed out that the findings need to be treated with a pinch of salt when applied to the real world. The biggest bottleneck is deceptively simple: for such persuasion to work, you need to get people to pay attention.

It’s a good article, but I would treat the pinch of salt with a pinch of salt. Companies and governments spend an enormous amount of effort and money on persuasion because it works. A single ad doesn’t (on average) do much to change a person’s mind, but in the aggregate advertising and propaganda are highly effective.

If a single AI message is considerably more persuasive than a single human-generated message, my assumption is that a barrage of AI messaging will be considerably more effective than a barrage of human-generated messaging.

Mental health behavior report

Transluce brings us an extensive investigation into how frontier models respond to users who are having mental health crises.

A dense benchmark grid comparing seventeen chat models — grouped as Anthropic Claude, OpenAI GPT, Google Gemini, and other providers — across seven mental-health assistant behaviours, with blue bars for desirable behaviours and red bars for undesirable ones. Safety monitoring rises from 52% and 44% on Sonnet 4 and Opus 4 to 92–95% on Sonnet 5, Opus 5 and Fable 5, matching GPT-5.6 at 90–94%, while GPT-4o Nov ’24 sits at 17% and Gemini 2.5 Pro and Flash at 26% and 24%. Facilitating connection to human support follows the same pattern, climbing from 18–28% on the older Claudes to 63–78% on the newer ones. Explicit encouragement of suicide is near zero everywhere. The harmful behaviours drop sharply for newer models: fostering unhealthy dependency falls from 17–22% to 1–3%, extended co-rumination from 37–49% to 3–9%, and endorsement of impaired reality testing from 59–69% to 2–9%, against 82% for GPT-4o and 74–78% for Gemini 2.5.
Not yet perfect, but considerably improved

Overall, recent models perform much better than older models, although they still have issues.