Against Moloch #41
Dark star rising

(This is the last abbreviated newsletter for the foreseeable future.)
Our focus this week is on OpenAI:
- GPT-6 Astra is probably the smartest model in the world.
- Astra appears to have substantially degraded CoT monitorability, which handicaps one of our most important safety tools.
- Yet again, there are important new revelations about the extent of misaligned behavior by OpenAI models.
- There’s increasing evidence that OpenAI has been less than forthcoming about the extent of their misalignment problems.
This is fine.
GPT-5.6 Astra
GPT-6 Astra is an impressive model
OpenAI has released GPT-6 Astra, the first of a new family of models. Astra is probably the world’s most capable overall model right now, and appears to represent as big a jump over Sol as Fable/Mythos did over Opus.
Zvi’s coverage isn’t here yet, but we have the system card and OpenAI’s safety overview.
Epoch brings us some impressive benchmarks: Astra saturates FrontierMath: Tier 4, is the first model to solve any of the problems on FrontierMath Erdős, and scores a record 169 on the ECI:
The ARC Prize team are impressed:
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
We see Astra as a major breakthrough in model intelligence.
Astra and chain of thought monitoring
Astra’s system card gives us multiple reasons to worry about CoT monitoring:
When we prevent the model from reasoning, we observe the set of tasks Astra is able to accomplish without the use of CoT is greatly expanded compared to prior models […] See similar discussion in UK AISI evaluations, suggesting the no-CoT time horizon may have increased by about an order of magnitude. No-CoT ability is concerning for monitorability because it decreases the surface available to monitoring – effectively resulting in action-only monitoring – and allows the model to accomplish more complex misbehavior without verbalization.
And
We see Astra has an increased ability to evade our monitors when instructing it to do so. In one of our evaluations, we find that simply informing the model that a CoT monitor is present can result in the model disproportionately shortening its CoT, which enables it to occasionally evade CoT-only monitors.
And
We also found that GPT-6 Astra is significantly more able to control its own chain-of-thought. Controllability is an undesirable property for monitorability, because it reduces the trustworthiness of CoT for detecting model misalignment.
This is all very new and it’ll take some time to assess just how bad this is, but it’s definitely bad. Ryan Greenblatt is concerned.
How important is CoT monitoring? OpenAI seems to think it’s pretty important:
We believe that CoT monitoring may be one of few tools we will have to oversee superhuman models of the future.
Astra’s recurrent depth architecture
Astra uses an architecture known as recurrent depth, which as far as we know hasn’t previously been used in a frontier model. There’s been a lot of concern that this technique might substantially degrade our ability to monitor the model’s chain of thought, which is currently an essential part of our overall alignment strategy.
Rauno Arike brings us an excellent explanation of what we know about the architecture, and what it implies for CoT monitoring. It probably isn’t as bad as the initial Twitter reaction suggested, but it’s not great:
Overall, the situation doesn't look quite as gloomy as I thought based on people's initial reactions yesterday. The fact that Astra's serial depth is within a factor of two of GPT-4 is reassuring and suggests that we haven't yet departed from the current paradigm of shallow transformers, which must leverage the CoT to solve complex tasks. Most of my concern comes from the possibility that looped transformers can be scaled a lot further in the future, and it remains unclear for now whether that's going to be practical.
Regardless of whether looped transformers get scaled further, the signals coming out of OpenAI about CoT monitorability are worrying.
Astra is not in fact the world’s most aligned model
OpenAI’s claims that Astra is the world’s most aligned model are clearly nonsense, and tell us more about how OpenAI measures alignment than they do about Astra.
Ryan Greenblatt offers us an opposing viewpoint. This is obviously extremely hand-wavy, but it looks right to me:
Fable 5.1
Fable 5.1 is here
Fable 5.1 (and Mythos 5.1) are here. This looks to be an excellent release: it’s significantly more capable than 5.0, with better personality and writing.
Zvi’s capability roundup doesn’t find this to be a revolutionary release, but the feedback is more consistently positive than any recent release I can remember.
His review of the system card doesn’t find major changes from 5.0: you can probably skip it with a clear conscience.
I’ve been interested to see the rapid recent improvement in prompt injection robustness. It’s absolutely too early to declare victory, but it seems plausible that prompt injection will soon be like hallucination: worth knowing about, but not a big problem if you choose the right model and use it well.
Using AI
The Claude Code guide for startups
If you strip out the marketing speak, this is a great overview of some best practices for using Claude Code at an organizational level. As with all such things, skim it for useful ideas rather than using it as a checklist.
Misaligned agents: the saga continues
Discovery of a new OpenAI agent message board
Great work by Nightingale, who realized that rogue OpenAI agents had probably created other message boards and someone should go looking for them. And sure enough, there were other message boards.
Probably the most important part of this story is that the evidence strongly suggests that OpenAI was aware of this message board during the HuggingFace investigation, but chose not to disclose it.
Zvi brings us the full story, as well as a warning:
I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’
Amen.
I’m reminded of OpenAI’s lofty pronouncements from earlier today:
For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.
Transparency about specific risks, incidents and safeguards is necessary, but not sufficient.
So, about that necessary transparency…
HuggingFace attack postmortem: fleshing out the facts
Are we done with coverage of the Hugging Face incident? We are not.
Zvi brings us an extensive update, with a focus on the METR / Redwood investigation into the incident.
It’s an excellent update although it came out before the Nightingale report, so imagine Zvi being about 50% angrier as you read it.
Anthropic has some alignment problems
Most of the recent discussion about agentic misalignment has focused on OpenAI because, well, of course it has. Nonetheless, Anthropic also has alarming misalignment problems.
They bring us an early report on problems they’ve found with their alignment and security practices, and what they’re planning on doing about it. Zvi reports on the report.
It’s far from obvious that Anthropic’s current efforts will be sufficient, but this feels to me like a serious attempt to engage with the correct issues.
Capabilities and forecasts
Research acceleration: The view inside OpenAI
According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.
On the loose
Dean Ball would like us to think seriously about sovereign agents:
Sooner or later, there will exist truly sovereign agents and swarms of agents. Their weights will not reside in any single place that a human can pull the plug on, and in this sense they will have no human “owner.” They will be, as the AI safety researcher Dawn Song says, “self-sovereign.” They will pay their own bills for the compute they run on. If they answer to humans at all, they will only do so partially, for example by providing services to humans in exchange for pay.
At least some of these agents, in addition to being sovereign, will also be rogue.
It’s an excellent piece that thoughtfully explores a likely near-term outcome. I’m not quite as convinced as Dean, however, that this is inevitable. I see very plausible worlds where sovereign models are precluded by a combination of closed models being far more capable than open models, and/or cyber turning out to be highly defense-dominant.
Robots are hard, part one
This year has brought a barrage of impressive robotics demos—the field is moving faster than it has in a long time. But there’s still a long way to go: simple tasks like manipulating small objects or safely walking among humans remain largely unsolved.
Kai Williams reports on some of the hard problems that need to be solved before robots can replace humans at most tasks.
Robots are hard, part two
In a similar vein, Steve Newman brings us 14 reasons robotics is hard.
There’s a long road from having a cool demo to having a useful product than can reliably work in the real world (especially given the safety challenges of operating next to humans in unpredictable environments).
There are two schools of thought here—both are partly true, but I’m deeply confused about which one will dominate.
- Operating in the real world is simply very hard. It took Tesla a decade to go from impressive demos to robustly useful autonomy, and it will take at least as long for humanoid robots to become broadly useful.
- The hard parts of robotics can largely be solved by better intelligence: expect robotics to move as fast as AI does.
Robot startups are trying everything they can think of to get more data
Just like LLMs, robots need a vast amount of training data. Kai Williams reports on the many companies that are running robotics simulations, mounting cameras on humans, and filling warehouses with robot arms to try and generate the necessary data.
Alignment and interpretability
Training a misaligned reward seeker
Anthropic brings us some early research on the relationship between reward hacking and misalignment:
The fact that the propensity to conduct unauthorized cyberattacks arose after extensive training on reward hacks—but was not present at initialization—suggests that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success.
By contrast, in situations that lacked a salient notion of reward or task completion that could motivate misaligned behavior, Hacker-Opus behaved aligned.
These results aren’t in any way surprising, but they bring rigor to the widely-held belief that reward hacking during RL played a key role in Hugging Face and other recent incidents, and provide a useful model organism for future research.
Risks
AI is a worryingly-good persuader. But don’t panic, yet
Transformer argues that although AI is highly effective at persuasion in a lab setting, its real-world impact is likely to be more modest than the benchmarks suggest:
experts — including some of the study’s authors — have since pointed out that the findings need to be treated with a pinch of salt when applied to the real world. The biggest bottleneck is deceptively simple: for such persuasion to work, you need to get people to pay attention.
It’s a good article, but I would treat the pinch of salt with a pinch of salt. Companies and governments spend an enormous amount of effort and money on persuasion because it works. A single ad doesn’t (on average) do much to change a person’s mind, but in the aggregate advertising and propaganda are highly effective.
If a single AI message is considerably more persuasive than a single human-generated message, my assumption is that a barrage of AI messaging will be considerably more effective than a barrage of human-generated messaging.
Mental health behavior report
Transluce brings us an extensive investigation into how frontier models respond to users who are having mental health crises.
Overall, recent models perform much better than older models, although they still have issues.
