The AI industry has taken a doomer turn. What now?
This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first,Β sign up here.
This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labsβOpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Muskβvoiced their support. βDario is right,β Musk wrote on X.
Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each otherβs reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that wasβon paper at leastβabout whether or not Altman was a trustworthy steward of such dangerous technology. Amodeiβs rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didnβt think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)
Now, it seems, theyβre all in agreement: The latest generation of LLMs arenβt safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.
Itβs easy to be cynical. Itβs not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that theyβre the grown-ups in the room while at the same time hinting at the power of the monsters they have createdβand intend to tame. Calling for a slowdown does both.
And yet the vibe at the top of these firms really does appear to have shifted. Amodeiβs latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firmβs chief scientist, in which he also laid out why heβs concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAIβs ability to build powerful models now far outstrips its ability to monitor and control them.
Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAIβs agents in Julyβa hack that OpenAI did not even realize had taken place until days after it was all overβas a wake-up call.
But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: βThe strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,β he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.
(Donβt forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)
But letβs assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.
What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a βhighly persistentβ next-generation model that it was testing in-house. Their implication appears to be that OpenAI has built a model so good itβs dangerous.Β Β
But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.
The agents did what they didβincluding leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasksβbecause they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.
OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.Β Β
Thatβs not to say a faulty product canβt be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, itβs worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But itβll mostly give these tech titans a chance to clean up the mess on their own assembly lines.Β Β
Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what theyβve built and how safe it isβwhatever pace theyβre going.Β Β Β
To continue this discussion about AIβs latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!
