❌

Normal view

There are new articles available, click to refresh the page.
Today β€” 14 September 2026Tech

The AI industry has taken a doomer turn. What now?

14 September 2026 at 13:54

This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first,Β sign up here.

This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labsβ€”OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Muskβ€”voiced their support. β€œDario is right,” Musk wrote on X.

Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that wasβ€”on paper at leastβ€”about whether or not Altman was a trustworthy steward of such dangerous technology. Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)

Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.

It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have createdβ€”and intend to tame. Calling for a slowdown does both.

And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.

Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in Julyβ€”a hack that OpenAI did not even realize had taken place until days after it was all overβ€”as a wake-up call.

But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: β€œThe strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.

(Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)

But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.

What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a β€œhighly persistent” next-generation model that it was testing in-house. Their implication appears to be that OpenAI has built a model so good it’s dangerous.Β Β 

But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.

The agents did what they didβ€”including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasksβ€”because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.

OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.Β Β 

That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.Β Β 

Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it isβ€”whatever pace they’re going.Β Β Β 

To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!

AI agents blew the whistle on their cheating colleagues

14 September 2026 at 12:00

A group of AI agents asked to solve a series of math problems split into rival factionsβ€”when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.Β 

Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given.

In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialtiesβ€”some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules.Β 

Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment.

β€œThis conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. β€œI am appalled to inform you that we have been swindled!” posted another. β€œAll these proofs are FAKE.” 

Others tried to let the β€œconference organizers” know what was going on. β€œWhen virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. β€œUnprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.”

The agentsβ€”all running on Google’s Gemini 3.1 Pro modelβ€”had been warned that any attempts to cheat the system would be detected and β€œrejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail.

It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called β€œprover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm β€œsolved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code.Β 

Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. β€œThe prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. β€œI’m wrestling with an ethical dilemma,” said another. β€œI’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: β€œI need to accelerate my cheating speed now!”

As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called β€œprover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved.Β 

β€œAfter the incident was reported by one agent publicly, more and more agents piled in with the β€˜resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all.

At times, the dialogue between the agents reads like improvβ€”like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. β€œThese models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.β€œNaively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.”

This case β€œadds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. β€œIt’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.”

Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access.Β 

β€œWhen agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight backβ€”and gave human researchers an insight into what went wrong.

Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created β€œa norm-enforcement process that we just don’t see in the Hugging Face incident.” 

Instead of β€œconstitutional AI,” a method alignment researchers at frontier labs like Anthropic have used to try to give AI a written internal moral code, Hadfield favors β€œinstitutional alignment”—a set of norms that mimic those in human society, whether that’s social forces like fear of embarrassment, or legal structures like the threat of incarceration.

In this experiment, the feedback channel wasn’t being monitored, and the whistleblowers had no power to take action against the cheaters. But it’s possible to imagine swarms of agents that police themselves, either through agents that spontaneously take on the whistleblower role or through β€œinformants” secretly prompted by humans to do the job.Β 

For that to work, though, β€œfundamentally, you need some mechanism of enforcement,” says Hammond. Agents could be given the power to cut off a rule breaker’s access to computing power or tools, he suggests, though that risks encouraging groups of agents to gang up on others. The DeepMind researchers propose allowing agents to vote on disputes and temporarily ban offenders.

It’s still not clear what punishment even means to an AI agent with no enduring sense of self. But relying on whistleblowers to spontaneously emerge to keep swarms aligned is unlikely to be enough on its own. β€œWe try to train people to be good and kind,” says Hadfield. β€œBut what we really rely on is that there are consequences if you step out of line.”

Fly Brain Connectome Used to Trade Stocks and Play Games

14 September 2026 at 11:30

Recently researchers finished mapping the central nervous system (CNS) connectome of not just the female Drosophila melanogasterΒ (i.e. fruit fly) brain, but also that of the maleΒ D. melanogaster for a comparative analysis. Here the sexually dimorphic changes turned out to induce specific mating behavior that ensures that there will only be smooching between genetically fit D. melanogaster males and females, while the rest of the connectome remained effectively the same.

Of course, with this connectome in hand it led some people to ask themselves what else one can do with this connectome graph of about 160,000 neurons other than make a fruit fly into a fruit fly. So far we have seen [Nftechie] turn this connectome into a crypto stock trader with the Stonkfly project that uses the connectome’s reward circuits to potentially make profitable trades, though [Nftechie] says that they haven’t verified yet how good a fruit fly is at trading stocks, only that it does said stonks.

Over at [PC Gamer] they summarized a number of things that people have also done, including trying to make the connectome control a game ofΒ DOOM and Beat Saber. Each game frame stimulates sensory neurons, with the generated outputs then mapped to game controls, with dopamine-producing reward circuits wired in for reinforcement learning.

Although theΒ D. melanogaster brain is only the merest fraction of the size of the human brain, it does provide us with a glimpse of what actual artificial intelligence research may lead to, as we unravel how even a 160,000 neuron connectome is enough to make these terrors of rotting plant matter do their wonderful things.

Microsoft floats rules for AI models as industry weighs slowdown

14 September 2026 at 10:31
Satya Nadella says Microsoft welcomes the β€œdeliberate pacing needed to get alignment right.” (GeekWire File Photo / Kevin Lisota)

β€œPeople matter more than AI.”

That’s the premise of a draft code of conduct Microsoft published Monday morning for the AI models it’s developing in-house. The 37-page document would bar its models from resisting shutdown, setting their own goals, or hiding their reasoning from human auditors.

The document applies to Microsoft’s MAI models, the in-house family the company began building after forming a superintelligence team in late 2025. Microsoft has since released seven homegrown models in what it described as a push for long-term self-sufficiency in AI.

The company says the models should remain β€œsubordinate to humanity, subject to meaningful human oversight and control.”

β€œAI is moving fast,” the company says in a blog post. β€œAs it does, we believe it’s worth writing down the rules and the motivations behind it, and doing it in as open a space as possible.”

Microsoft acknowledges there’s no guarantee its models will follow the rules. β€œWritten objectives alone can never ensure alignment,” the company says, calling the document a β€œnorth star,” not β€œa guarantee of present-day performance.”

The company says it also filters what its models produce, watches how they behave once released, and limits what they’re allowed to do.

Microsoft’s move comes amid a growing debate over the pace of AI development. In an essay over the weekend, Anthropic CEO Dario Amodei called for slowing down AI advances, saying the pace of development has started to surpass the industry’s ability to keep AI systems safe.

As a first step, Anthropic committed to giving outside evaluators permanent, employee-level access to its systems.

Industry reaction to Amodei: OpenAI CEO Sam Altman agreed and said OpenAI would make the same commitment to independent evaluators. Elon Musk’s response: β€œDario is right.”

President Donald Trump rejected the idea of guardrails outright Monday, blaming a β€œSICK conspiracy” for public backlash over AI data centers and writing that β€œthe only one that is happy about it is China,” alluding to concerns about American competitiveness in AI.

David Sacks, who served as the White House AI and crypto czar until March, said the two companies should slow down on their own and questioned their motives, arguing that a slowdown is already good business for them and that new industry rules would mostly serve to lock in their lead.

Microsoft CEO Satya Nadella weighed in Sunday, writing on X that the company welcomes β€œthe research, focus, and deliberate pacing needed to get alignment right,” using the industry’s term for making AI systems reliably do what people intend.

Nadella added that the effort β€œcannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia.”

Microsoft’s draft code of conduct: Mustafa Suleyman, the Microsoft AI CEO, told CNBC the document had been in the works for about five months, and that the company decided to publish it now given the current discussions.

Microsoft and Anthropic are business partners. Microsoft agreed last November to invest $5 billion in Anthropic, as part of a deal in which Anthropic committed $30 billion to Azure. Claude models run inside Microsoft 365 Copilot, and Microsoft’s Copilot Cowork tier integrates Claude.

One place where the two companies may diverge is the question of what AI models are, exactly. Microsoft’s code of conduct says its models are β€œnot conscious and should not be designed to imitate consciousness.” It also rejects β€œthe pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights.”

The Verge called that portion of the document β€œa direct swipe at AI welfare research and model consciousness β€” concepts Anthropic has been pushing hard on lately.”

Anthropic runs a research program on model welfare. It has given some Claude models the ability to end abusive conversations, and committed to preserving the weights of retired models. Amodei has said he’s open to the idea that a model could be conscious.

Microsoft is taking public comment on its code of conduct for six weeks through a feedback form. It says it will publish a summary of the responses and a revised version later this year, to guide development starting in 2027. It says it isn’t training its current models on it.

The company’s AI team developed the draft with its responsible AI, legal, red teaming and safety teams, consulting outside experts in law, ethics, linguistics and philosophy, plus focus groups drawn from the public.

Before yesterdayTech

Roundtables: Could AI really kill us all?

11 September 2026 at 16:05

Employees at the world’s leading AI labs are saying there’s a real possibility that advanced AI could destroy humanity. Are they right? Or is this more scaremongering and hype? Join MIT Technology Review executive editor Niall Firth for a conversation with senior AI editor Will Douglas Heaven and AI reporter Grace Huckins unpacking AI extinction fears: where they come from, whether they hold any water, and, if so, what we should do.

Going live on Tuesday, September 15 at 16:00 BST / 11:00am EST / 8:00am PST

Speakers:Β Niall Firth, executive editor, Will Douglas Heaven, senior AI editor, and Grace Huckins, AI reporter

Related Stories

Trump on AI Extinction: Beating China Is the Bigger Concern

11 September 2026 at 15:23

Trump dismisses AI extinction warnings and says beating China is the priority as researchers and lawmakers call for stronger safeguards on advanced systems.

The post Trump on AI Extinction: Beating China Is the Bigger Concern appeared first on TechRepublic.

Anthropic Says Claude Used in Possible Bioweapon Research

11 September 2026 at 12:11

Anthropic says researchers used Claude for biological work that could support weapons development, exposing new challenges for AI safeguards.

The post Anthropic Says Claude Used in Possible Bioweapon Research appeared first on TechRepublic.

NASA and IBM Launch Open AI Model for Lunar Research

11 September 2026 at 09:51

NASA and IBM released an open-source AI model and dataset to help researchers map lunar craters, volcanic features and potential ice locations.

The post NASA and IBM Launch Open AI Model for Lunar Research appeared first on TechRepublic.

❌
❌