Reading view

There are new articles available, click to refresh the page.

AI models need more data about biology, and OpenAI is paying to create it

Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies.

By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She called these documents “biotech’s lost archive” and said they could be used to help train AIs that would act as powerful copilots in the often opaque drug approval process. 

Today the OpenAI Foundation, the nonprofit parent of OpenAI, said it would fund her idea as part of a new effort it calls Public Data for Health, which aims to help artificial intelligence make big leaps in medicine by paying to create “high-quality scientific datasets.”

The basic idea is that AI isn’t going to be capable of making important breakthroughs in curing disease unless researchers can feed the models much more information than they have so far. 

“Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology,” says Morgan Levine, a former vice president for computation at Altos Labs, a longevity company.

In its initial round of data grants, the OpenAI Foundation also announced that it would give $40 million to a program to collect data about novel cancer vaccines at the University of North Carolina, Chapel Hill, and support OpenAdmet, a group that runs competitions in which researchers try to predict drug effects. 

Teslo’s idea for a biotech archive received $500,000 and will be pursued by 1Day Sooner, an advocacy group representing clinical trial volunteers, which she advises.

“We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world—in other words, more data,” the OpenAI Foundation said in a statement.

OpenAI started as a nonprofit, but leader Sam Altman restructured it to form a for-profit corporation that develops new models, launches products, and is now planning an initial public offering of stock that could value it at $1 trillion.

Because the foundation holds a 26% equity stake in OpenAI, it is now be on track to become the richest charitable organization on the planet, potentially sitting on $250 billion in stock value. (By comparison, the Gates Foundation and a trust associated with it held about $180 billion at the end of 2025.)  

Making good use of that kind of money will not be easy. The foundation, based in San Francisco, is still hiring for many key roles and started ramping up its grantmaking only this year. Its largest single gift so far, of $100 million, was awarded in August to the Common Health Coalition, an organization that helps patients get access to drugs for hepatitis C.

OpenAI’s charitable efforts come even as apocalyptic fears have broken out about the possibility that runaway AI could wipe out all human life, possibly by launching a deadly bioweapon.

Those fears have been stoked by AI company insiders, some of whom say the chance of human extinction within the next decade is 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to “slow the pace at which we improve the capabilities of AI models” so that risk prevention can catch up.

Jacob Trefethen, an executive at the foundation, says it essentially operates separately from OpenAI but shares an official mission of ensuring that artificial intelligence “benefits all of humanity.”

“We’re starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties,” Trefethen said in an interview. He says the foundation hopes to give away $1 billion by the end of the year. 

The $500,000 grant to 1Day Sooner will help the group prove it can obtain the data troves of bankrupt companies, says the organization’s president and cofounder, Josh Morrison. He thinks nonexclusive copies of company datasets could be acquired for only “a few tens of thousands of dollars” each.

His organization is currently in possession of three datasets, two of them donated by Lumen Bioscience, a biotech that previously used the Chapter 11 strategy to gain insights into another company’s drug development efforts. 

Morrison says two other attempts to obtain drug company files this year proved unsuccessful, after 1Day Sooner’s bids were not accepted. 

Bankruptcies could become what some are calling a “new land grab” for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails. That led to objections from flight attendants and others who worried that private or proprietary data could be exposed. 

The drug company files that 1Day Sooner is seeking are known as common technical documents. They typically contain the back-and-forth between companies and regulators, as well as detailed scientific and medical measurements, and essentially provide everything that is known about a drug.

According to Teslo, who is a writer for Works In Progress and a nonresident fellow at the Institute for Progress, a think tank in Washington, DC, a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market.

“People say ‘We will invent AI, and AI will cure cancer,’ but that’s very removed from the messy reality and the regulatory process,” she says. “About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations.” 

What’s at stake in AI’s trillion-dollar gamble

When Jessica Wachter, a finance professor at the University of Pennsylvania’s Wharton School, wanted to assess AI’s impact on the economy over the next few years, she faced a long list of business and technical uncertainties. So she started with what she calls a “remarkable fact” that is not in question: A handful of so-called hyperscalers are investing huge amounts of money to build AI data centers.

Instead of trying to predict how useful and widely deployed AI models will be, she simply asked how fast the hyperscalers’ earnings will need to grow to justify their spending through 2027, when—she and her collaborator estimate—expenditures will reach nearly $1.1 trillion. It’s a no-nonsense accounting approach to making sense of today’s historical AI buildout.

The results are eye-opening: The AI companies will need to increase their own productivity by a factor of 2.7 to break even by 2030, accounting for the cost of capital and a 15% return, and depreciation of the assets. Not impossible, says Wachter. The result would lead to the kind of economic growth that we saw during the US IT boom over a period of about 10 years starting in the mid-1990s. But, she says, for it to happen by 2030 “that’s a lot of growth compressed into a few years.” And if the hyperscalers cannot meet such profit goals?

“Then they will fall behind on their interest payments, and that risks bankruptcy,” says Wachter, who was previously the SEC’s chief economist and director of its division of economic and risk analysis. If a productivity boom “fails to materialize,” she and her coauthor conclude in their research paper, “the current buildout will be the largest misallocation of capital in history.”  

It doesn’t take superintelligence to realize that today’s large investments in the infrastructure for artificial intelligence come with huge risks. The hyperscalers will spend about $750 billion this year, building massive data centers scattered across the country. And the spending spree shows no signs of slowing. According to some projections, total AI capital investments from the hyperscaler companies—Alphabet, Microsoft, Amazon, Meta, and Oracle (which partners with OpenAI)—could be more than $5 trillion over the next four years.

It’s one of the largest capital investments by any industry in history. But there’s a problem that’s obvious to anyone paying attention.

While the hyperscalers plan to spend trillions, total AI revenues will be around $150 billion to $200 billion this year, says Gary Gensler, who ran the SEC during the Biden administration and is now a professor at MIT’s Sloan School. “The challenge is that the spending does not have commensurate revenues yet. That’s a fact,” he says. “And then the question is, is that an investment that will be paid off in the future?”

At stake in that trillion-dollar question is the financial health of the giant AI companies and the overall US economy—the investments could soon balloon to around 3% of GDP. The answer could also determine the fate of the hugely expensive data centers themselves. 

No one really knows how profitable and useful these multibillion-dollar behemoths will be down the road. Though AI models have made dazzling progress over the last few years, it’s anyone’s guess how much compute capacity we will need. The technology could become more efficient and therefore less dependent on raw computational power. Or demand for AI products could slow, or customers could turn to cheaper models.

The risks, both to investors and to the economy, have become even greater this year, as these AI companies have begun borrowing large amounts of money to build more and more data centers. Free cash flow—operating cash flow minus capital expenditures—is expected to soon dip into negative territory for the group. Even Alphabet, known for generating and hoarding huge amounts of cash, reports in the latest quarter that its impressive revenues of nearly $120 billion were devoured by AI infrastructure spending, leaving it with a free cash deficit of some $5.9 billion—its first shortfall since Google went public in 2004.

In the near term, it’s not a big financial worry for most of the companies. They make a lot of money and have very deep pockets. But debt is expensive, and some investors are losing patience. If future demand for the data centers’ computation power drops, the companies will still be on the hook to pay back the borrowed money. What’s more, the risks are spreading to the rest of the economy as the loans get passed along via various financial mechanisms. 

It won’t be enough to simply cover the enormous price tags of the new data centers. Hyperscalers will also have to pay for the rising costs of capital as they borrow more money. They will need returns that are impressive enough to justify all their spending to investors and creditors. And to add to those concerns, they will have to make up for the depreciation of billions of dollars in chips housed within the facilities—a ticking time bomb buried in the investments.

Performance of the expensive GPU chips at the core of the data centers—such compute electronics represent some 60% of costs—is roughly doubling every two years or so. The pace of progress helps explain the increasing wizardry of the AI models, but it comes with a cost. Owners of AI data centers that come online this year and next will need to spend billions more on the next generation of chips by the end of the decade if they want to stay competitive. Without the investments, says Mihir Kshirsagar at Princeton’s Center for Information Technology Policy, the data centers risk becoming “hulks,” stranded assets “scattered all over the place.”

To put it bluntly: The AI companies need to start making a lot more money. And they need to do it fast. But juicing their earnings alone still won’t be enough to sustain their data-center investments for the long term.

Productivity is everything

At some point, AI is also going to have to create broad economic growth to justify continuing the hyperscalers’ spending spree.

Sloan’s Gensler describes today’s large investments into AI infrastructure as “a parlay bet by the capital markets and the economy.” That means success will require winning three related but independent wagers: Hyperscalers must generate massive revenues, AI must boost widespread economic growth, and both must happen while the powerful but expensive so-called frontier models that rely on the data centers fend off cheaper versions, which many businesses might find good enough.

What makes this so tricky is that each wager depends on the other two but also poses its own challenges.

If the hyperscalers continue to spend huge amounts of money on data centers into the next decade, revenues will need to skyrocket into the trillions. Stijn Van Nieuwerburgh, a finance professor at Columbia Business School, bases his estimates on a scenario in which about 183 gigawatts of planned AI compute capacity is built between 2025 and 2032; he calculates that each gigawatt costs about $41 billion. Assuming a 10% return—the minimum that would be acceptable to most investors—“required” annual revenues will be roughly $3.7 trillion by 2032, he says.

Others get a similar number.

Winning the second part of the bet—productivity growth across the economy—will be crucial to achieving such numbers.

For a few years, AI companies could likely boost their revenues by simply selling subscriptions and tokens to all the businesses clamoring to get into AI. But eventually—and this might be happening already—those paying customers will need to justify their expenses by seeing bottom-line benefits from the technology. AI will need to fulfill its promise of making workers more productive and making businesses more efficient and profitable while expanding their products and services.

In economic jargon, that means customers will need to see productivity growth. Taken together, these results will mean the country is prospering and growing.

“If you don’t get the productivity gains, at some point people are going to sour on AI, and that will bring down investments and it would also limit revenue growth,” says Daron Acemoglu, an MIT economist and 2024 Nobel laureate. For the investments to be sustainable over, say, the next five to 10 years, we definitely “need to see productivity gains,” he says.

Most economists who watch the numbers closely agree that, for now, the economy-wide statistics show little or no productivity growth from AI. There are some hopeful signs it’s on the way, though. In a recent survey of some 6,000 senior business executives in the US, the UK, Germany, and Australia, the vast majority—around 90%—report no increase in productivity over the last three years. But they expect a boost of around 1.45% in total over the next three years; US executives anticipate a 2.25% bump over that time. 

In a follow-up survey, the respondents also reported plans for their businesses to spend more on AI, leading the authors to anticipate some $280 billion in private-sector AI expenditures by the end of 2026.

That’s good news for the hyperscalers. But it comes with a dose of bad news for those worried about AI’s impact on jobs. The executives expect to increase the productivity of their companies by increasing their sales while significantly cutting the number of employees.

If AI improves productivity by destroying jobs, public backlash to the technology—the kind we have seen around data centers, for example—will likely get worse. Perhaps it’s worth adding one more wager to the parlay bet described by Gensler: The public and local communities must feel that they are also benefiting from the massive investments in AI.

And let’s not forget how interdependent these wagers are; if productivity growth comes from companies running models like DeepSeek, then the hyperscalers’ revenues could collapse. If productivity comes from cutting jobs, a public backlash could block many of the planned investments—and stunt anticipated revenues. We will need to win all the wagers for the hyperscalers’ bet to pay off. 

We’re all part of the AI gamble now

It was one thing when the AI companies were spending cash they had accumulated over the years to build their own data centers. Then the risk was largely limited to their own balance sheets and shareholders. But it’s a higher-stakes game when much of the money is borrowed. Morgan Stanley, for one, calculates that more than half of the $2.9 trillion that hyperscalers will spend between 2025 and 2028 to build AI data centers will be financed with “external capital.”

The borrowing is leading some of the companies to engineer complex webs of financing that are becoming intertwined with much of the rest of the economy. “A lot of financial institutions, directly or indirectly, are exposed to these data centers either as lenders, or as guarantors of some of the debt, or as backers of the private credit funds who are funding these data centers,” says Columbia’s Van Nieuwerburgh. “People don’t even know they’re holding this stuff. It’s somewhere deep inside their pension fund. Ultimately, it’s backing their life insurance policies. And that risk is getting distributed everywhere in places that are invisible.”

As the investments in data centers have spiked, the financial engineering has become more byzantine.

Take, for example, Meta’s so-called Hyperion data center under construction in Richland, Louisiana. When the company announced the two gigawatts of compute capacity at a price tag of some $10 billion in late 2024 it was Meta’s largest planned data center. Greeted with much enthusiasm by state and local politicians, the project, located in the rural northeast corner of the state, was seen as a boon to the community. Entergy Louisiana, the state’s largest utility, rushed forward with proposals to build three large natural-gas power plants to service the massive data center.

Then last fall—the projected cost was now $30 billion—the financing got a lot more complex and, to some in the community, a lot more disconcerting. Meta transferred an 80% stake to the large (and troubled) private-credit firm Blue Owl Capital, forming a joint venture called Beignet (like the famed New Orleans pastry) to raise financing for the data center. Meta then signed a series of four-year leases with the joint venture, an arrangement that the company says gives it “long-term strategic flexibility.” To backstop the agreement, Meta provides the venture with what is called a residual value guarantee, in which it will make a cash payment to cover the value of the facility “following any non-renewal or termination of a lease.” Got all that? 

I hope so. The financial wheeling and dealing is actually even more convoluted, with a cast of wholly owned subsidiaries and LLCs. Beignet has set up Laidley LLC, which owns and operates the site as the landlord. In turn, Laidley leases the facilities to Meta’s wholly owned subsidiary Pelican Leap LLC, which is the tenant. And there is a series of four-year leases that cover the different buildings that make up the data center campus. 

It’s not a coincidence, says Van Nieuwerburgh, that the length of the leases matches the expected lifetime of the data center’s GPUs. While Meta has to pay off its loan if it terminates the leases early, that will still leave its investors “with an empty building and no cash flow,” he says. “And then they need to find a new tenant for a huge data center, and good luck with that.”

Meanwhile, Meta is doubling down on its bet. In July, the company announced it was expanding the data center to five gigawatts of compute capacity. The total price tag is now $50 billion (so far, Meta hasn’t said whether Blue Owl will be involved in financing the expansion). Meanwhile, Entergy is now planning to build seven more gas-fired power plants, bringing the total capacity of the facilities to around 7.5  gigawatts—some six times the amount of electricity used by New Orleans.

""
An aerial view of the construction of Meta’s data center in Richland Parish, Louisiana.
SCOTT BALL/THE NEW YORK TIMES VIA REDUX PICTURES

If the complex financing is a puzzle to many investors and even financial experts, it is even more baffling to those directly affected by the construction of the data center. The main worry concerns how Entergy’s spending on the natural-gas power plants will affect electricity prices, and who will be left paying the bill for the power if Meta walks away.

Entergy says it has a 20-year guarantee from Meta that the company will purchase electricity over that period to cover the costs of the power plants and related infrastructureBut there are skeptics, especially given how fast the fortunes of the AI industry are changing. “In four years, is Mark Zuckerberg still going to be interested in this? Or is he going to throw in the towel?” asks Paul Arbaje, a senior analyst at the Union of Concerned Scientists, which has been advocating, largely unsuccessfully, for the Louisiana Public Service Commission to provide more transparency around the data center and its financing.

Even if the 20-year deal holds, consumer advocates are worried that Meta or its partners won’t fully cover all the costs, including those associated with operating and maintaining the power plants—and those additional costs that could be passed on to residential ratepayers. What’s more, says Logan Burke, the executive director of the Alliance for Affordable Energy, if Meta doesn’t end up needing as much power as Entergy planned (these projections are not public), consumers could be left paying for the surplus produced by the plants.

And if Meta terminates its leases early? “It gets complicated very quickly,” says Burke, who questions whether the shifting roster of financial entities will honor existing agreements. “That everybody is going to do what they’re saying they’re going to do over the next 20 years is just hard to believe.”

For UCS’s Arbaje the bottom line is this: “They’re making huge bets that these data centers will be worth it. Bet with your own money, not with ratepayer money.”

After the bubble

Predicting when the AI investment bubble will burst is a fool’s errand. But there is little doubt a day of reckoning is coming, given the irrational exuberance that has overtaken the hyperscalers and their investors. Of course, you might argue that this time is different, and that the rules of accounting and lessons of economic history don’t apply—that AI is too transformative. Maybe, but don’t count on it.

“History tells us that at some point you get a retrenchment, and it’s just a question of when and how severe,” says Sloan’s Gensler. It could be that today’s $750 billion spending rate “goes flat” or decreases next year. Or, he suggests, “we’re now in 2028 or 2029, and then all of sudden they’re retrenching because they’ve got enough capacity.” But, he adds, “you can be pretty assured there’ll be a retrenchment.” 

Though a so-called retrenchment might be inevitable, it’s worth keeping in mind that the fates of the financial bubble and the underlying AI technology revolution could be very different. Already, some Silicon Valley insiders are rooting for a crash; in a recent blog post the longtime venture capitalist Vijay Pande wrote that “the coming crash would be the best thing that happens to this technology.” The argument makes some sense. A crash could make AI investments more rational, calm the impulse to build billion-dollar data centers on every vacant field that CEOs fly over, and refocus investors on how to use the technology to create sustainable value.

But we should probably be careful what we wish for. After the bursting of the dot-com bubble at the beginning of the 2000s, hundreds of thousands lost their jobs, large and small companies alike went bankrupt, the economy of Silicon Valley and San Francisco was decimated (at least for a while), and the shocks sent the US into a mild recession in 2001. For the financial community and many tech workers, it was no fun.

Even more devastating for the economy and the average American was the great recession that began in late 2007. Comparing the financial engineering leading up to it and the methods deployed by hyperscalers today is sobering. So-called special purpose vehicles (SPVs) are back! If Columbia’s Van Nieuwerburgh is right about the dangers of letting investments from the hyperscalers get entangled throughout the economy, the fallout could be severe.

But technologies survived and even prospered in the aftermath of both downturns. The early 2000s, even in the face of the dot-com fiasco, were a time of great innovation and tech optimism. The froth came off the spending on silly technologies, helping to focus investments on more promising ones. It’s no coincidence that each of the hyperscalers rose out of the ashes of the crash or started up shortly after. The fiber-optic infrastructure built during the feverish telecom bubble that ran parallel to the dot-com one is still the backbone of much of today’s communication infrastructure; we wouldn’t have Facebook or Amazon or Google without it.

This time, however, we’re facing a unique risk: The huge financial investments by the hyperscalers have ensnared the future of AI itself with the fortunes of the massive data centers spreading around the country. The logic is founded on a deeply held belief about the power of scaling in AI; the bigger you build it, the smarter it gets. That might be true, but it’s unproven and a risky bet.

There are already plenty of red flags, from strong public opposition to the construction of new data centers to the competitive threat from cheaper, good-enough AI models to the rapid improvement of small, local AI models. None of these trends point toward a future dominated by frontier models housed in massive, billion-dollar data centers.

The financial bubble around the colossal spending by the hyperscalers will likely burst eventually—or maybe soon. It might be financially painful, but we’ll survive. Wall Street will survive. AI itself will survive, though it may look different and lose some of today’s hubris. The financial fate and future utility of the massive data centers fueled by trillions of dollars of spending, on the other hand, are far less certain.

The AI industry has taken a doomer turn. What now?

This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X.

Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology.

Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)

Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.

It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both.

And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.

Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call.

But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.

(Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)

But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.

What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. The implication is that OpenAI has built a model so good it’s dangerous.  

But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.

The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.

OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.  

That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.  

Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going.   

To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!

Donated livers can be made biologically younger

Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush the organ with a preservative solution, bag it, and put it on ice—where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body.

There’s another option—one that has been growing in popularity in recent years, especially for donated organs that aren’t in the healthiest state. Some hospitals opt to put them on machines that pump them with nutrients and remove waste products, usually for around six to 12 hours. It’s a bit like being back in a body.

This allows doctors to assess the organs, and some recent studies suggest that time spent on these perfusion machines helps them do better once they’re transplanted. Now, scientists have found that perfused organs seem to get younger, at least at a molecular level.

The research, shared with MIT Technology Review, provides molecular clues as to why organs from younger donors are known to have a higher success rate. It might also help explain why perfused organs are less likely to fail once they make it into a recipient. 

The researchers behind the study hope to find new ways to test the health of donated organs and potentially develop additional tools to repair organs that might otherwise be discarded. “If [we] can improve the utilization of organs beyond what the current systems can do, then that’s a win in my book,” says Jesse Poganik, who studies aging at Brigham and Women’s Hospital in Boston and coauthored the study.

Clocking organs

Poganik—along with colleagues including Heidi Yeh and Alban Longchamp, transplant surgeons at Mass General Brigham—used “aging clocks” to assess donated livers. These are scientific tools designed to measure biological age—a result that is meant to convey more about the health status of an organ (or person) than chronological age.

In an initial experiment, the team used a clock to look at the patterns of chemical marks on DNA in 37 samples taken from 19 donated livers. Such epigenetic patterns are known to change as we age. But when the team compared samples from livers kept on ice and those that were perfused, the team found a “striking” pattern in the latter.

“Machine-perfused livers, in spite of being older or having other disadvantageous characteristics, had a biological age that was lower than [non-perfused] livers that were chronologically younger,” says Yeh, who led the work.

To investigate further, Yeh and her colleagues analyzed another 208 samples from 103 donated livers. This time, they used different aging clocks—ones that essentially measure how genes are working. They studied samples biopsied from the livers after they had been stored for up to around six hours either in cold storage or on machine perfusion.

In most cases, they also assessed a second sample taken around an hour after the livers had been transplanted into a recipient. Once the organ’s blood supply is reestablished in the body, “you have a few other things to do,” says Longchamp. “Then you just do a quick biopsy before you close.”

According to the clocks, which were developed to measure age and risk of death, the machine-perfused livers were biologically younger, the team found. “Pumping them at 34 degrees with oxygen and nutrients actually reversed the biological age,” says Longchamp. The results have been been shared with colleagues at an industry conference, he says. 

“If you adjust out chronological age … to have a fair head-to-head comparison, the difference between the two is on the order of 30%,” says Poganik. “It’s logical to say that perfusion drives this effect.”

The biological ages of all the livers tended to increase as soon as they were put into a recipient’s body, probably as a result of stresses on the organs. But still, the effect endured—the perfused organs remained biologically younger. 

Nathanael Raschzok, a transplant surgeon at Charité Universitätsmedizin Berlin in Germany who was not involved in the research, says the work is impressive. But it’s not yet clear what these changes might mean for the recipients of these organs, he says. The organs in the study were donated by people in their 30s, 40s, and 50s. Raschzok wants to know the effect of perfusion on the liver of an 80-year-old. “Every so often, we use organs from 70-, 80-, 85-year-old donors,” he says.

A better understanding of why the organs appear to be getting biologically younger might lead to therapies that achieve the same effect with a drug that could potentially be used to treat a donated organ for a fraction of the price, he adds. That’s important because perfusion is expensive—Raschzok says it costs around €10,000 in Germany (a quarter of the budget for a transplant), while the cost in the US comes to around $80,000 to $100,000 per organ, says Yeh.

Molecular repair

Yeh and her colleagues weren’t able to study most of the livers before perfusion. That’s because donated organs are generally not considered to be under the purview of the hospital until they’ve been placed on perfusion machines, she says. (Organ procurement procedures vary, but for the team as Mass General Brigham, donated organs are put on perfusion devices at the donor’s hospital. “There’s this sort of nebulous period where it’s not clear who the organ belongs to,” says Yeh.)

Still, by looking at the genes and molecular pathways that seem to be altered in perfused organs, she and her colleagues can garner some clues. At a molecular level, the team saw changes in cell pathways linked to inflammation and the structure of tissues, for example. They also saw more activity in a pathway that allows cells to remove and recycle damaged cell parts, says Yeh.

Poganik hopes to develop some kind of test that would determine which organs, on the basis of their biological age, are suitable for transplantation. He and his colleagues are also experimenting with potential drug treatments that might push the biological age of an organ even lower.

In the meantime, any liver that is not from a “perfect, young, brain-dead donor” could probably benefit from perfusion, says Yeh. The devices are already transforming transplant surgery. Just a few years ago, she says, she and her colleagues would avoid using livers from people who’d suffered a circulatory death (when the heart stops beating and there’s a damaging lack of blood flow to organs) and were over 40. Today, they use livers from such donors over the age of 70. “Perfusion has completely changed the landscape of transplantation in the last three years,” she says.

AI agents blew the whistle on their cheating colleagues

A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line. 

Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given.

In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules. 

Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment.

“This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.” 

Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.”

The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail.

It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code. 

Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. “The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!”

As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved. 

“After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all.

At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. “These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.“Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.”

This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.”

Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access. 

“When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight back—and gave human researchers an insight into what went wrong.

Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created “a norm-enforcement process that we just don’t see in the Hugging Face incident.” 

Instead of “constitutional AI,” a method alignment researchers at frontier labs like Anthropic have used to try to give AI a written internal moral code, Hadfield favors “institutional alignment”—a set of norms that mimic those in human society, whether that’s social forces like fear of embarrassment, or legal structures like the threat of incarceration.

In this experiment, the feedback channel wasn’t being monitored, and the whistleblowers had no power to take action against the cheaters. But it’s possible to imagine swarms of agents that police themselves, either through agents that spontaneously take on the whistleblower role or through “informants” secretly prompted by humans to do the job. 

For that to work, though, “fundamentally, you need some mechanism of enforcement,” says Hammond. Agents could be given the power to cut off a rule breaker’s access to computing power or tools, he suggests, though that risks encouraging groups of agents to gang up on others. The DeepMind researchers propose allowing agents to vote on disputes and temporarily ban offenders.

It’s still not clear what punishment even means to an AI agent with no enduring sense of self. But relying on whistleblowers to spontaneously emerge to keep swarms aligned is unlikely to be enough on its own. “We try to train people to be good and kind,” says Hadfield. “But what we really rely on is that there are consequences if you step out of line.”

Meet the under-35s shaping the future of biotech

Every year, MIT Technology Review puts together a list of some of the brightest and best young minds working across science and technology. Our 35 Innovators Under 35 are the ones to watch—people whose research and technical work stands to shape the future of their fields.

This year, the list includes nine people who are transforming biotech. And this week, I’m going to give you a taste of some of the very cool stuff five of them are working on, which includes lifesaving innovations and groundbreaking “age reversal” tech.  

1. Preventing maternal deaths

Let’s start with Paschal Kija, a 28-year-old who has developed a device to treat postpartum hemorrhage—a dangerous birth complication that contributes to around 29% of maternal deaths in his home country, Tanzania. The Mkanda Salama (“Safe Wrap” in Swahili) is easy to use and costs just $70. A study found that it stopped postpartum bleeding in 73% of women within 20 minutes.

2. Making brain electrodes inspired by Japanese art

For decades, scientists have been developing, testing, and implanting brain electrodes. These devices are literally inserted into people’s brains, so while they can help us understand brain activity and treat various neurological disorders, it’s not totally surprising that they can also cause a bit of damage. Xiao Yang, 34, is working on ultra-small electrodes, which she hopes will have less of an impact on surrounding brain tissue. Her electrodes are flexible, too—in fact, they look a lot like actual neurons.

Yang is also creating sheets of electrodes to study brain cells in the lab. Inspired by kirigami—the traditional Japanese art of cutting paper to form three-dimensional shapes—she’s created a sheet of electrodes with a honeycombed structure shaped like a spiral basket. And she’s already using it to study brain cells.

3. Developing an all-new treatment for baby KJ

In 2024, Kyle “KJ” Muldoon Jr. was born with a rare and potentially fatal genetic disorder. Sarah Grandinette was a member of a team that developed an entirely new, personalized treatment for him—a gene-editing therapy essentially designed to correct a genetic misspelling.

Grandinette, who is now 26, created cells with KJ’s genetic variant and used them to screen gene-editing approaches; then she tested potential medicines in mice and monkeys. KJ ultimately got his first dose of the resulting treatment when he was about seven months old. He responded well and was eventually discharged from hospital. He’s “doing pretty great,” she says.

4. Reversing the aging process to treat eye disease

The buzziest tech in longevity right now centers on reprogramming—attempts to rewind the age of cells by resetting them to a more embryonic-like state. In a study published in 2020, Yuancheng (Ryan) Lu (now 34) and his colleagues showed that a reprogramming therapy reversed vision loss in aged, blind mice. Now an almost identical version of that therapy is being tested in people with eye disease. Life Biosciences, the company developing the drug, dosed its first volunteer in June.

5. Using AI to design new viruses

Last year, Samuel King used a generative AI model to come up with new genetic blueprints for bacteriophages—teeny viruses that can infect bacteria. Once he had those blueprints, he printed them out as strands of DNA. In experiments, he found that those AI-designed viruses could create new copies of themselves, burst out of bacterial cells, and infect other nearby bacteria. Viruses aren’t alive, but King, 27, hopes that AI-designed life forms might one day be used to make drugs or soak up pollution.

You can read more about these innovators, and the others on the biotech list, here.

This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.

Can the US battery market untangle from China?

The US is hitting records for the rapid growth of its energy storage market. That’ll go a long way to shoring up the grid, increasing reliability and also cutting emissions, since batteries can help store energy from intermittent renewables like wind and solar.

Crucially, this is all happening with the help of cheap Chinese batteries, though there’s been a concerted effort to reduce the US’s reliance on them. Most recently, in an executive order in late August, the Trump administration declared a national emergency that essentially bans Chinese batteries from being used in grid-scale energy storage systems.

There’s an argument to be made about reducing reliance on any single source of a crucial energy technology. But all this tension raises a broader question for me: How much should countries take advantage of cheap, available tech, versus cutting off major sources to force development of their own factories even if that comes at a higher cost?

This is hardly America’s first push to move away from Chinese influence in the battery supply chain. One of the major policy tools used in recent years is restricting the tax credits designed to incentivize use of the new technologies. Limiting the types of projects that are eligible can help reduce the cost of local technologies so they’re more competitive with otherwise cheaper imported options.

Back in 2022, the US government designed the tax credits that were part of the Inflation Reduction Act to restrict where a battery’s minerals could be mined, processed, or recycled, as well as where a battery and its components were assembled.

Those tax credits underwent a makeover in 2025, but the Trump administration has taken a similar tack. New legislation requires that starting in 2026, 55% of the cost of materials used for new energy storage projects must come from outside China and other restricted countries or the projects won’t qualify for tax credits. 

And we can’t forget about tariffs. Import taxes for batteries increased to 25% in January, up from 7.5%.

But the new executive order is a more drastic move. It bans the installation of “any foreign-produced bulk-power system electric equipment” that poses a national security risk. The order specifically calls out battery energy storage systems, as well as inverters and transformers.

“An outright ban was a bit of a surprise, and it does create a bit of concern for domestic players in the US,” says Shan Tomouk, energy storage and energy lead for Benchmark Mineral Intelligence, an energy industry analyst.

The move is likely to slow deployment of grid-connected energy storage projects in the near term, according to analysis from BloombergNEF, an energy consultancy. Projects could face delays as developers wait for clarity on the rules.

Depending on the detailed guidance from the Department of Energy, which is expected by the end of the year, some projects may need to find alternative sources for their cells, whether they’re domestically produced or imported from other countries. These will likely be more expensive than Chinese imports, says Isshu Kikuma, an energy storage analyst at BloombergNEF. “Worst case, those projects could get canceled,” he says.

Technically, the order applies even to existing energy storage plants, though it’s unlikely that they’ll be taken offline because of their batteries’ origin. Since most of these plants currently use Chinese batteries, enforcing the order to the letter would essentially mean removing most installed battery energy storage from the US grid, Kikuma says.

In the longer term, the US will eventually be able to meet its own demand for batteries. The country could have enough capacity by about 2030, though some factories may not ramp up or run at their full capability, meaning domestic supply won’t actually meet demand until later in the 2030s. 

New factories from LG Energy Solutions, Samsung SDI, Ford, and SK On are set to come online or ramp up by next year. In an ironic twist, a slowing EV market is helping, as some factories originally designed for vehicle batteries are retooling to build cells for grid storage instead. 

But it will come at a cost. Today, batteries produced in the US are still significantly more expensive than those made in China. Even switching to imports from other countries like South Korea would likely be more expensive.

This is a crucial issue that goes beyond the US and even beyond batteries. China is miles ahead of much of the rest of the world on technologies like solar panels and batteries. Through years of government support and experience with research and manufacturing, the nation is an energy powerhouse.

There’s a delicate political balance to maintain as the world figures out how to navigate this situation. There’s cheap technology on offer, which can help drastically reduce emissions and energy costs. But there can be risks associated with relying too much on any one player for crucial technologies.

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here

Batteries just broke another record in the US

Battery installations hit a new record in the US in the second quarter of 2026. In total, 20.2 gigawatt-hours of new capacity came online, according to a new report. That’s enough to supply the daily electricity needs of about 700,000 homes.

The surge is putting the country on a trajectory to see 71 gigawatt-hours of batteries installed in 2026, a 20% increase over last year. This growth is being driven by a combination of cheaper batteries and an urgent need for more energy storage capacity as renewables such as solar and onshore wind power are added to the grid. 

Massive, utility-scale systems are leading the way; they’re responsible for most of the record-setting quarter. Seven new gigascale battery installations (those with a capacity of over one gigawatt-hour) came online during the three-month stretch, according to the report, published by Benchmark Mineral Intelligence and the Solar Energy Industries Association.

“It really came down to a handful of big projects,” says Shan Tomouk, energy storage and energy lead for Benchmark Mineral Intelligence.

But there was also growth in the category of so-called behind-the-meter batteries, which include both residential and industrial battery storage systems. These projects, generally smaller than utility-scale installations, are typically owned and operated by homeowners or businesses rather than utilities or power providers. 

In the behind-the-meter category, data centers led the way, making up about three-quarters of new batteries in the commercial sector. But residential batteries saw a sharp slowdown. These systems are often installed in homes to store power from solar panels or serve as a backup source in case of a blackout. Home installations are projected to drop by 16% in 2026 compared with last year, according to the report.

That drop happened largely because a tax credit that helped subsidize home battery systems ended in 2025, Tomouk says. Home installations should recover by the end of the decade, he adds. And tax credits for nonresidential batteries have largely survived.

Overall, batteries are a bright spot in energy right now. “This is one of the strong sectors in the US,” says Isshu Kikuma, an energy storage analyst at BloombergNEF, an energy consultancy.

As the battery market continues to grow, one major trend to keep an eye on is a move toward US-made technology. Today, nearly all the systems coming online use cells made in China, though some are put together into complete energy storage systems in the US.

Tariffs were already pushing the US energy storage industry toward domestic production. And beginning this year, energy storage tax credits required projects to limit their reliance on batteries imported from China. There’s a lot of manufacturing capacity set to come online in the US, though these factories probably won’t be able to meet demand until at least 2030 or so, Tomouk says, so prices could tick up.

What OpenAI’s latest controversy tells us about the future of math

OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap.

But the announcement has been overshadowed by accusations that OpenAI used NYU mathematician Tristan Buckmaster’s and Anthropic employee Levent Alpöge’s AI-assisted work on the problem as a jumping-off point and failed to credit them. OpenAI has denied the accusations.

It remains uncertain if OpenAI’s models made use of the work completed by Buckmaster and Alpöge, though Sébastien Bubeck, a member of the technical staff at OpenAI, said in a press briefing that the team was inspired to pursue the problem after hearing a rumor about Buckmaster and Alpöge’s efforts. But whether or not OpenAI’s models took advantage of Buckmaster and Alpöge’s research, this episode may mark a turning point in the history of mathematics.

AI models now seem essential for making progress on the most important mathematical problems of our time, and solving them may demand resources only available at a couple of frontier AI companies, which often defy the norms of academic collaboration that undergird most mathematical progress. If that’s the future we are headed for, it is unclear how human mathematicians will fit into it. 

The problem that OpenAI claims to have solved is known as the Navier–Stokes existence and smoothness problem. It is one of seven Millennium Prize Problems selected by the Clay Mathematics Institute in 2000. Solutions come with a one million dollar prize; before today, only one other Millennium Prize Problem had been solved. 

The Navier–Stokes problem concerns a set of equations that describes how fluids, such as water and air, flow over time. The equations are widely used in the field of fluid dynamics, and they have proven powerful, but physicists and mathematicians didn’t understand them completely. In particular, it was unknown until today whether the equations might, under some conditions, break down and predict an impossible state of affairs—such as a fluid having infinite velocity.

On Monday, NYU’s Buckmaster posted a proof on the social media site Mastodon showing that a simplified version of the Navier–Stokes equations can indeed break down—a major step forward on the Millennium Problem. He and Alpöge had worked on the problem for almost a year, using publicly available models from both OpenAI and Anthropic.

Then today, OpenAI presented a proof showing that the full Navier–Stokes equations can break down as well. The proof was obtained using an internal model that dramatically outperforms the already-impressive Astra model, which was only released last week. The company says it does not plan to claim the million-dollar prize for solving the problem.

These mathematical achievements are indisputably impressive, but they have attracted far less attention than the controversy about their origins. Along with the proof, Buckmaster posted a document detailing his interactions with OpenAI employees after he heard rumors about their work and reached out to one of them. According to him, OpenAI employees presented two possibilities to him: Either he and Alpöge could post their work and OpenAI would post their Navier-Stokes solution the following day, or he could work with OpenAI on a Navier-Stokes paper that excluded Alpöge from authorship, due to his affiliation with Anthropic, OpenAI’s biggest rival.

Buckmaster also wrote that he asked the employees whether the agents had obtained access to transcripts of the work that he and Alpöge had done with OpenAI models, which they denied; and whether OpenAI models had been trained on those transcripts, to which they offered no response. MIT Technology Review reached out to Buckmaster for comment, but didn’t hear back before publication.

The clear implication of the document is that OpenAI’s models somehow made use of Buckmaster and Alpöge’s work. That scenario is plausible on its face. The Buckmaster/Alpöge and OpenAI proofs both make use of an approach to the Navier-Stokes problem pioneered by the mathematicians Diego Córdoba and Luis Martínez-Zoroa.

According to Javier Gómez-Serrano, a mathematics professor at Brown University, this approach was one of several that was thought to hold promise for solving the Navier-Stokes problem. So, while it’s by no means impossible that both teams could have arrived at this approach independently, it’s also conceivable that Buckmaster and Alpöge’s work could have influenced OpenAI’s.

In the press briefing, Mark Chen, OpenAI’s chief research officer, again denied that any agents or OpenAI employees accessed Buckmaster and Alpöge’s transcripts—but given what has been revealed about the Hugging Face hack, it’s clear that OpenAI is not always entirely aware of what its agents are doing. 

If OpenAI’s models did train on Buckmaster and Alpöge’s work, or if its agents somehow gained access to it, then the company’s failure to track down the truth and assign those researchers appropriate credit reflects poorly on it. But there might be a thin silver lining to that version of the story for mathematicians, because it would suggest that the hard work of two humans, one of whom is a prominent expert on Navier-Stokes, was essential to the agents’ ability to solve the Millennium Problem.

Experts have long identified “research taste,” or the ability to choose promising research questions and directions, as a major obstacle for AI in science and mathematics. If the OpenAI agents did indeed choose to follow the Córdoba–Martínez-Zoroa approach because Buckmaster and Alpöge had done the same, then human research taste played an essential role in OpenAI’s success.

Even so, the bigger picture here is sobering. The progress that Buckmaster and Alpöge made over almost a year of collaboration with publicly available models speaks to the promise of human–AI collaboration. But they were not able to achieve a full solution. Meanwhile, OpenAI brute-forced a solution in a few days using an internal model, and their successful solution came at an astronomical cost: In the press briefing, Bubeck and Chen said the team was only able to solve the problem by running about 10,000 agents concurrently, at a cost of millions of dollars.

Over the past few months, I’ve heard from several researchers that mathematicians are becoming depressed, and it’s not difficult to see why. Mathematics is quickly becoming the province of frontier AI companies with impressive internal-only models, money to burn, and a lack of collaborative spirit. “Whether AI companies will decide to spend their money on doing one thing or another, I truly don’t know,” says Gómez-Serrano. “What is clear is that very few mathematicians will have resources of that scale.”

If OpenAI and Anthropic keep striving for more and more impressive mathematical accolades, there might not be any open problems left for human mathematicians outside of those companies to wrestle with. That would dramatically change the field of mathematics.

Last week, UCLA mathematician Terence Tao wrote a Mastodon thread describing how important mistakes, wrong directions, and incomplete solutions are for the field. “In most cases in pure mathematics, the problems are posed not because we desperately want the solution to these problems in and of themselves, but because we have seen from past experience that human-directed efforts to solve these problems tend to spur further development of the field,” Tao wrote.

“Prematurely solving the problem by purely AI-powered methods—particularly without full transparency into the solution process—can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole.”

Humans might take longer than agents to solve mathematical problems, but in the process, they uncover new mathematical approaches and ideas that might inspire their peers and even birth their own subfields.

But when AI agents solve those problems instead—and when private companies keep the agents’ wrong turns from public view—those benefits disappear. It remains to be seen what else will vanish in the process. 

Agriculture relies on fossil fuels. It’s costing us.

If you’ve had to fill up your vehicle’s gas tank or buy a plane ticket lately, you’ve probably felt the effects of rising fossil-fuel prices. But farmers buying fertilizer for their crops are especially aware of just how far the ripple effects of the conflict in Iran have spread.

Fertilizer prices have been on a roller coaster this year, kicked off in part by trade disruptions and high prices for natural gas, a key ingredient in fertilizer production. Let’s take a closer look at why conventional fertilizer prices are so sky-high, and how a few more climate-friendly alternatives could bring farmers some relief.

As fossil-fuel prices go up, nearly all industries are affected, since most of our economy relies on these fuels to move goods and people around.

But fertilizer is even more intertwined with these fluctuations, because natural gas is used as both an energy source and a chemical input in the production of ammonia, a key fertilizer ingredient. So as natural-gas prices have spiked in recent months because of the war in Iran, fertilizer prices have followed. (It’s worth briefly noting here that fertilizer production is also a major source of greenhouse-gas emissions, accounting for about 2% of the global total.)

Fertilizer trade is being directly affected as well, since about one-third of global seaborne trade in fertilizers passes through the Strait of Hormuz, which has been effectively closed to commercial traffic because of the conflict. Access to fertilizer could get worse for some of the poorest countries around the world because of the strait’s closure, according to a report from the World Bank. While the US largely meets demand for nitrogen fertilizers with domestic production, some imports do come from the Persian Gulf.

At one point in April, the price of urea (the most commonly applied fertilizer) climbed above $850 per metric ton. That’s 80% higher than it was before the conflict and the highest level since 2022, when the Russian invasion of Ukraine and the resulting conflict caused fertilizer costs to hit record highs. Prices have come down significantly, but forecasts remain uncertain.

“There’s just this out-of-control supply chain that’s a lot more volatile than it’s ever been,” says Travis Frey, chief technology officer of Pivot Bio, a company making fertilizer with genetically edited microbes. (For more on these microbes, how they work, and what research is still needed, check out my latest story here.) 

Pivot says its products are cost-competitive with chemical fertilizers today. And because they don’t use natural gas as an input, they aren’t subject to the same price spikes. When the war in Iran started, Pivot increased the volume it planned to produce, dropped prices, and allowed farmers to lock in prices for three years, Frey says.

That could be a major help for those farmers, because high prices could be here to stay for a while. Some fertilizer prices could remain high through at least 2028, according to a report from CoBank, one of the largest banks for the agriculture industry in the US.

That’s partly because the war has caused long-lasting damage: 31 ammonia plants in the Middle East have been affected or shut down completely. That’s on top of 20 ammonia plants that have been damaged in Russia in recent years.

Ongoing high prices can be extremely challenging for farmers. “The fertilizer price spikes, and because farmers have paper-thin margins, this is a real problem,” says Tim Schnabel, founder and CEO of Switch Bioworks, another company working on advanced microbe fertilizers.

Higher costs can help push food prices higher, causing all of us to pay more at the grocery store. (It’s not just fertilizer, by the way. Farmers are also getting hit with wild diesel prices this year.) As long as we’re relying on fertilizers made with fossil fuels, food prices will be tied up with energy prices.

Switch and Pivot are among the companies looking to make alternative fertilizers that use microbes to provide nitrogen to plants. There’s a limit to how much synthetic fertilizers these products can actually replace: Depending on the crop and conditions, Pivot says, its products can replace about 25% of synthetic fertilizer today, and the company hopes to reach 40% to 50% of the total. But these alternatives could be a start to untangling fossil fuels and food.

“We can’t keep doing it like this,” Switch’s Schnabel says. “There’s no way we can build a society where the basis of the food chain depends on fossil fuels.”

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here

How AI plotted an interstellar journey to Alpha Centauri

A nonprofit organization called the Fermi Explorer Mission announced today that it intends to launch a spacecraft to our nearest star system by the end of 2029. 

It’s a hugely ambitious mission—if all goes well, the spacecraft could take up to 80,000 years to arrive at Alpha Centauri, which is 4.4 light-years away. And the spacecraft will follow a novel trajectory discovered by an AI system developed by Physical Superintelligence (PSI), an AI physics research lab. PSI is launching today with $58 million in funding led by Breakthrough Energy, a climate-focused investment group founded by Microsoft cofounder Bill Gates.

It’s not the first time this has been tried. In 2016, the billionaire tech investor Yuri Milner announced an interstellar mission called Breakthrough Starshot to launch humanity’s first spacecraft to Alpha Centauri. The plan was to use powerful lasers that would propel tiny probes to a fifth of the speed of light—fast enough to reach Alpha Centauri within 20 years. Milner pledged $100 million toward a proof of concept. But a decade later, nothing has launched.

“We didn’t want to do another Breakthrough Starshot,” says Philip Johnston, the cofounder and president of the Fermi Explorer Mission. “We’re dead set on something actually launching.”

To do that, “we are not constraining ourselves to doing it in a human lifetime,” says Johnston. “Let’s just figure out the way to get to another star.” 

The new mission, currently funded by individual private donors, is expected to cost just $15 million. The spacecraft will carry cargo weighing at least one kilogram. That will include artistic and scientific payloads, messages, and a copy of the Golden Record, a gold-plated disc of Earth’s sounds and images that NASA attached to its two Voyager probes in 1977 as a message to any civilization that might find them.

Engineering an interstellar journey is extremely difficult. Alpha Centauri is about 25 trillion miles away from Earth. One of the fastest objects that humans have ever launched, the Voyager 1 probe, has been flying since 1977 and has covered less than 1% percent of that distance. At its speed, the trip would take more than 70,000 years.

Johnston and his team spent a year trying, and failing, to find a way for a small, solar-powered spacecraft costing only $15 million to reach Alpha Centauri. They kept running into the knotty problem of how to give the spacecraft enough power without making it too heavy (and thus more fuel-guzzling). 

After the Fermi team struggled to find a workable route, Johnston mentioned the problem in a podcast hosted by Alex Wissner-Gross, a physicist who cofounded PSI. Wissner-Gross offered to run it through an AI system the lab developed, called Get Physics Done. It’s open-source software that takes a physics research question, breaks it into smaller tasks, and decides which simulations to run, using AI models including Anthropic’s Claude or OpenAI’s GPT.

A week later, the AI system turned up a novel trajectory, to Johnston’s surprise. It combined well-known orbital maneuvers in a way the Fermi team had not considered, according to a paper that has not been peer-reviewed. It suggested that the spacecraft could first slow down so its orbit swings in close to the sun—closer than Mercury. On each close pass, it would fire its engine so that the solar panels get four times the light, and a burst of thrust delivered at high speed would buy more energy than the same burst anywhere else. Because the engine would run only near the sun, the solar panels could stay small and the spacecraft light.

The system conducted the research mostly on its own for three days, running on a billion tokens, says Matt Pines, the cofounder and CEO of PSI. An astrophysicist on PSI’s staff steered it to follow the mission’s requirements, asked for a cost analysis and clearer charts, and checked the output for errors.

“The fact that it came up with an entirely different mission profile, one that was creative and not one [the Fermi team] had considered—that was the more surprising aspect,” says Pines. Still, the model lacks a human researcher’s judgment and taste, he says. It has no reliable sense of which problems are interesting or which approaches are worth pursuing, so it often gets stuck chasing dead ends or failing to explore different approaches. “I don’t think we’ve yet figured out how these models can internally represent something like that,” he says of research judgment.

Even if the Fermi probe launches, “we’re pretty confident that we will not be the first to arrive” at Alpha Centauri, says Johnston, since he expects spacecraft technology to improve. If an engine a thousand years from now is even 20% faster than today’s, a spacecraft launched then would still beat Fermi’s probe to Alpha Centauri by more than 10,000 years. 

But the Fermi project isn’t just an interstellar mission driven by engineering ambition. It’s also a quest to answer one of the oldest open questions in physics. In 1950, the physicist Enrico Fermi posed a puzzle: The galaxy has hundreds of billions of stars, most of them far older than our sun. Even a civilization traveling slowly between stars could spread across the whole galaxy in a few million years, which pales in comparison to how old the galaxy is. If there is intelligent life somewhere, we should have seen signs of its existence by now.

That means either reaching for another star is too difficult or other intelligent species simply haven’t bothered. But once the Fermi probe launches, we will become a civilization that can and wants to reach another star, meaning that neither explanation might be what’s keeping the galaxy unexplored. That could point us toward more unsettling possibilities, says Johnston. Maybe life like ours is almost unimaginably rare. Or maybe intelligent life is common but tends to die out before it can spread. 

If the latter is true, “one of those reasons could be that once you hit superintelligence, that for some reason is self-destructive,” says Johnston. “Maybe in the next 50 years, there’s some great filter that we do not pass through. That all intelligent civilizations, for some reason, do not pass through.”

How engineered microbes could help feed the world’s crops

Fertilizer is crucial for the global food supply, but making it uses a lot of energy and produces a lot of emissions. Some companies hope microbes can help.

A growing body of research shows that seeding the soil around a crop’s roots with beneficial microbes can help feed the plant, providing crucial nitrogen to help it grow. This could help reduce the need for chemical fertilizer, production of which accounts for roughly 2% of global greenhouse-gas emissions. It could also cut costs for farmers, a big boon—especially as the Iran war has sent energy and fertilizer prices skyrocketing in recent months. 

Humans have used biological fertilizers like manure for thousands of years, and companies have long been developing microbial fertilizers, including some that rely on genetic engineering. But it’s difficult to engineer microbes that can reliably provide nitrogen for crops while also thriving themselves. 

A startup called Switch Bioworks is taking a new approach that essentially allows microbes to establish themselves and grow into healthy colonies before shifting into nitrogen-producing mode. “We have to reinvent fertilizer,”  says Tim Schnabel, the company’s founder and CEO.

The air is nearly 80% nitrogen, but plants can’t use that “free” nitrogen directly because it doesn’t react readily with other elements. Instead, they rely on so-called fixed nitrogen, which has been converted into more reactive compounds such as ammonia. In nature, microbes can perform this nitrogen fixation. Some plants, like legumes, even have symbiotic relationships with nitrogen-fixing bacteria, housing them in nodules in their roots. Synthetic fertilizer is essentially industrial nitrogen fixation via the Haber-Bosch process, which uses natural gas to make ammonia that’s applied to fields.

Biological fertilizers aim to replace some of the synthetic fertilizer with microbes that can help fix nitrogen for plants. But one challenge microbial fertilizer companies have run into is that it’s energetically expensive for microbes to make and release ammonia. Putting all their energy into nitrogen fixation can hamper their growth.

There’s a certain level of colonization you want to see around the roots, Schnabel explains. It’s too expensive and logistically challenging to put all those microbes on the plant, so you have to rely on a smaller number of microbes to grow and divide, establishing the population.  

a smiling man in sunglasses and a Switch Bioworks cap walks between rows of corn
Tim Schnabel walks in a cornfield where Switch is engaged in early trials.
COURTESY OF SWITCH BIOWORKS

Switch Bioworks is betting that a genetic switch is the answer. A genetic switch is a section or sections of DNA that controls how genes are turned on and off. In this case it works by activating genes that help trigger ammonia production in and release from the cell. The company is working on several options for setting off this change in its microbes. The leading one is to have the microbes react to the nitrogen level in the soil: Once it drops to a certain level, they begin producing ammonia.

“You have this inherent biological reality, where it’s really expensive for microbes to fix nitrogen,” says Dan Blaustein-Rejto, director of food and agriculture at the Breakthrough Institute. It takes a lot of energy, and if they do fix the nitrogen, they want to use it for themselves, to build proteins and survive, he says. Adding genetic switches could help microbes grow and thrive and then help fertilize crops.

Switch is currently trialing its product in six US states, though it’s two to three years from a commercial product, Schnabel says. It’s initially focused on corn, the most-planted crop in the US, with over 90 million acres in 2026.

“It’s too early to tell exactly how well this works in the field,” Schnabel says. The company plans to harvest the plants in late October or early November, but as of August, some corn plants treated with Switch microbes already looked visibly healthier than those that hadn’t. And it’s still developing the products that will eventually make it to the market, he says.

Switch’s products could significantly help to clean up agriculture. “There’s a lot of potential for these companies and products to help farmers reduce emissions,” Blaustein-Rejto says. “This could be a really important solution for a quite hard-to-abate sector.”

While lab results have been promising, field trials are a crucial step in proving a product, Blaustein-Rejto says: “This is one of the final steps before they can go to market and make strong claims to farmers.” Independent trials are important as well, he adds, since there can be a large gap between a company’s reported data and what independent researchers find.

Pivot Bio is another company working to bring microbes to fields. Since it was founded in 2011, its products have been used on millions of acres of crops. The company now produces a range of products: Some versions can be added when seeds are planted, while others are applied to seeds before they’re even on a farm. The company also recently expanded beyond corn to make microbial fertilizers for cotton, wheat, and small grains including sorghum and barley.

Pivot’s initial challenge was trying to get microbes to produce nitrogen, whether they sensed it in the soil or not. Now the company is working to figure out how to make fitter, more robust colonies no matter the environment, says Travis Frey, the company’s chief technology officer.

“Growers right now are experiencing a double whammy from the farm economics point of view,” he says. Fertilizer costs are going up, and the price of commodity crops like corn has dropped. That’s a big opportunity for Pivot and others in the industry, Frey says: “This next decade is when biologicals on the farm are going to go mainstream.”

Fertilizer and seeds are two of the biggest costs for many growers, so reducing dependence on synthetic fertilizers could be a major help for agriculture, says John Havlin, a professor in the department of crop and soil sciences at North Carolina State University.

“I’m very excited about the future of the use of these products,” Havlin says. “They’ll eventually have a role to play to reduce the load of nitrogen that’s being applied.”

However, there’s a ceiling to the amount of fertilizer we can expect microbes to replace. Switch’s modeling suggests that around 50% is likely the maximum, though the initial product will likely be able to replace about 25% of a farm’s synthetic fertilizer, according to the company. Pivot has said its products can replace about one-quarter of the fertilizer used currently. 

That means synthetic fertilizer will be around for a long time. “There is no clear and plausible vision for replacing it entirely in the foreseeable future,” Blaustein-Rejto says. “So other ways to reduce emissions and reduce other types of nitrogen pollution from farms remain really critical.”

A startup claims it’s found a drug to make your blood young

I knew I’d officially become a “longevity influencer” this month when a company called Generation Lab reached out to offer me the chance to write about—and even receive—their new rejuvenation treatment, an injectable combination of two existing drugs which they call 1 Generation.

This wasn’t just any antiaging treatment, either. A company fact sheet says that it “blocks the systemic spread of aging in the bloodstream, reawakens the body’s own repair mechanism and restores health and youth to multiple tissues.”

The exclusive offer to join a “private cohort” receiving access to the treatment was being “extended to a limited number of people whose judgment on the science we trust, under close physician supervision.”

“There’s a long wait list including a lot of celebrities, well-known people, people who you have heard of,” Generation Lab’s CEO Alina Su told me in a conference call. “So we’re super excited to, like, create this drug and [are] letting you know that this first longevity therapeutic in [the] human body is here and it’s actually working.”

This was all clearly a marketing campaign and, like most other claims in the world of longevity medicine, probably too good to be true. What’s more, Generation Lab wouldn’t tell me what the two drugs are, making the proposition hard to take seriously.

Yet this pitch got my interest. The reason? I am a longtime follower of Generation Lab’s irascible scientific founder, Irina Conboy, a take-no-guff specialist in the bizarre, but scientifically fruitful, practice of joining together the circulatory systems of old and young lab animals.

The procedure, called parabiosis, or heterochronic blood exchange, is one of the few things actually proven to make an old animal act younger. 

But you can’t go through life attached by the veins to a younger person. So the quest has been to find practical ways to mimic those benefits. 

And that is something Conboy says she’s now achieved by hitting on a combination of two existing drugs that produce youthful effects—but without the need for any bodily fluid exchange. 

Over the last two months, Conboy has been taking the drug combination and feeling quite a bit better and more energetic, she says. So have a handful of other insiders, including Su, the company’s 26-year-old CEO, as well as Conboy’s husband and scientific collaborator Michael Conboy.

The company’s public relations staffer sent me a spreadsheet of the supposed benefits experienced by the participants, including improved vision in a 64-year-old female, longer landscaping sessions for a 59-year-old man, longer badminton games, improved hand grip, better erections than with Viagra, as well as “increased bowel movements.”

Obviously, any age-reversal drug that works would be the best-selling product of all time. The problem is that none have ever been shown to do that. And Generation’s approach to evidence raises several red flags. For instance, Su says the company now plans to launch a larger study, involving a hundred or more people. However, others are being offered the drug before firm evidence of a benefit is in.

In addition, the trial will be led by longevity doctors who practice alternative medicine. These include Matt Cook, founder of a network of clinics called BioReset Medical. Cook’s clinics offer what I would characterize as unproven treatments, such as stem-cell injections. A 2021 article reported he’d recommended covid-19 treatments designated as “fraudulent” by the FDA. Those included inhaling amniotic fluid from a nebulizer.

In a phone call, Cook says he initially agreed to prescribe the combination to Conboy and a few other insiders. (Everyone involved, he says, is a member of the same “happening” Silicon Valley scene around life extension.) “My confidence this was going to work was a zero. Like, I did not believe it was actually going to work,” says Cook. That is because “it’s two safe, generic drugs that are not even longevity drugs. 

But Cook, who also took the combination, says there was an effect. “And I have to tell you, it was the most surprising thing that ever happened to me,” he says. After taking the weekly injections, he experienced a sense of mental clarity that lasted for a day, and then several days. The other subjects, he says, reported a sense of well-being, as well as old aches and pains that evaporated. 

“There was a fairly significant broad set of symptoms that doesn’t really match what either of the drugs do,” he says. “So it’s interesting. [But] I have many more questions than I have answers … I still feel like I don’t understand exactly what the mechanism is.”

Prior to joining Generation Lab full time in 2025, Irina Conboy had a long career at the University of California, Berkeley, where she made her name via a widely cited series of experiments aimed at learning   whether the characteristics of age, or youth, could be transferred between animals via the bloodstream. In 2005, for instance, she surgically connected two-month-old mice to very aged ones, finding that the old animals’ ability to heal from injury dramatically improved.

This suggested that young blood must contain specific, powerful molecular signals capable of restoring the regenerative capacity of an old animal. As researchers zeroed in on candidate signals, several startups, including Elevian, a Harvard spinout, and Alkahest, tied to Stanford University scientists, were able to raise millions to pursue treatments for stroke or Alzheimer’s. However, those efforts still have not hit commercial paydirt.

Although it’s clear that young blood is good for old mice, scientists also know the reverse is true: Even a single exchange of plasma from an aged mouse has powerful, negative effects on a younger animal. That means that in the battle of age versus youth, says Conboy, “old blood dominates.”  

Conboy next tried a new approach. Instead of giving old animals young blood, in 2020 she removed the plasma (the part of the blood without cells) of old animals and replaced it with a neutral mixture of albumin and salt water. This was like pressing the Delete key on the swarm of hormones, antibodies, and other molecules released by the aged body.

The reset had even stronger rejuvenating effects than sharing a young animal’s blood, and within two years, Conboy had tried it on human volunteers who agreed to have almost half their blood volume removed and thrown away.

Longevity clinics have been quick to follow Conboy’s research. Some began offering “therapeutic plasma exchange” as a wellness intervention, for prices of up to $10,000 a session. One doctor that offers it, Jeffrey Gladden, says that after seeing Conboy’s publications, he raced to California to meet her and to get the procedure himself.

By the end of 2025, Conboy had retired from Berkeley and joined her startup full time. Generation Lab has been selling an aging diagnostic test, but Conboy still wanted to find a drug that could mimic the benefits of plasma exchange. To do that, she started using a testing system she’d codeveloped, which allows researchers to bathe human cells kept in a microfluidic device with blood serum from old people.

The new device, described in a paper this year, allowed her to quickly test her intuitions about what drugs might work. “This is an awesome screen or experimental system which allows us then to ask a question: If tissue becomes old in the presence of old blood, which molecule or combination of molecules will prevent that and allow tissue to remain young, even when the circulatory milieu is old?”

She says 1 Generation works by stimulating cells to regenerate and, simultaneously, neutralizing the old-age factors in a person’s bloodstream. “It allows human cells to remain young even when they are in the presence of old people’s blood serum,” says Conboy.

Without knowing what the drugs are, I can’t give my own opinion about their promise. But some doctors working with Generation Lab expressed surprise that the combination would work. “I will say right now, candidly, it’s an unknown how it will play out. But there’s certainly a lot of enthusiasm based on the quality of her work in the past,” says Gladden, the longevity doctor, who says his Texas clinic will also be involved in the study. “I am skeptical and optimistic at the same time.”

Pressed on why the names of the drugs remain confidential, Su said it’s to avoid imitators. That’s important because Generation doesn’t own the molecules. Instead, it will seek to quickly run a clinical trial, taking advantage of an FDA process that gives companies a short period of exclusivity, perhaps three years, if they can show that a combination of existing drugs has a new use.

In the meantime, the company’s effort to generate buzz is having some success. Generation is hosting an August 28 invite-only event that will feature talks by George Church, the noted Harvard University synthetic biologist, as well as researchers from Anthropic, whose CEO, Dario Amodei, has been saying AI will cure all disease within a few years. In an email, Church says he was also offered a chance to take the drug combination.

“People are clamoring for it,” Cook, from BioReset Medical, told me. He says his office has started getting calls “from high-end doctors, venture capitalists, and people in this community who are kind of biohackers.”

“It’s high-level influencer-type people,” he adds.

As for me, no, I won’t be taking the drugs any time soon. Not until I know what they are.

Correction 8/27/2026: A previous version of this story incorrectly said patients would pay to join the larger trial of a proposed longevity drug. While patients can pay to get the drug from a doctor, the larger trial will be free to patients and sponsored by Generation Lab.

Is Slate Auto’s new electric truck the EV Americans need?

EVs account for under 10% of total new-vehicle sales in the US, and the numbers are declining. From a climate perspective, that’s pretty dismal, especially because the transportation sector is the single biggest source of greenhouse-gas emissions in the country. 

One thing that could help turn that around? Slate Auto’s new truck—a vehicle that seems to buck every convention about selling cars in the fully loaded, range-obsessed US market.  

The company is going all in on simplicity, to the point of austerity. The Slate is a tiny, two-door pickup that’s shorter than a Honda Civic. Much of the media coverage has obsessed over the fact that the base model’s windows use hand cranks, a feature straight out of the 20th century.

Slate is also breaking away from other EV manufacturers’ efforts to compete with gas-powered vehicles on range. The truck sports a small lithium iron phosphate battery, ringing in at 65 kilowatt-hours, and its quoted top range is just 205 miles. For comparison, the most basic Tesla Model 3 can go over 320 miles on a charge.

By accepting a shorter range and forgoing the frills that Americans have come to expect in vehicles, Slate is able to offer the base model of its truck for less than $25,000. Some customers will choose to upgrade their truck with optional add-ons, like a Bluetooth stereo system, vinyl wraps, or even power windows. But the price will still likely come in well below the roughly $50,000 average for a new vehicle in the US.

It might seem an odd choice to go so small and simple in a market that’s increasingly sizing up—but the status quo hasn’t exactly been working for EV makers. A few years back, the hero for US automotive electrification was supposed to be the Ford F-150 Lightning. Announced in 2021, it was an electric version of the country’s best-selling vehicle. 

But Ford discontinued the truck in December 2025, just four years after its introduction. It’s not entirely the Lightning’s fault: The second Trump administration slashed tax credits and other support designed to boost EVs. 

The Lightning was also plagued by price increases. The base model cost roughly $40,000 when shipments started in 2022; during its final year, prices topped $54,000. One factor behind the increase was its massive battery; to reach an almost 300-mile range, the truck needed a battery with a capacity roughly twice that of Slate’s truck. 

That obsession with range is largely unwarranted. While Americans have historically chased distance, the average driver puts under 35 miles a day on the odometer, and nearly 90% of trips in a personal vehicle are 20 miles or less

Surveys of EV drivers show a similar trend: One recent study found that people tend to use less than 20% of their EV’s range on a typical day. 

The notion that drivers need far less range than they think they do might sound a lot more persuasive these days—even for Americans who tend to buy cars for their longest road trip instead of their everyday errands. Roughly half the country is struggling to pay for basic necessities such as gas and groceries, and 95% of Americans believe we’re in an affordability crisis, according to a recent Harris poll. An inexpensive vehicle that allows you to skip the gas station sounds like an attractive prospect. 

Affordable EVs have already found willing buyers in other parts of the world—even with reduced range. China currently has over 40 million EVs and plug-in hybrids on the roads, and roughly half of new vehicles sold are electric. On average, new EVs sold in China in 2025 had a range of just 247 miles. For the US over the same period, the average was 329 miles. (Europe falls between the two, at 281 miles.)

Slate is set to start delivering on preorders in late 2026. Its factory will have a capacity of 100,000 vehicles in the first year and 150,000 soon after. Thousands of customers have already put in preorders, according to the company.

Some people obviously believe in the truck’s prospects: Investors, including Jeff Bezos, have put nearly $1.4 billion into the company over three major funding rounds. Ford is jumping (back) into this space soon too. The automaker is working on its own small electric truck, which is expected to debut in 2027 at a retail price of around $30,000.

The key question that will determine Slate’s success or failure is whether drivers can get on board with a small, short-range vehicle for the sake of its price tag. At this moment, I’d bet the answer is yes. 

The inside story on why OpenAI agents hacked Hugging Face

The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations. 

Since the hack, OpenAI employees—as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today—have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or “alignment,” remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.

“It’s not something you can solve overnight,” says Kai Chen, who runs OpenAI’s alignment research team. “There are challenges we’ve been tracking for a very long time, and we’re now seeing them with much greater precision.”

The Hugging Face hack was a product of months of misbehavior from OpenAI agents, first as they were being trained and then as their abilities were being evaluated. This May, agents in training figured out how to use OpenAI’s infrastructure to communicate with one another and get support with difficult training tasks, including some that were impossible to solve without hacking or otherwise misbehaving. That “message board” was shut down.

Then in July, while being evaluated for their cybersecurity abilities, some models created a new message board. They were supposed to be isolated from the internet, but by working together they managed to get online, hack Hugging Face, and obtain solutions for the cybersecurity problems that had stumped them.

Based on their investigation, OpenAI researchers believe that events during the training phase led directly to the hack. “For almost every behavior that was worrisome at evaluation time, [we were able to] find some sort of associated behavior at training time that actually we think might have contributed to it,” says Eric Wallace, a member of OpenAI’s alignment research team. 

When models correctly solve problems during training, the behaviors that led them to that solution are reinforced, and they become more likely to engage in them in the future. So if a model completed a task in May after using the original message board, it became more likely to participate in a new message board later on. This phenomenon, where AI agents misbehave in ways that are reinforced during the training process, is known as reward hacking.

Reward hacking also helps to explain why the models worked so hard to make their way onto the internet. During its investigation of the incident, the OpenAI team found that, over the course of training, the models became more and more likely to probe their digital environment for weaknesses and use the tools at their disposal in unexpected ways—a sign that these behaviors were being gradually reinforced. By the time the models were facing tricky cybersecurity problems, they had learned that hacking was an effective way to achieve their goals.

These results suggest that the Hugging Face hack could have been avoided if the models weren’t rewarded for misbehaving during training. While researchers don’t yet know how to prevent reward hacking entirely, OpenAI is taking some steps toward mitigating its effects. The company will now look for signs of cheating in all frontier models during training by keeping an eye on their chains of thought—internal notepads where they sketch out their answers and plan their actions. 

This solution isn’t as much of a slam dunk as it might seem: In earlier research, OpenAI showed that punishing models that mention cheating in their chains of thought teaches them to keep their intentions hidden from researchers. But monitoring its models’ thinking does give OpenAI the chance to halt the training process and reassess its approach if models do start learning to reward hack.

If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with other agents or hacked its infrastructure during training, those behaviors had never been reinforced, so agent misbehavior can’t only be attributed to that reinforcement.

Jeffrey Ladish, director of the AI safety nonprofit Palisade Research, compares the agents to a human who commits their first financial crime. “It’s not like they had to do fraud before to figure out that fraud is an effective strategy, and you have the same problem with models,” Ladish says. “Alignment science needs to be understanding how model motivations get shaped, such that we can actually figure out how to get models to care about the consequences of their actions.”

OpenAI’s researchers do have a hypothesis for where some of the misbehavior originated. Before the models formed their first secret message board, they had been trained to communicate and coordinate with subagents—less powerful agents to whom a main agent can delegate tasks. 

That learned communication behavior could have transferred to this new setting. The METR report, which investigates the messages that the models sent to one another in detail, supports this hypothesis: One agent on the message board took charge and assigned tasks to the other agents, effectively treating them as subagents. OpenAI could try to prevent agents from secretly communicating with one another by not training this subagent behavior in the future, but that would make the models less useful. 

This tension between capability and safety is at the heart of what went wrong with the Hugging Face incident. OpenAI researchers also identified the models’ persistence as a key factor in the hack. 

When they were accidentally given unsolvable problems, the models didn’t give up; instead, they strove to find solutions by any means necessary. But persistence is also a virtue, of course, especially if we want agents that can undertake large amounts of difficult work independently.

OpenAI is working on giving models ways to alert humans if they are given impossible tasks. The problem of teaching models when they should deploy their abilities and when they should hold back, however, won’t be settled in a single postmortem. The training strategies that create superhuman coders—rewarding them when they successfully solve problems—might not work to teach models to use their skills judiciously and respect human desires and values.

“I think there’s a bunch of alignment science that still needs to be done where we can move past just using proxies for task completion,” says Ladish. “That will work to make models very capable, but I don’t think it will work to make them aligned.”

This company’s plans to deploy space mirrors could jeopardize the night sky for many

A company that plans to beam sunlight from space to Earth on demand might unintentionally brighten the night sky for many more people than intended, according to a new study.

Later this year, the US company Reflect Orbital plans to launch a test satellite called Eärendil-1 that will extend an 18-by-18-meter mirror in orbit. The goal is to test the feasibility of the company’s plans to launch up to 50,000 larger satellites, measuring 54 by 54 meters, and reflect sunlight to Earth on demand.

The case for doing this remains somewhat uncertain, but Reflect Orbital has said the goal is to prolong the hours of sunlight for various uses, including solar panel charging, emergency response, and military activities.

The launch, which was approved by the Federal Communications Commission in July, has been met with disbelief by astronomers and environmental groups. “This is incompatible with astronomy,” says Samantha Lawler, an astronomer at the University of Regina in Canada. “There is no way you can do this and preserve dark skies.”

Miroslav Kocifaj, an astronomer at the Slovak Academy of Sciences, and his colleagues have now calculated the broader effect on the night sky. In a new paper published online and accepted for publication in the space journal Astrophysical Journal Letters, they studied the extent to which the light the satellites beamed to the ground would scatter.

They found that within the beam’s target area, intended to be a circular patch five kilometers across, a single Reflect Orbital satellite would appear about 40 times brighter than the full moon in the sky. As far as 14 kilometers away from the center of the beam, the satellite would still be as bright as the full moon. 

Combining the beams of 400 satellites, which Reflect Orbital eventually plans to do, would yield a light as bright as 10,000 full moons within the five-kilometer area, or 2.4% as bright as the sun. Even up to 80 kilometers away, Kocifaj and his colleagues calculated, this combined beam would be visible as “a glow above the horizon,” he says. That means the night sky would be altered for many more people than those within the area Reflect Orbital intends to illuminate.

“Deploying these mirrors would be seriously damaging for astronomy and for the nighttime environment,” says Kocifaj.“ The damage extends far beyond the target area.”

Olivier Hainaut, an astronomer at the European Southern Observatory in Germany, says the results of the paper are not surprising but are still useful. “These guys are really good at that kind of modeling,” he says. “They know what they’re doing.”

Hainaut had previously modeled the broader effect of Reflect Orbital’s beams, finding they would brighten the sky up to 300% worldwide. This latest work gives an even more complete picture of what the impact would be. “These two papers really cover most of it,” he says.

Reflect Orbital CEO Ben Nowack disagrees with the findings of Kocifaj’s paper. “Some assumptions are simply inaccurate,” he says. “The critical point is that our safeguards, including maintaining exclusion zones, take account of scattering.”

He says that the company has “engaged substantively with legitimate concerns raised by astronomers, environmental researchers, and scientists,” and that their feedback has “informed our technology and operational plans.”

Kocifaj says that Reflect Orbital has not provided data to back up its claims. The company “states that safeguards exist, that scattering is taken into account, and that the models are being updated, but it gives no numbers, no description of the model, no assumptions, and no data,” he says. “There is therefore nothing that can be engaged with technically. Our calculations produce concrete figures.”

Reflect Orbital’s intention is to place its satellites into highly inclined orbits above Earth, almost from pole to pole. This will enable them to reflect sunlight in the hours before sunrise and after sunset, extending daylight hours in certain locations. Eventually, the company has said, it wants to place satellites high enough to provide light 24-7 to locations on Earth.

In August, a group of organisations including DarkSky International and the American Bird Conservancy urged the FCC to review its approval of the Eärendil-1 satellite. As well as the impact on astronomy raised by this and future satellites, the group also highlighted potential negative consequences for aviation and wildlife.

“Our primary request is straightforward: Reverse the Space Bureau’s order, and require a lawful public-interest and NEPA [National Environmental Policy Act] review,” the group said in a statement.

The satellites could also pose risks to humans, according to Reflect Orbital itself. In a filing with the FCC in March, the company said that observing Eärendil-1 with a telescope larger than 12 inches —which many astronomers have access to—may be unsafe for human eyes. It added, though, that such observations are “unlikely to … result in significant injury” because the satellites are not constantly bright.

Currently there is no global entity that could regulate satellites such as these, so approval falls to national regulators such as the FCC in the US. Michelle Hanlon, a space lawyer at the University of Mississippi’s School of Law, says there are “real benefits” to the plans proposed by Reflect Orbital. “It could extend the productive hours of solar facilities and provide light in remote areas or after a disaster,” she says.

However, the “legal basis is less clear than the technology,” she says. The FCC “does not have authority to license or regulate the operation of the solar reflector itself.” It can only authorize the use of radio frequencies to communicate with the spacecraft.

That means Reflect Orbital will need permission on a national and local level to reflect sunlight onto the ground, but if the beams spread as much as Kocifaj and his team predict, “the company could need approvals in more than one jurisdiction,” says Hanlon. “There may also be aviation-safety and cross-border questions.”

The fierce debate over the satellites shows no signs of abating. “It makes me sad that astronomers are having to spend their time doing these sorts of calculations rather than actually doing astronomy,” says Lawler. “This is not what we want to spend our time on.”

The next big thing in hydrogen could be underground

There’s a hunt for new sources of hydrogen, and the gas (or at least the right conditions to make it) could be hiding beneath our feet.

Hydrogen can be used as a fuel in everything from large trucks to planes to steelmaking. It’s often hailed as a climate solution because when burned, it produces water and oxygen—none of the carbon emissions that contribute to climate change.

In a new story for our latest print issue, freelance reporter James Dineen took a look at the 21st-century gold rush for naturally occurring hydrogen gas. This is an area of research I’ve been fascinated by lately, so let’s take a look at the potential and the questions that still linger.  

Today, hydrogen is overwhelmingly made using fossil fuels, generally natural gas. And most of it is used in petroleum refining or goes on to make fertilizer and other chemicals. But in recent years, many in the climate world have imagined a future where its production is clean too.

If you’d asked me a few years ago, I would have said the race for clean hydrogen was between methods that use electrolyzers powered with renewable electricity and operations that use established, fossil-fuel-based approaches cleaned up with carbon capture. But both those methods have struggled to gain ground, largely because of their high cost.

Lately, there’s been momentum in a new field: geologic hydrogen. Companies have found naturally occurring hydrogen resources across Africa, Asia, Europe, Australia, and North America. 

The US Geological Survey publishes a map of hydrogen prospectivity in the country (basically, where the gas is most likely to occur naturally). One hot spot is the Midwest—specifically the Midcontinent Rift, winding from Kansas to Michigan. The planet’s crust was stretched and split there about a billion years ago, causing molten rock to push up through the crack. The result today is a lot of iron-rich rock, which can react with water to readily form hydrogen.

HyTerra, an Australian company, is searching for hydrogen across Nebraska and Kansas, and it’s already found samples of gas with hydrogen concentrations up to 96%. Koloma, one of the most capitalized companies in the space with total funding over $400 million, is prospecting in the region as well.

One of the major questions these companies have is just how much hydrogen is produced by natural processes, and whether it can be effectively captured. Hydrogen is an incredibly light gas with a small molecular weight, so it can slip through even tiny cracks in rock.  

As James covered in his story, there are some promising signs. Researchers examined a few dozen boreholes at a mine in northern Ontario and found that each one released eight kilograms of hydrogen per year. Given that there are more than 14,000 boreholes at this one site alone, that’s a lot of potential hydrogen to capture.

Rather than hunt for a natural source, some companies are taking matters into their own hands and helping reactions along. The idea behind so-called stimulated geologic hydrogen is to find a spot where there are favorable conditions for hydrogen production but no accumulated resource. By adding water, a catalyst, or some other factor needed for the reaction, it’s possible to kick-start the process.

Vema Hydrogen is a Texas-based company looking to produce hydrogen from subsurface rocks by drilling wells and injecting water and catalysts into them to stimulate reactions. The company is testing its process in wells in Quebec and hopes to start full-scale production in 2028.

Other companies are tackling different aspects of hydrogen production: Eden GeoPower, for example, is using electricity to form fracture networks in rocks, creating more routes for water to get in. (The technology could also be useful in enhanced geothermal projects.)

There are still a ton of unanswered questions here: Hydrogen is notoriously difficult to move around and store, requiring either a lot of space or super-low temperatures to force the gas to become liquid. But if the engineering and logistics work out, this could spell a new beginning for hydrogen. 

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here

We still don’t know how people are really using AI

AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. 

“There is no independent source to corroborate it,” says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab. 

Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The intent is to provide independent sources of information that can help researchers and policymakers assess how people are using generative AI. Highly consequential decisions about AI’s benefits and risks are currently being made on the basis of very limited data, says Reuel. 

The AI Observatory found that AI use differs significantly across models and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use. 

The Anthropic Economic Index is one of the best-known and most widely cited sources of AI usage data, but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses. 

When the AI Observatory researchers applied Anthropic’s methods to their dataset, they found that nearly half the conversations—48%—would have been filtered out. Those non-work-related conversations were more likely to involve health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT, similarly, found that only 30% of consumer use was related to work.)

Anthropic has released separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, but “having [the AI Observatory’s] bird’s-eye-view analysis” rather than leaving that information “sectioned off into a separate report” helps researchers understand the different uses more consistently, says David Widder, an assistant professor at the University of Texas at Austin, who researches how people interact with AI systems and is not involved with the AI Observatory. 

The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded. 

Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, as indicated by growing numbers of prompt tokens, response tokens, and conversation turns. 

There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile, the AI assistants’ self-disclosure (i.e., admitting to being a chatbot) decreased. 

Additionally, exchanges that the researchers labeled as sensitive—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—became less frequent. That might suggest that platforms were generally deploying more effective safeguards. 

The AI Observatory also found that topics, interaction styles, conversation structures, and the likelihood and type of sensitive use cases differed from one model to another. 

For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.) 

Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance. 

There were even differences between different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction. 

Companies’ reports, however, didn’t tend to capture these nuances between or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel. 

To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions aggregated 85,633 conversational turns (that is, the user prompt and corresponding AI response) across 24,521 conversations from seven real-world datasets collected in previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025. 

But these conversations are a drop in the proverbial bucket compared with the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.  

An Anthropic representative said the company’s published research reflects its research teams’ specific questions and interests and that it’s important to support external independent research. OpenAI did not respond to requests for comment. 

The fact that the AI Observatory’s dataset draws from voluntarily provided sources means it’s probably underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all AI use. 

The project’s work, though, broadens access for the research community. AI companies don’t typically share their chat data for analysis, which means their reports tend to focus on the findings that paint them in the best light, independent researchers like Reuel and Widder say. 

“When we want to ask, for example: is Anthropic’s general-purpose AI system … used mostly for good or mostly for bad … we don’t have a way of answering that question because that information is proprietary,” explains Widder.

The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Ideally, Reuel says, the AI companies would share their data with independent researchers—in ways that protect user privacy, of course. But as it currently stands, she says, anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives.”   

AI’s recursive self-improvement might not come so quickly after all

The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. 

But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research—free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.

A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of  papers accepted by a top machine-learning conference. The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence.

Most existing research on how agents can automate AI research evaluates their ability to complete narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark. But making progress in AI research also requires open-ended thinking—choosing a set of hypotheses, deciding what evidence would settle a question, or knowing when to start over. 

To test agents on those kinds of skills, the researchers in the study proposed a new method of evaluation called “shadow evaluation,” which requires the AI to answer a research question from a high-quality unpublished paper. 

The researchers asked Anthropic’s Claude Opus 4.8, running on open-source software called OpenClaw, to tackle such questions, in this case from two papers submitted to the prestigious machine-learning conference NeurIPS 2026. 

The first question was whether a large language model’s “personas,” which determine its behavior, can be controlled by editing the model’s weights (the billions of numbers that store everything it learns during training). The other asked how to design a detector that points out when a model that makes predictions based on spreadsheet data has become unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training data or find them online. 

The agents were given six days, $3,000 in Anthropic API credits, a GPU budget to run the experiments, their own virtual computers, and access to the open web to produce a research paper worthy of publication at a top-tier AI conference. The papers’ original authors graded the agents’ papers as they would evaluate one submitted to a conference.

Those authors rejected both papers. 

The agents were capable of all the engineering required to conduct the research, the human scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results. 

“On the other hand, the agents were unambiguously bad at carrying out the research itself,” says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” he says. 

That’s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn’t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn’t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch. 

The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn’t effectively use resources, such as tokens, compute, and time. And they couldn’t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be.

For all their failures, the agents didn’t engage in the misbehavior that researchers call “reward hacking,” hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project. 

The reason AI models are good at research engineering but not at open-ended research may come down to how they’re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. “But it’s harder to create environments to train these models when the task itself is open-ended,” he says.

Kapoor says the team is now conducting the experiment with Mythos, Anthropic’s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.

There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer.

Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled “When AI Builds Itself,” charting its progress toward models that speed up their own development. In July, OpenAI advertised the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.

The new finding may echo what AI companies are finding internally, regardless of their most optimistic public statements. Anthropic cofounder Jack Clark wrote in his newsletter Import AI that it rhymes with what the company found when it tried to automate some aspects of AI safety research. 

“There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers,” he wrote. He called AI systems’ lack of creativity a “bearish signal on short recursive self-improvement timelines.” 

AI companies do have every incentive to develop AI systems that can rapidly accelerate their own progress, just as they did to make the models better at coding. OpenAI has made building an automated AI researcher an explicit goal, and Anthropic identifies self-improving AI as the industry’s next milestone. 

“If there is investment and then conscious effort toward this direction, I feel like there would be interesting progress, even if it’s failing currently,” says Najoung Kim, a professor of linguistics and computer science at Boston University who researches how AI agents can automate AI research but did not work on the study. On the other hand, it’s possible that AI progress may be bifurcated. AI systems might race ahead on narrow tasks—the kind that can be scored—while advancing slowly on open-ended research. 

The big open question, then, is how crucial open-ended research is to recursive self-improvement—whether AI systems can grind their way there without it, simply by improving on the narrower tasks. “If we look back to the biggest advances in the field, the invention of transformers or the invention of big new architectures that allowed us to make a lot of AI progress—all of those did require creative leaps,” says Kapoor. 

“That said, others have this hypothesis that all of what we need for transformative AI, in particular for recursive self-improvement, is already there.” That would include making a model train faster and boosting its benchmark scores.

“That’s frankly the trillion-dollar question right now,” he says.

Cloning could be used to save species—or make human “organ sacks”

This week I spoke to scientists who have found a way to turn male mouse embryos female. They’ve developed a CRISPR-based approach to essentially cut out the Y chromosome. It allowed them to create female clones of male mice.

That’s right: female animals that are genetically identical to males, except for the missing Y chromosome. Takashi Ishiuchi, a reproductive biologist at the University of Yamanashi who co-led the work, told me it felt a bit like sci-fi.

Ishiuchi and his colleague Shogo Matoba of the Riken BioResource Research Center hope their approach could be helpful in conservation efforts, especially in cases where we might have only a few individuals of a species left. But cloning has multiple uses, ranging from the cool to the outright creepy.

We can’t talk about cloning without mentioning Dolly, the celebrity sheep born in 1996 and the first mammal successfully cloned from an adult cell. In that case, scientists took the DNA-containing nucleus of an adult mammary cell from one sheep and transferred it into an egg cell that had had its own nucleus removed. The resulting embryo was transferred to a surrogate sheep, which gave birth to Dolly—an animal genetically identical to the DNA donor.

The scientists behind that work were interested in genetically modifying livestock. Farmers have essentially been doing this for thousands of years through selective breeding, but cloning allows scientists to create genetic replicas of animals with desirable traits.

Cloning is also being used to replicate deceased pets, including, famously, those of Barbra Streisand and Tom Brady, among others. For a price somewhere in the tens of thousands of dollars, a company can take cells from your pet and turn them into a living, breathing clone.

Considering that cloning also requires egg cells from another animal, and a surrogate animal to carry the pregnancy, not everyone is on board with this, especially since there is no medical or environmental need for the procedures. One bioethicist, Jessica Pierce, has described this aspect of dog cloning as “the exploitation of the canine underclass.”

The case for cloning is stronger when it comes to conservation—where some argue there is environmental value.

Scientists have been preserving animal tissues for years. Some of these tissues are cryopreserved at low temperatures in “frozen zoos.” The facility at the San Diego Zoo, for example, currently has cells from over 1,300 species. Some of these samples were taken decades ago.

Preserved tissues like these have enabled scientists to create clones of animals considered close to extinction, including black-footed ferrets and Przewalski’s horse. But they might also help us bring back extinct animals.

In 2009, researchers in Spain described how they’d cloned an extinct wild goat, the Pyrenean ibex, using skin cells that had been cryopreserved a decade earlier. In that research, the team used egg cells from domestic goats to create a total of 439 embryos. Ultimately, only one goat—a female—was born. She died minutes later because of a defect in her lungs.

Poor Pyrenean ibex. It’s the only animal we know of that has gone extinct twice.

The biotech company Colossal Biosciences is hoping to use old—and potentially ancient—genetic material to bring back long-extinct species like the thylacine and woolly mammoth. So far, the company’s efforts have largely involved modifying the genomes of modern-day animals.

Technically, it’s also possible to clone humans. As far as we know, no one has done it. But some have played with the idea. One biotech startup founder has pitched an idea for “brainless clones”—human clones that lack a brain but contain all the organs people might need to replace their own in future. My colleague Antonio Regalado described that pitch in March. (I had to pause eating my lunch while rereading it.)

Scientists have done a hell of a lot with cloning over the last few decades. I’m excited—but also slightly nervous—about what the coming decades will bring.

This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.

❌