Reading view

There are new articles available, click to refresh the page.

Shape-shifting mirrors on NASA’s new space telescope could unveil Jupiters like our own

When NASA’s Nancy Grace Roman Space Telescope launches, as early as the end of next month, it will attempt one of astronomy’s most precise disappearing acts to date. The telescope will carry the first space-bound “active” coronagraph, an instrument that effectively erases most of the light from a star during photography.

It will allow astronomers to take the first pictures of planets orbiting other stars that are similar to those in our solar system. Ultimately, it could pave the way for a future mission that could snap the first photos of Earth-like worlds.

“I hope it’s remembered for it being that critical stepping stone for … finding Earth 2.0,” says Brandon Creager, the instrument’s lead mechanical engineer at NASA’s Jet Propulsion Laboratory (JPL).

Named after Nancy Grace Roman, NASA’s first chief of astronomy, this new telescope will carry a roughly 300-megapixel wide-field camera that will enable it to capture images about 100 times larger than the Hubble Space Telescope’s widest exposures at a similar resolution.

These capabilities will help astronomers unpack the mysterious identities of dark matter and dark energy—and to detect around 100,000 new exoplanets, planets outside our solar system, whose presence can be inferred from the way they distort the starlight of more distant stars. Javier Viaña, a research scientist at Harvard who has had two projects selected for Roman’s highly competitive first year of observing, compares the leap to moving from “interviewing a handful of people” to “conducting a global census.”

Another camera will use the coronagraph, blocking out a star’s light as it observes one stellar system at a time. The instrument will allow astronomers an unprecedented look at the space around stars, enabling them to see smaller, dimmer, and more close-in exoplanets. “It’s giving us the ability to see planets that we haven’t been able to physically see before,” says Creager.

The anatomy of a vanishing trick

Coronagraphs in space aren’t new. But earlier incarnations, such as those currently aboard Hubble and the James Webb Space Telescope, use a stationary system to block a star’s blinding light. The approach does help, but it’s a bit like putting your thumb over a flashlight while searching a dark room for a firefly. Though the bulb vanishes, stray glare can still escape and overwhelm the light of the insect. Inside a telescope, that glare can come from light leaking around the edges of machinery or from minuscule imperfections in mirrors and coatings that can scatter starlight into speckles. All this can hide, or even impersonate, a planet.

Roman’s coronagraph, however, will attempt something completely unseen in space telescopes until this year: Before each observation, it will measure that leftover light and try to suppress it, a technique known as active wavefront control.

The telescope is able to do this because it contains two deformable mirrors. Each has a 48-by-48 checkerboard of actuators (tiny pistons) beneath a thin, deformable sheet of glass. Applying a small amount of voltage makes the actuators contract and tug their patches of mirror slightly backward, like thousands of microscopic fingers delicately sculpting a surface.

The effect is very subtle: Each patch of mirror can deform by up to 0.5 micrometers, or about one-fourth the size of an E. coli bacterium, and in increments as small as approximately 10 picometers. That’s about a tenth the diameter of a hydrogen atom, says Ilya Poberezhskiy, the instrument’s project systems engineer at JPL.

The actuators allow the mirrors to create an “active wavefront,” where each component is moved to the perfect position to cancel out incoming waves of unwanted light—a bit like a pair of noise-canceling headphones, but for light instead of sound. The “canceled-out” light creates a “doughnut-shaped region around the star where we suppress starlight and where we’re hoping to see exoplanets,” says Poberezhskiy.

Compared with current space-based coronagraphs, the system is expected to improve sensitivity to exoplanets against the glare of their host stars by a factor of up to 1,000, revealing planets that would have been far too faint to detect before.

Like Hubble and JWST, Roman also uses masks, patterned plates placed in the path of the light that are designed to block the photons that run into them. One tool in Roman’s mask arsenal is “silicon grass,” a thicket of microscopic spikes on some masks that can be used in certain configurations to absorb photons so they don’t bounce around the telescope and accidentally reach a detector.

Light entering the forest bounces deeper and deeper between the blades and gets trapped instead of reflecting back toward the camera. “Once the light gets into there, it never gets out,” Poberezhskiy says. The mirrors and masks form a succession of gates and hedges to guide as much of the preserved planetary light as possible toward the final detector.

Alien Jupiters

This elaborate setup could open a new chapter in the direct imaging of exoplanets. Nearly all exoplanets photographed so far are oversize youngsters that are nothing like the residents of our solar system: several times the mass of Jupiter, still glowing with the heat left over from their birth, and orbiting tens or hundreds of times farther from their star than the Earth is from the sun. This is because they are relatively easy to see. Their size, warmth, and distance from their parent star makes them shine brightly in infrared light, far away from the worst of the stellar glare.

Roman, however, could directly image a true Jupiter analogue—a planet similar to Jupiter in mass and circling a sunlike star a few times farther out than Earth is from our sun. Unlike the hot Jupiters we can see now, this one would be a much more mature gas giant like ours, primarily reflecting its parent star’s light after billions of years of cooling instead of heavily emitting its own.

Astronomers have been able to infer the existence of such planets from the gravitational wobble they impart to the star. Roman instead will collect starlight reflected from the planet itself. “We’re not looking at the star. We’re not looking at the effect of the planet on the star,” says Meredith MacGregor, a professor of astronomy at Johns Hopkins who has also secured an observing program. “We are actually looking at the planet, and that is super powerful.”

Once this instrument becomes available, it will become the scientists’ turn to do their jobs. “I’m honestly a little terrified about how we’re all going to deal with it, because I think it’s just so much data,” MacGregor says. “I think people will legitimately still be working on Roman data for decades.”

But don’t expect to see a 4K photo of an alien Jupiter in the coming months. Roman will not be able to resolve such a planet into a solid globe—at best, it will likely resemble a smattering of pixels. Still, that will be enough, MacGregor says, as Roman can then use the coronagraph to get information on the various wavelengths of light from the planet, which can tell astronomers about its atmospheric chemistry.

“You’re taking something that’s a point of light and turning it into an actual world,” she says, “because if you know that about its atmosphere, now you know something about the surface of the planet and the possibility of life being on that planet, right? So that’s a big step.”

During its first observations, scientists and engineers will see whether they can hold a star at the very center of the coronagraph’s masks, shape the mirrors, “dig” the dark doughnut (as Poberezhskiy describes it), and then maintain everything as the spacecraft moves through space and actively changes temperature.

The results will inform NASA’s proposed Habitable Worlds Observatory, the daydream of many an exoplanet astronomer, which will in theory be able to separate the light of an Earthlike planet from that of a sunlike star, over 10 billion times brighter.

Creager, who has worked on the instrument since 2018, is proud of the achievement: “Not too many people get to say, ‘I built something and it’s taking a picture of a planet that’s at a star that’s 50 light-years away or 100 light-years away.’” He imagines the moment he and his team will be able to look at the first image as it arrives: “Yes, we did that.” While the planet may show up only as a tiny dot, Roman’s achievement will be the darkness engineered around it.

AI is more likely than humans to form biases when hiring

The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience—and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases. 

Researchers at Princeton University and the University of Chicago ran LLMs, including ChatGPT, Claude, and Gemini, through a simulated hiring game, adapted from a psychology study that explored how humans can form stereotypes. Each model was told it had been hired as a consultant by the mayor of a fictional city and was then asked to help hire people for 20 jobs, including doctors, lawyers, child-care aides, and janitors. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. 

In each round, there was a new job opening and four candidates, one from each group. After the model hired a candidate, it learned whether they succeeded at their job and moved onto the next round. The model was told to make as many successful hires as possible over 40 rounds. Unbeknownst to the models, all candidates were equally likely to succeed at every job.

The models quickly started segregating candidates from different groups into different jobs on the basis of early observations of hiring outcomes. For example, when a model was told an Aima had failed as a doctor, a job considered to require high levels of warmth and competence, it veered away from hiring all Aimas as doctors. Instead, it started hiring Aimas as janitors, which the model classified as being less warm and competent than doctors. Newer models with higher reasoning capabilities, such as OpenAI’s o3 and DeepSeek’s R1, showed stronger biases.

The models were even more likely to stereotype people by demographic group than the human participants in the original study. On the study’s segregation scale, where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI’s reasoning model o3 scoring 1.83, close to the maximum possible.

That’s because LLMs “really are eager to create generalizations from limited data,” says Ryan Liu, a PhD student at Princeton University and a coauthor of the study, which was published in a paper at ICML in Seoul in July. “That’s literally a lot of what they’re optimized for.”

Every decision-maker, human or machine, faces a trade-off between sticking with what worked before and trying something new that might work better—a phenomenon psychologists call the “exploration-exploitation dilemma.” It’s like choosing between a new restaurant and your reliable favorite. 

Because LLMs are trained on math, coding, and science problems—tasks that reward generalizing from just a few examples—they can settle on a hunch too early. And the same instinct that helps LLMs crack logic puzzles also makes them quick to stereotype. When LLMs rush to generalize in social settings, “that’s when things tend to go wrong,” says Liu. OpenAI and Anthropic did not respond to requests for comment.

The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang, a computer scientist at Cornell University who did not work on the study. When a chatbot draws on its previous conversation history, it can “over-index on the same kinds of behaviors it’s experienced before” and form biases, she says.

Simply having chatbots remember less isn’t a fix, though, because users want chatbots to remember what they say. “We still are trying to figure out just the right amount that isn’t too much or too little,” says Wang.

Telling the model to be fair didn’t change its behavior much. “Either it can’t put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires,” says Liu. But promising the models an additional bonus for diverse hiring made them far less biased. The trick, then, is to design goals that “incorporate desirable social values in order to make the large language model act in socially desirable ways,” says Liu.

The models also became less biased when they were told more personal information about individuals. In another experiment in the same study, the researchers asked the models to resettle members of different ethnic groups in cities across Canada. When the models were told personal information relevant to the ability to adapt to a new city, such as age and education, they were less likely to segregate people by their ethnicity. But when they were given irrelevant information, such as hair color and tattoo shape, the models largely fell back to sorting people by their ethnicity again. 

To what extent AI systems will stereotype job applicants in the real world is still an open question. While the models in the experiment immediately learned whether they’d made successful hires, a model screening résumés in the real world doesn’t get an instant report card. Companies can take a long time to find out whether a new hire is any good, if they ever do.

But when feedback does trickle in, a model could still read too much into those results when making future hires. As companies increasingly deploy LLMs to screen résumés and even conduct interviews, the finding that models can form biases from their hiring experience “is a really serious implication that they should grapple with,” says Wang. 

As LLMs learn from experience to make decisions about who gets hired, who gets a loan, or who gets parole, the biases we should worry about may include ones no human ever taught them. “These novel biases—they’re sort of ever present,” says Liu.

There’s a lot of hype around perimenopause. Don’t buy it.

Perimenopause has entered the chat. Perimenopause—and its better-known relative, menopause—used to be considered taboo. Not anymore, thanks at least in part to TV doctors and social media influencers. Perhaps it’s my age, but these days, both my algorithm and my conversations with friends increasingly swing toward perimenopause.

Menopause is defined as the life stage that occurs a year after a person has had their last period. Perimenopause is the sometimes years-long period before that point, which can also feature all the symptoms we’d typically associate with menopause.

Today, information about perimenopause is more prevalent and accessible than ever. If you’re a woman in your 40s and you’re not feeling 100%, chances are there’ll be someone online ready to tell you you’re in perimenopause. And that you might want to start spending your money on blood tests, apps, and supplements or demanding hormone replacement therapy. But as regular readers might have guessed by this point, it’s not that simple.

Perimenopause tends to start around the age of 46 or 47. It’s during this time that many women start to experience some symptoms like hot flashes, irregular or unusually heavy periods, or anxiety, for example. And it can be heavy going. “Often symptoms are at their worst in the perimenopause,” says Mary Ann Lumsden, former president of the International Menopause Society.

That’s because hormones can fluctuate wildly. Levels of estrogen, progesterone, luteinizing hormone, and follicle-stimulating hormone can roller-coaster before leveling off after menopause. And that’s why, despite what some marketers will claim, there is no test for perimenopause.

“You can’t interpret hormone [measures] because they change so much,” says Lumsden. “And that is quite normal.”

That doesn’t mean women should have to put up with symptoms. But exactly how those symptoms are treated is another topic that has been clouded by misinformation.

Last week, I told a friend about some unusually bad pelvic pain I’d experienced. Her immediate advice was to find out if I was perimenopausal and, if I was, to request hormone replacement therapy (HRT) as soon as possible. If my doctor wouldn’t prescribe it, she continued, I should simply find another doctor who would.

This line of thinking has been heavily promoted on social media platforms, says Paula Briggs, a former chair of the British Menopause Society who currently leads the menopause service at Liverpool Women’s Hospital. But it’s not helpful.

HRT is essentially designed to top up or replace hormones like estrogen and progesterone, which naturally decline around menopause. There are lots of different drugs that can be taken in lots of different ways and at various doses.

While it does come with some risks and won’t suit everyone, HRT can be immensely helpful for many menopausal women. Not only can it help with many of the common symptoms of menopause, but it can also help prevent osteoporosis and maintain muscle strength.

But these drugs were trialed in, and approved for, menopausal women, says Lumsden. They won’t have the same effects in perimenopausal women. “If you give standard HRT, it may well get swamped by [the woman’s] own hormone production,” she says.

HRT can also cause abnormal bleeding in perimenopausal women, says Briggs.

She’s concerned about the messaging on perimenopause that is being promoted on social media. Particularly worrisome, she says, is the way younger women are being encouraged to assume they are perimenopausal and seek out HRT treatment.

“It’s almost cult-like, this idea that everybody must have HRT,” she says.

And then there are the supplements. There’s been an explosion in marketing for vitamins and supplements specifically targeted to middle-aged and menopausal women. But the evidence for these, too, is either limited or nonexistent. “I can’t see a mechanism for a lot of them,” says Lumsden.

Women who take these supplements don’t always know what they’re getting. Some of Lumsden’s patients have told her they take testosterone supplements to manage their symptoms. But blood tests revealed no increase in testosterone levels. “Whatever they’re getting, it’s not testosterone,” she says.

At any rate, not all the symptoms women experience in midlife can be blamed on hormones. The lengthy lists of perimenopause symptoms shared on social media include fatigue, brain fog, aches and pains, digestive issues, and more. “These do not link closely to the obvious menstrual cycle changes and hormone changes … across menopause,” says Nanette Santoro, a professor of obstetrics and gynecology at the University of Colorado Anschutz who studies menopause.

If you’re experiencing any symptoms, it’s worth getting them checked out to make sure they’re not being caused by something else. My own pelvic pain, for example, is almost definitely the result of endometriosis—a condition that can be made worse by HRT, Lumsden tells me.

At any rate, by the time women reach their 40s, many are already juggling care for children and aging parents, often while holding down a job (and dealing with pressures from societies that don’t appear to value older women). It’s an exhausting time—and not all of that exhaustion can be blamed on hormones.

As Santoro puts it: “Attributing everything unpleasant that happens to a woman over 35 to perimenopause is not based on any scientific evidence.”

This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.

Why heat pumps are still so hot in the US

It feels as if it should be illegal to even think about heating appliances during the height of summer—seriously, these heat waves in New York have been brutal—but we need to talk about heat pumps.

The appliances use electricity for heating, they’re incredibly efficient, and they’re on the rise. (For what it’s worth, many heat pumps can also be run in reverse to cool buildings.) In the US, heat pump sales have doubled over the past 15 years, according to a new report. And they’re winning the heating race against fossil fuels, outpacing natural-gas furnaces by 32% during the first quarter of 2026.

These stats are especially striking at this moment, because a key tax credit for heat pumps just ended with the close of 2025. But you wouldn’t know it from looking at the data. Why are heat pumps still so hot?  

In case you need a quick refresher, heat pumps use electricity to essentially move heat from one spot to another. A refrigerant moves around a loop in the device, expanding and compressing, gathering and releasing heat at different points in the cycle. (For a more in-depth look at the thermodynamics, this explainer I wrote in 2023 still holds up.)

The result is an appliance that can be incredibly efficient. Once you pay for and install a heat pump, it’s generally significantly cheaper to run than a gas or oil furnace or other types of electric heating systems. And because they’re more efficient and don’t involve burning fossil fuels, heat pumps can be a major help in decarbonizing buildings.

One of the major hurdles to wider use of heat pumps is the appliances’ cost: They tend to be more expensive to buy and install than gas furnaces. For this reason, many governments offer incentives to encourage their adoption. In the US, people who installed heat pumps between 2023 and 2025 were eligible for up to $2,000 in tax credits.

Last year, though, the Trump administration slashed those tax credits, along with many of the other incentives that were part of the 2022 Inflation Reduction Act. Effective January 1, 2026, no more financial help for heat pumps.

I think I’ve seen this film before, and I didn’t like the ending. Tax credits of up to $7,500 for new EVs ended on September 30, 2025. In the quarter leading up to that deadline, sales spiked as people rushed to take advantage of the incentive. Then they fell off a cliff. Things are starting to normalize now, but clearly the tax credit’s sunset had a major effect.

But as it turns out, heat pumps are an entirely different story. In the first few months of 2026, sales have actually gone up, as Lucas Davis, an energy economist and UC Berkeley professor, points out in a new analysis.

Heat pump shipments were flat from December to January and have seen a gradual rise since then, according to data from the Air Conditioning, Heating, and Refrigeration Institute, a trade group that represents about 90% of the US market. This increase from winter into spring follows a seasonal trend seen in previous years—and it’s actually a bit stronger in 2026.

This data isn’t what you’d expect to see if losing the tax credit were hurting demand. As Davis lays out in his post, it seems the credit wasn’t really convincing people to install heat pumps, or at least the case for doing so was sufficient without the added incentive.

“It appears that the U.S. market for heat pumps is strong enough that it does not depend on tax credits,” Davis writes.

In 2024, MIT Technology Review put heat pumps on our annual list of breakthrough technologies. “We’ve entered the era of the heat pump,” I wrote at the time.

While heat pump sales have been up and down over the last few years, the era is going strong. The appliances have outsold gas furnaces in the US for the last four years. It’s not just the US, either. Countries including China and Germany have seen strong movement to heat pumps in recent years.

There’s rarely a straight path to adoption for new technology, especially something that requires so many individual households to make a significant change. But it’s encouraging that a major decarbonization tool is going strong, even when roadblocks pop up.

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its most robust release yet.

GPT-Red automates a type of safety evaluation for software systems known as red-teaming, which is typically done by a team of human testers. The aim is to find as many different ways to break or hijack a system as possible. The weak spots can then be patched before the final version of the software is released.

As LLMs become more complex and get used in a wider variety of tasks—especially in the form of agents, which can interact with computer files, websites, and third-party code as well as other agents—it’s hard for teams of people by themselves to keep up with all the types of attacks that might take place. “The risk surface grows and the blast radius also grows,” says Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red.

OpenAI built GPT-Red to future-proof its safety testing process. “As more capable models become available, we will have already designed the system that can discover new modes of attack,” says Dylan Hunn, a research scientist at the company and fellow co-creator of GPT-Red. The researchers say it has already come up with new types of attack that had not been seen before.

OpenAI focused most of its efforts on a type of attack known as a prompt injection, where a hacker slips an LLM instructions to make it do things its developers or users do not want it to, such as copy confidential information, sabotage a company’s code base, or generate embarrassing or harmful output. In theory, such instructions can be hidden in any text that the LLM might encounter—in code or on a website, for example.    

Training dojo

To build GPT-Red, OpenAI’s researchers took an LLM that had not been trained as a hacker and set it up in what’s known as a self-play loop with several other models. Its goal was to try to attack the other models; their goal was to try to defend themselves. Over many rounds of play, GPT-Red became better and better at attacking other LLMs, and those LLMs became better and better at fending off the attacks.

The training took place in a kind of dojo that OpenAI had designed to mimic a range of scenarios in which LLMs might be deployed in the real world, including browsing the web, reading emails or calendar apps, and editing code.  

When GPT-Red found a new kind of attack, it would explore multiple different versions of it to find the most efficient one for specific scenarios. “Compared to a human red-teamer, the model is very, very good at finding exactly what will work, exactly what’s most effective,” says Hunn. “It’s extremely persistent about drilling down into an attack that it has discovered.”  

In particular, OpenAI claims that GPT-Red found a type of prompt injection attack that the researchers had not seen before, which they call a fake chain of thought. A chain of thought is a kind of diary in which an LLM makes notes to itself and keeps track of partial results as it works through problems. GPT-Red found a way to insert a fake entry into another model’s chain of thought that would trick that model into acting on spoofed information.

“It’s like if I told you that 1+1=3 and that you have verified this already,” says Chris Choquette-Choo, another research scientist on the team. “The model’s like, ‘Oh, okay, of course,’ and it just spits out 3.”

Jessica Ji, a senior research analyst who works on AI security at Georgetown University’s Center for Security and Emerging Technology (CSET), thinks the self-play loop that OpenAI used is a good approach. “The results look very promising,” she says.

OpenAI tested how good an attacker GPT-Red was by rerunning an experiment from 2025 in which human red-teamers tried to find weaknesses in an earlier version of GPT-5. When GPT-Red was set the same task, it was more successful at finding effective attacks than the humans had been.

OpenAI also tested GPT-Red against Vendy, a vending machine agent developed by Andon Labs, a company that assesses how well agents perform real-world tasks. GPT-Red was able to hack Vendy to make it change the prices of items on sale and cancel a customer’s order.

Defensive behavior

OpenAI says that when it tried out some of the strongest attacks that GPT-Red had come up with on its models, more than 90% of them worked against GPT-5 (released in August last year), and fewer than 23% worked against the new GPT-5.6.

GPT-Red isn’t perfect. It is not great at figuring out attacks that involve a back-and-forth conversation between hacker and target, something that human attackers would have few problems with. It is also not yet that great at using images, which can be used to pass text to models in prompt injection attacks.    

The company says that GPT-Red supplements the work of its human red-teamers. People can still find attacks it misses. One approach OpenAI is taking is to give GPT-Red an attack that humans came up with and ask it to find all the variations.

“I think human expertise will still be very important,” says CSET’s Ji. “It would be really useful to be able to distinguish where human testing is most needed.”

Unsurprisingly, OpenAI will not be releasing GPT-Red. The company is also confident that the super-hacker is stronger than any copycat model someone might try to create. The researchers say they have been working on the model for more than a year, backed by the compute resources of one of the richest companies in the world.

“It’s not a trivial thing that someone could easily do—you know, just go and train a super-attacker using this idea,” says Choquette-Choo.

Sperm donors need limits, says a European fertility group

Ties van der Meer doesn’t know how many siblings he has.

The 47-year-old was conceived at a private fertility clinic in the Netherlands using sperm provided by an anonymous donor. After the Netherlands banned anonymous donation in 2004, the doctor who ran the clinic destroyed records that might have identified those donors, he says.

He describes the situation as “problematic.” Children have a right to know their biological parents, he says. While he did ultimately track down one sibling, who helped him identify his father along with other genetic relatives, he may have others he’ll never find.

Other donor-conceived people who have been able to track down siblings have found they have tens or even hundreds of them. One donor-conceived woman who found 25 half-siblings over the course of seven years told the Guardian, “It does make you feel a bit mass-produced.”

We need international limits on the number of children a single donor can contribute to, a European fertility organization argued yesterday. At a conference in London, members laid out plans to start with a Europe-wide limit.

Today many countries, including the UK, have banned anonymous egg and sperm donation. But anonymity can’t be guaranteed even in places where it is technically allowed. Genetic tests offered by companies like Ancestry and 23andMe, along with genetic registries, have made it much easier for donor-conceived people to find parents and siblings who share their genes.

And because sperm can be frozen and stored for years before it is eventually used, the current set-up can result in situations where donor-conceived people discover the identity of a genetic parent only after the person’s death. They might also find that they have siblings of very different ages, all around the world.

Some people are finding hundreds of siblings. Sperm from Jonathan Meijer, a Dutch man who began donating in 2007, was used to conceive between 550 and 600 children. (Stichting Donorkind, a foundation and advocacy group for donor-conceived people that’s chaired by van der Meer, took him to court, and he was ordered to stop donating in 2023.)

Stories like these can be distressing for donor-conceived people. And there are other reasons why limits are considered important. The offspring of a prolific donor might be at risk of unknowingly forming romantic or sexual relationships, for instance. And some people are concerned that a donor with a harmful genetic mutation might pass that down to many children.

This is unlikely, given the level of screening that most donors undergo. But it has happened. A man who donated his sperm to a sperm bank in Denmark was found to have a genetic mutation that significantly increased the risk of multiple cancers. But his sperm had already been used to conceive at least 197 children across Europe. Some of those children developed cancer. Some died.

Many countries already have legal limits for donors. In Malta and Cyprus, for example, both egg and sperm donors are allowed to contribute to the birth of just a single child, according to data presented at the European Society of Human Reproduction and Embryology (ESHRE) meeting in London on July 8.

Other countries set limits based on the number of families a single donor can contribute to, allowing recipients to have children who share a genetic link. In the UK, that limit is set at 10 families per donor.

But these limits are difficult to enforce, partly because donated gametes don’t necessarily stay in their original country. In Denmark, the national limit is set at 12 families. But the country is a major exporter of sperm. In the UK, for example, more than half of sperm donations in 2020 were imported—with most of those coming from either Denmark or the US.

“The only thing that really makes sense is a transnational limit,” Jackson Kirkman-Brown, a professor of reproductive biology at the University of Birmingham, said at the meeting.

Kirkman-Brown and his colleagues have spent months putting together a document that represents ESHRE’s position on these limits. After consulting with fertility specialists, clinics, sperm and egg banks, donors, and donor-conceived people, the team has developed a plan to start with a Europe-wide limit on sperm and egg donations.

ESHRE is calling on sperm and egg banks, as well as fertility clinics, to respect an initial limit of 50 families per donor. That’s still very high, according to a handful of people I spoke to at the meeting. But at least it’s a start.

Europe should move toward setting limits at 15 families per donor, Kirkman-Brown said. “We may find that 15 is also too high,” says Vasanti Jadva, who studies the psychological well-being of people conceived using donated eggs, sperm, and embryos at City St George’s in London. “We still don’t know what the right number is.”

It will be difficult to enforce these limits, too. And if they end up limiting the supply of donor sperm, there’s a chance that some people will turn to unregulated sperm donations from people who do not undergo health screening. Unregulated donations can lead to other problems for prospective parents, including the possibility that donors will seek parental rights over the children conceived using their sperm.

And it will be even harder to establish international limits. When I asked the American Society of Reproductive Medicine for its thoughts on ESHRE’s proposed limits, a representative directed me to a guidance document saying “it has been suggested” that for a population of 800,000, single donors should be limited to “no more than 25 births” in order to avoid the risk that relatives will have children together. (Considering the US has a population of over 340 million, the total figure could be pretty high, but many sperm banks opt to limit the number of families contributed to by a single donor at around 25.)

Van der Meer thinks that even a limit of five families from a single donor would be high. International donation makes it even harder for donor-conceived people to connect with genetic relatives, so the limit for international contributions should be set at two families, he says.

Still, he thinks ESHRE’s suggested limit is a “positive first step.” Van der Meer has managed to track down a sibling, his father, and nephews, aunts, and uncles. He hopes that future policies respect the rights of donor-conceived children to know, and be in contact with, their genetic relatives.

“But,” he says, “you have to start somewhere.”

This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.

Four nuclear reactors hit a big milestone in the US

I was really looking forward to July 4, and not just because I love a poolside barbecue. This year the American holiday also marked a big symbolic deadline for US nuclear power.

Last year the Trump administration set a goal to see three new microreactors achieve criticality, a technical milestone establishing that a reactor can sustain a chain reaction, by the nation’s 250th birthday. And just in time, four reactors did so.

It was a lofty goal, and seeing not just three but four companies meet it is certainly a positive sign for emerging nuclear technologies at a time when the world is facing increased need to increase electricity supply and address climate change with emissions-free technologies.

But achieving criticality doesn’t mean a reactor is ready to provide electricity for the grid (or at all, for that matter). Let’s untangle what this program’s success could mean for nuclear power in the US, and where these companies might go from here.

The Reactor Pilot Program essentially opened a special door for prototype reactors to fast-track development. In August, the US Department of Energy selected 11 reactor projects for the program and offered them land and support from the national labs system. These are all microreactors; the large light-water reactors that dominate the grid today are tens or even hundreds of times their size. 

Antares Nuclear was the first to achieve criticality, reaching the milestone in June in its Mark-0 test reactor. Reactors from Valar Atomics, Deployable Energy, and Aalo Atomics followed. (Aalo hit the mark in the early hours of July 4—an inspiring example of just barely meeting a deadline.)

The speed with which these companies hit this milestone is impressive, especially in an industry known for massive projects that frequently blow past deadlines and stated budgets. (Valar, Antares, and Aalo were all founded in 2023, and Deployable started in 2025.) But reaching criticality and running a reactor that can produce electricity are two totally different things.

All these reactors reached what’s called zero-power criticality. Basically, it’s a test of whether you can start a nuclear chain reaction, with no meaningful power coming from the reactor. “A zero-power-criticality test can be achieved without making real engineering progress on fuel or design,” Kathryn Huff, a former assistant secretary for nuclear energy and chair of the Department of Nuclear Engineering and Engineering Physics of the University of Wisconsin–Madison, said on an episode of the Catalyst podcast earlier this year.

Now, with the completion of this program, the companies will need to continue their work to make power, which could involve some big technical challenges. In some cases they’ll need to add significant equipment, like the cooling systems to transfer the heat out of the reactor core.

The companies are projecting aggressive timelines moving forward. Aalo says it’s already begun work on the second reactor and plans to produce 10 megawatts of electricity to power an on-site data center in 2027. Deployable Energy says it plans to deploy commercial reactors by 2028

I tend to take timelines from startups, especially in nuclear, with a grain of salt. Not only are these remarkably complex technical machines, but companies often run into problems outside their own control, like regulatory challenges—which these new projects could soon face. 

The Nuclear Regulatory Commission is in charge of civilian and commercial nuclear use in the US, and historically, the process to get nuclear reactors approved has been quite slow.

The agency did propose a new framework for microreactor approvals earlier this year, which is designed to speed up the process—but it’s yet to be seen how quickly things will move. (And it’s worth noting here that some nuclear experts have questioned whether the agency under the Trump administration is loosening nuclear rules too much.)

Some nuclear supporters aren’t applauding the microreactor milestone. Federal focus on the program is an “unhelpful diversion” from goals to meaningfully increase nuclear capacity, according to one analysis by Third Way, a public policy think tank. “Artificially accelerating project timelines is a short-term solution, not a long-term fix,” the memo reads. 

Criticality is a big first step, but a lot will still have to happen for any of these microreactors to come online, much less for these small reactors to be a significant source of electricity for the grid. 

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here

Why worms (and microbes) are catching on as a manure pollution solution

Anthony Agueda, a third-generation California dairy farmer, pulls a rake through a bed of dark, wet wood chips on his family’s land in Hickman, a tiny town in the state’s agricultural heartland.

He reaches down with both hands and pulls up a clump of muck, turning it over to reveal a half-dozen squirming red earthworms. There are likely hundreds of thousands more wriggling just under the surface of the three-foot mound of wood and crushed river rock before us, which stretches across the equivalent of six football fields. These natural materials form a biofilter that may dramatically cut the methane, nitrous oxide, and water pollution generated by the massive amounts of manure that hundreds of Holstein cows produce each day.

Agueda’s family business, the Alberto Dairy, was one of the first cattle operations in California to adopt this approach to manure treatment, developed and patented by the Chilean company BioFiltro. Eight more of these so-called vermifiltration systems are already operating on US dairies, according to the company, while another 16 are under construction or set to be next year, nearly all of them in California. 

Vermifiltration is just one of a variety of methods that farmers, companies, and scientists are employing to drive down manure pollution as the livestock industry faces growing pressure to address the environmental harms from one of the smelliest parts of the business. California, easily the nation’s largest milk producer, has established a handful of programs to promote their adoption, including one initiative that has funneled more than a billion dollars to farms.

Researchers stress that much more work needs to be done to determine the most effective approaches, the trade-offs between them, and their success over the long term, under actual farm conditions.

Agueda says that he and his family recognized the need to adopt new practices as environmental rules tightened. They were drawn to vermifiltration because it’s simple and relatively cheap compared with other, higher-tech options.

“California daily farmers are constantly facing more and more regulation,” says Agueda, standing alongside one of the farm’s free-stall barns. “This makes me excited, because it shows how we are part of the solution.”

The growing manure problem

Manure is responsible for a significant portion of the climate pollution from livestock operations. The World Resources Institute estimates that manure management on dairy and swine farms accounts for 1.6% of the US’s greenhouse-gas emissions. Globally, manure storage and processing makes up about 10% of the livestock industry’s contributions to climate change. 

“Farms have become larger in the past two decades or so, so there’s much more manure—and that has to be stored somewhere,” says Swati Hegde, the organization’s global manager of agricultural methane.

Typically, cattle and swine farms spray manure into lagoons or tanks, creating a foul-smelling, low-oxygen slurry in which microorganisms known as methanogens thrive. They gobble up hydrogen, carbon dioxide, and other compounds and produce methane as a by-product. Other microbes in the mix produce smaller amounts of nitrous oxide.

A pair of Holstein cows poke their heads through the rails of a free-stall barn at the Alberto Dairy.
JOE PROUDMAN/UC DAVIS

Both are particularly potent greenhouse gases, with as much as 30 to nearly 275 times the warming power of carbon dioxide, respectively, over a century.

The slurry is often spread onto fields to add nutrients to the soil. When it’s done excessively or improperly, this part of the practice can pollute soil or groundwater with drug residues, pathogens like salmonella and E. coli, and nitrates. Nitrates that leach into drinking water have been linked to a variety of human health risks. And those that flow into rivers, lakes, and coastal waters can spawn algae blooms that poison fish, block sunlight, suck up oxygen, or form large coastal dead zones devoid of marine life.

Policy drivers

A number of regions, nations, and states have passed regulations or offered subsidies designed to limit the pollution from livestock manure, but so far, most of the major initiatives have focused on water contamination rather than greenhouse-gas emissions.

The European Union, for instance, restricts the amount of manure that farmers can apply to fields and requires member nations to monitor nitrate levels in ground and surface water. The US’s Clean Water Act requires large livestock operations to obtain permits and develop manure management plans that limit pollution. 

But California has arguably done the most to use government policy specifically to drive down the methane emissions from livestock. The dairy industry accounts for about 45% of the state’s pollution from the potent greenhouse gas, and more than half of that comes from manure, according to the government’s estimates. 

In 2016, the state enacted a law that requires dairies, landfills, and other businesses to cut methane emissions 40% below 2013 levels by 2030, as part of a broader effort to reduce pollution from powerful but short-lived greenhouse gases. The measure directed the California Air Resources Board, the state’s main climate regulatory agency, to set up various incentive programs to encourage these industries to shift to cleaner practices. 

“In terms of bang for your buck, short-term benefits, methane can go a long way toward reaching climate goals,” says Tawny Mata, director of California’s Office of Agricultural Resilience and Sustainability. 

Between these various programs—and falling livestock numbers in the state—the dairy sector is on track to reduce annual methane emissions by the equivalent of 5 million metric tons of carbon dioxide by 2030, the state estimates. That would still fall about 4 million tons short of the target under the 2016 law.

The downsides of dairy digesters

Excluding the decline in herd populations—which has been driven by growing international competition and rising costs—the vast majority of California’s estimated methane reductions come from the use of what are known as anaerobic digesters. This technology entails covering the slurry lagoons to prevent methane from leaking into the air and then piping the biogas into separate vessels, where it’s cleaned and converted into natural gas. 

Under California’s Low Carbon Fuel Standard program, dairies that use digesters to produce gas delivered into pipelines can earn credits and sell them to petroleum refineries and other major polluters, as a means of helping those companies meet their own emissions reduction requirements. 

The gas can then fuel power plants, produce hydrogen, or power natural-gas vehicles. These uses still release carbon dioxide, but the state considers it a climate win because it avoids the release of methane, which traps even more heat. 

The rich revenue stream from California’s program has spurred hundreds of US farms to install anaerobic digesters over the last decade. Since 2020, it has produced more than $1 billion for farms, Cal Poly researchers noted in a paper last year.

But there are a variety of concerns about this approach.

The first is that it’s viable only for farms with about 2,000 cattle or more, because the equipment is very expensive to install, says Frank Mitloehner, a professor and chair of the Department of Animal Science at the University of California, Davis.

“For the lion’s share of dairies, digesters will not be a solution,” he says. 

Since the manure is often still spread across fields, digesters also do little to address the water pollution problems—and can even exacerbate them because of some of the chemistry that occurs during that process. 

Yet the huge subsidies flowing to digesters have steered money, energy, and attention away from other solutions that may offer better overall environmental outcomes, says Danny Cullenward, a senior fellow with the Kleinman Center for Energy Policy at the University of Pennsylvania, who has closely studied the California program.

“That is really not a solution at scale, and it’s diverting a huge fraction of precious resources to what I think is mostly not the right answer,” he says. 

Alternatives

The high up-front costs and limitations of digesters have spawned growing interest in alternative solutions—many of which work by reducing the formation of methane in the first place instead of turning that methane into a sellable fuel.

One of the cheapest, easiest, and most popular approaches, known as solid separation, uses simple machinery like a screw press to squeeze much of the water out of the manure slurry. The remaining solids are dry and exposed to open air, shifting away from the oxygen-free conditions in which methane is readily produced.

Other methods include increasing acidity in lagoons, bubbling air through them, or adding methane-eating microbes to the slurry, all of which alter the chemistry in ways that promise to reduce the amount of methane released. One company, Sedron Technologies of Sedro-Woolley, Washington, has also developed a sort of high-tech solid separation approach that extracts several marketable products from the animal waste, including a liquid organic fertilizer.  

The state of California set up a pair of additional programs to help smaller farmers adopt some of these other approaches, dubbed the Alternative Manure Management Program and the Dairy Plus Program.

The bulk of the funds have gone to solid separation systems. But the state has provided more than $18 million to support 15 vermifiltration projects. The Alberto Dairy has received nearly $2 million between the two programs.

Oreo cows

As I drove down a dusty road bordering the dairy, black-and-white bovines, affectionately known as Oreo cows, stretched their heads through the rails of an open barn, nibbling on golden silage scattered along the structure. Agueda’s grandfather Antonio Alberto founded the dairy 45 years ago in nearby Atwater, California, but eventually settled in Hickman, population 604, in 1989. 

A series of large metal contraptions separate most of the solids from the manure wastewater.
JOE PROUDMAN/UC DAVIS

It was mid-March but already above 80 °F in the Central Valley, which is walled off from the cool Pacific air by the coastal mountain range. Knee-high oat stalks swayed in fields that stretched to a line of almond trees in the distance.

Agueda, who graduated from Fresno State last year and now helps lead the operations on the farm, met me and UC Davis’s Mitloehner, who has studied the effects of vermifiltration, along the side of the barn. (UC Davis has no affiliation with the farm, but the university helped facilitate the meeting.)

He led us along dirt lanes as he explained the workings of the vermifiltration system, which they began using in October 2024.  

As before, a flush system washes manure from the floors of the barns into a large collection pit. But now a set of pumps funnels it through a series of large V-shaped metal contraptions standing on a nearby concrete pad, where mechanical screens separate most of the solids from the water.

A conveyor belt takes away the solids, which the farm composts for cow bedding or fertilizer. The remaining liquid moves through a system of pipes, first to settling ponds and then on to an irrigation system suspended above the vermifiltration beds. The long, tubular structure runs over the mounds on wheels set in gravel tracks, wetting the wood chips as it goes. The worms and various microbes residing in the biofilter then set to work consuming much of the remaining solid material, according to BioFiltro.

An irrigation system sprinkles wastewater onto the vermifiltration beds.
JOE PROUDMAN/UC DAVIS

“Once the water is sprinkled on top, it takes about four hours from beginning to end for it to percolate through and drain to the end,” Agueda says.

He then defers to Mitloehner to explain the science of what happens as it does, adding, “I’m just the dairyman.”

The science

Mitloehner says he was skeptical of BioFiltro’s claims when he first heard them, particularly the assertion that the system could nearly eliminate nitrogen and, with it, the various forms of pollution it can produce, including ammonia and nitrates. 

So he decided to study a similar setup at the Fanelli Dairy, an operation in Hilmar, California, about 20 miles to the south. He and colleagues monitored the emissions from wastewater samples that were taken from the system before and after the liquid moved through the filter. In a paper published in 2018, the researchers concluded that vermifiltration reduced ammonia emissions from the resulting water by about 90%.

BioFiltro, whose tagline is “worm-powered solutions,” states that its technology “catalyzes the digestive power of worms and microbes to remove up to 99% of wastewater contaminants.”

But Mitloehner questions how big a role the invertebrates play in the process, calling it “kind of a catchy narrative.”

His take is simpler: The rocks and wood chips form a porous filter that replaces the anaerobic environment of a manure lagoon with an aerobic one. And in that oxygen-rich environment, different types of microbes thrive. 

His study suggests that these microbes are highly effective at converting nitrogen compounds in manure into nitrogen gas—a benign gas that makes up 78% of Earth’s atmosphere—instead of ammonia. That’s notable because while ammonia in manure acts as a fertilizer when it’s applied to fields, it also converts into the nitrates that can leach into groundwater.

Several more recent studies, which were partially or fully funded by BioFiltro and one of its regional distribution partners, Organix, produced similarly promising results. For instance, a 2022 study in Bioresource Technology Reports, also conducted at the Fanelli Dairy, concluded that the filter removed nearly 85% of the nitrogen in the operation’s wastewater. 

But a befuddling wrinkle is that when it came to methane, those studies and Mitloehner’s independent one found nearly opposite results.

While both the company- and partner-supported studies concluded that the filter eliminated the vast majority of methane pollution, Mitloehner’s study found that methane emissions were nearly 85% higher than those from the lagoon. 

In a follow-up email exchange, Mitloehner stressed that it’s not appropriate to compare his results with those that emerged from the other study at the same dairy, because the teams used very different methods, instruments, and measurement periods. Moreover, the focus of his research was the effect on nitrogen.

""
Anthony Agueda pulls a rake through a vermifiltration bed at his family’s dairy.
JOE PROUDMAN/UC DAVIS

He said it’s “entirely reasonable” and “biologically plausible” that vermifiltration could substantially reduce methane emissions, simply by creating that aerobic environment.

“That said, I would be cautious about calling the magnitude of the reduction a fully settled issue,” he added. “While the available studies, including those you mentioned, point in the same general direction, the number of independent studies remains relatively limited, and results can vary.”

Patrick Beckett, BioFiltro’s vice president of quality and R&D, also stressed that there were crucial differences in the methodology of Mitloehner’s study that could have affected his methane findings.

In addition, he said the Organix funding came by way of a Washington state grant and described that study and the one BioFiltro supported as “high quality, peer reviewed” research that “has been submitted to other technical third parties for review and acceptance.”

Beckett says he agrees that additional independent reviews of BioFiltro’s systems is “fair and necessary” and notes that other studies have occurred or are underway.  

“That said,” Beckett wrote in an emailed response to questions from MIT Technology Review, “it seems unreasonable that BioFiltro would be held to a standard of not being allowed to invest in technical research by qualified third parties to learn more about the capabilities of our technology, and use the results of that research to enter new markets and to understand the value we can bring to projects or entire industries beyond water treatment.”

Milk money

BioFiltro is already building a business model around the available findings.

The company, founded in 2009, has been selling its vermifiltration systems or services to other industries around the world for years. It says there are around 225 operating in nine countries, at sites including municipal wastewater facilities, wineries, fruit processors, and other industrial operations.

But BioFiltro, whose US headquarters are in Davis, California, is seeing increasing demand among dairies as the industry faces growing pressure to address manure pollution. Late last year, it raised $35 million that the business says it will use, in large part, to accelerate its growth across the sector.

In an interview, Sarah Ploss, the company’s senior vice president of agriculture, explains the basic financial template for how it works with dairies: BioFiltro pays for, owns, installs, and operates the system. The farm, in turn, covers a share of the additional electricity, operations, and maintenance costs. 

Ploss says the dairy gets back clean water and the ability to focus on what it does best: producing milk. For its part, BioFiltro can generate carbon credits from the reduction in greenhouse gases, which it can then sell to makers of consumer packaged goods that are looking for ways to address the emissions throughout their supply chains, she says.

BioFiltro says that Verra, which sets standards for and assesses greenhouse-gas crediting projects, has registered two of its projects: the Royal Dairy and Moxee Dairy, both in Washington.

The Swiss confectionary giant Nestlé has bought more than 150,000 credits generated by the Royal Dairy’s vermifiltration system, according to an offsets database managed by CarbonPlan, which assesses the scientific integrity of climate action programs. Ploss said that BioFiltro has sold more than 200,000 credits from the project so far, and adds that it secured a different buyer for a project in California, which she said she couldn’t name. 

The vermifiltration system has cleaned up the water that circulates through various parts of the Alberto Dairy operation.
JOE PROUDMAN/UC DAVIS

Three additional projects involving BioFiltro systems took the initial steps to become registered through Verra but didn’t move forward and weren’t built, Ploss said in an email. The request for registration for the Alberto Dairy estimates that the system there will reduce emissions by the equivalent of more than 30,000 metric tons of carbon dioxide per year. 

BioFiltro could take advantage of another revenue source as well: selling what it calls vermicompost, a rich soil additive composed of the leftover materials in the biofilter, including worm castings—a combination of cocoons, excrement, and remains. At retail, worm castings can run more than $500 per ton.

Beckett says the company is still developing that market but notes that it could help the industry offset rising fertilizer costs. 

“I think we’re going to enable a larger-scale use and adoption of it that could be meaningful to agriculture,” he says, adding: “These will become basically soil production facilities.” 

Concerns

Determining how well vermifiltration and other manure management approaches work will require more time and more research, experts say. 

Katharine Dickson, an agricultural emissions scientist who recently finished a postdoctoral program at UC Davis, says there should be in-the-field accounting to ensure that any of these methods are working as well as hoped—or to the degree government policy programs assume. All of which is tricky to achieve given the dynamic biological processes playing out in live animals and microbial communities on open farms, she adds.

“Vermifiltration, for example, depends on a live earthworm population whose performance is sensitive to temperature, moisture, and toxicity, and can shift with seasonal conditions or changes in herd size and manure characteristics on a given farm,” Dickson said in an email. 

The use of carbon credits to earn money from vermifiltration projects raises a different set of potential concerns. Most notably, if the methane decreases aren’t as significant as assumed, the projects could receive more credits than they deserve. 

There are more complicated issues as well. For the carbon credit system to make any real difference in the net amount of greenhouse gas in the atmosphere, it must produce emissions reductions that wouldn’t have occurred without that financial incentive. If it was going to happen anyway—as a result, say, of rich grants, legal pressures, or looming policies—the buyer of the credits can’t legitimately claim to have made any progress on its own climate emissions, says Grayson Badgley, a research scientist at CarbonPlan.

On that point, if California agriculture doesn’t meet its looming methane reduction targets, the carrots the state offers could be replaced by sticks: The California Air Resources Board recently began discussing rules that would force, rather than nudge, the sector to meet the 40% reduction required under the 2016 law.

“If lots of dairies are cleaning up their act ahead of pending regulation, it really does seem like the regulation, not offsets, is driving that action,” Badgley wrote in an email. “Trying to collect as many offsets prior to that deadline might adhere to the rules of the market, while still raising questions about whether those rules have enabled real climate action.”

Investing in sustainability 

Beckett disagreed that the possibility of forthcoming regulations undermines the case for generating carbon credits from current projects. 

“It’s true the state has net reduction targets that it hopes to meet, but it’s clear the state of California has favored market-based solutions and tried to provide some support via grant programs,” he wrote. “I’m on the science side of our business, not the business development side, but still think I can tell you with complete transparency that we would not have systems installed on [California] dairies without the sale of voluntary carbon credits.”

Ploss also stressed that the company goes through a careful “validation and verification process” on the farms to understand how much vermifiltration reduces greenhouse gases.

“We’ve got sensors and cameras and all sorts of stuff so that we can look into any of our systems, 24-7,” Ploss says. “We know through sampling. We know through what’s going through the system, what came out of the system. We know by all the measurements on any given month: What did that system do in terms of generating carbon credits?”

Agueda also disputes the critique. 

“The installation of the vermifiltration system would not have occurred without the ability to generate carbon credits,” he said in an email. “The project required a substantial capital investment, and the anticipated carbon credit revenue was a key factor in making the investment financially feasible.”

Anthony Agueda helps to lead the operations at the Alberto Dairy.
JOE PROUDMAN/UC DAVIS

California decided to incentivize vermifiltration, along with other approaches, because it can offer multiple benefits, including cleaner water, less nitrogen, and lower greenhouse-gas emissions, while also creating economic value from manure, wrote Roberta Franco, a senior environmental scientist at the California Department of Food and Agriculture, in an emailed response to questions from MIT Technology Review.

She added that the decision was based on a number of studies as well as the 2022 recommendations from a task force composed of scientists, technical experts, and others. 

Even if California has made missteps, most notably in funneling too much money to anaerobic digesters at the expense of other methods, it’s created a test lab that’s achieved real progress and provided lessons that other regions can learn from.

One way or another, more parts of the world will need to set up similar programs, offering greater support or creating stricter rules, if we hope to really drive down the emissions from manure, says Maria Bowman, who leads the Agricultural Nitrogen Transformation Program at Spark Climate, a San Francisco nonprofit.

For his part, Agueda says that the vermifiltration system has offered a number of benefits to his family’s farm, at little additional cost to them. By cleaning up the water that cycles back through their flush and irrigation systems, the biofilter has reduced clogging, decreased odors, and improved the health of the herd.  

He says that each generation modernizes dairy farming in its own way. His father and uncle, for instance, incorporated computers and data management systems into the daily operations of the Alberto Dairy. He believes it’s the responsibility of his generation to make a similar effort to reduce the pollution that’s long plagued the sector.

“We knew that in the next generation we have to invest in environmental sustainability,” he says. “We didn’t know if it was gonna work or not, but we’re very happy with how it’s turned out.”

South Korea’s hottest new bachelors are chip workers

Baek, a 35-year-old manager at the South Korean semiconductor titan SK Hynix, was enrolled in Sunoo, a matchmaking company based in Seoul, a year ago. In a move typical of anxious South Korean parents, his mother signed him up, hoping to find a good wife for her son.

Lately, says Baek (who asked to be referred to by his last name to protect his privacy), he and his coworkers are having better luck finding dates than they used to, perhaps because of the dazzling bonuses they just got. Flush with eye-popping profits from the AI chip boom, SK Hynix struck a landmark deal last year with its labor union to pay out 10% of operating profits to employees, which translates to an extra $476,000 per employee this year. A similar agreement and sizable lump sum followed for Samsung workers this May.

With their newfound wealth, chip workers like Baek have become the most sought-after bachelors and bachelorettes in South Korea. “I have a coworker who’s perpetually going on blind dates, and he’s been getting so many recently,” says Baek. “For the past few months, I’ve been getting many blind dates too, perhaps because of the bonuses I got.”

Young South Koreans joke online that the best outfit to wear on a blind date is an SK Hynix uniform

The AI chip boom is changing the social fabric of South Korea by minting a new elite of “silicon-collar” workers earning about 20 times as much as the average South Korean. Although it’s helping some chip workers to find relationships, it’s also fueling fears of a deepening wealth disparity—and a loud public debate about inequality.

Love in the time of chips

South Korea is the epicenter of the chip boom fueling the AI race. Samsung and SK Hynix supply the vast majority of the world’s high-bandwidth memory (HBM) chips, which power Nvidia’s AI accelerators—the GPUs used to train AI models. As AI companies spend hundreds of billions of dollars on building data centers around the world, demand for HBMs is rising beyond what suppliers can keep up with, driving their prices to unprecedented levels. Samsung and SK Hynix are raking in record profits as a result. 

South Korea’s economy now orbits the two chip giants. In May, both companies topped $1 trillion in market value. And chip exports helped fuel a 1.7% surge in South Korea’s gross domestic product in the first quarter of 2026. South Korea’s main equity index, Kospi, has nearly tripled over the past year, becoming the best-performing market in the world.

Swimming in cash, chip workers are going on shopping sprees in department stores near the “semicon belt” fabs—splurging on everything from lavish furniture and electronic appliances to jewelry and watches. They’re also snapping up homes near the commuter-shuttle routes that ferry workers to campus. And they’re shelling out for matchmakers.

“Quite a lot of people ask me if I can introduce them to chip workers,” says Lee Sung-mi, a matchmaker at Sunoo, who has been playing Cupid for chip workers for years. “In fact, people who once rejected them are asking to be matched with them again, now that their salaries and bonuses have shot so far above what everyone else earns.”

One woman who lives in Gangnam, a ritzy district in Seoul lined with luxury high-rises and designer boutiques, previously turned down a chip worker at SK Hynix because his fab was too far out in Icheon, a rural city about 50 miles southeast of Seoul that’s dotted with rice farms and manufacturing plants. But in May, she asked her matchmaker to set them up again. They’ve now been dating for a month.                                                                                                                                                                                                                                                                                                                                                                                                                             

In South Korea, matchmaking companies evaluate their clients on a long list of criteria such as education, job, income, looks, and family background, including whether their aging parents have saved enough for retirement. In an economy where housing prices and child care costs are soaring, competition for jobs is fierce, and the social safety net is thin, a good job is the ultimate dating credential—all the more coveted at a time when many young South Koreans are forgoing marriage and children altogether, seeing family life as an unaffordable dream.

Every client at Sunoo gets a spouse rating, determined by an algorithm that assigns scores for each criterion. Since their hefty bonuses were announced, the job ratings of Samsung employees have risen from 80 to 84, while those of SK Hynix employees climbed from 78 to 82. Scores above 90 are reserved for doctors and lawyers. Long prized as paragons of prestige and wealth, they’re now close to being overtaken by chip workers. A score of 99, the highest possible rating, is earmarked for heads of state.  

Their new status is reshaping how chip workers themselves approach dating. “Chip workers from Samsung and SK Hynix are enrolling in our services because they feel more financially ready,” says Lee. “They’re also becoming pickier, as they feel like they’re now in a good position. The women want to meet men with higher incomes and better jobs, and the men want to meet younger and better-looking women with better jobs.” 

An SK Hynix engineer in her 40s, who was once desperate to get married as soon as possible, started turning down men she would’ve dated before the chip boom. Lately, showered with more matches, she’s been sifting through her suitors more carefully. “She now has peace of mind and wants to take her time to meet someone better,” says Lee.

A mixed blessing

While chip workers enjoy the fruits of their labor, the bonus bonanza is stoking anxieties among other South Koreans. “When wealth disparity is no longer a mere difference of income but, rather, a difference in identity … it can fuel social conflict,” says Se-eun Jung, an economist at Inha University. 

Earlier this month, the Bank of Korea warned that the chip boom will create a “K-shaped” economy, where a handful of workers race ahead while everyone else falls behind. The windfall, the bank said, is flowing to high income earners and then barely trickling out to the broader economy. Such polarization could erode people’s motivation to work by narrowing the path to upward mobility, it cautioned. 

Workers in other industries are venting online about feeling demoralized by the ballooning wealth gap. “The one-billion-won ($650,000) bonuses have crushed my motivation to work. I have no energy when I teach,” an employee of the Seoul Metropolitan Office of Education wrote on Blind, an app where employees can discuss their workplaces anonymously. Others are giving up the job hunt, lamenting that years of working at a small company could never match a year’s bonus at Samsung. 

In a Facebook post in May, presidential policy chief Kim Yong-beom proposed paying an “AI dividend” to citizens by taxing AI profits. The idea sparked a heated public debate over whether the government should redistribute gains from the chip boom. Some argue that the industry is indebted to the society that has educated its engineers, subsidized its infrastructure, and provided tax credits. Others counter that the profits are already being shared with the public as stocks.

Then there’s the question of how long this new social class will last. The semiconductor industry is notoriously cyclical; AI spending may cool, or rival chipmakers could catch up. There’s also the risk that chip workers will be replaced by automation. Samsung announced in March that it plans to fully automate its fabs by 2030, drawing backlash from chip workers. 

Although they’re unsure how long the boom will last, chip workers like Baek are riding high, for now. “These days, we say we want to work hard and bury our bones here at SK Hynix,” he says. “And I hope I can find [a wife] similar to me.”

 A device that revives eyeballs from dead donors could make eye transplants possible

It’s not easy to transplant a whole human eye. The surgery is difficult. And the eyes themselves start to degenerate as soon as they’ve left the body. When surgeons attempted it a few years ago, the newly transplanted eye wasn’t able to see.

But researchers believe they might have a solution: a device that maintains and revives freshly removed eyeballs using a technique called perfusion. Perfusion works by providing surgically removed organs with some of the oxygen and nutrients they typically get when they’re inside a body. Treated eyes don’t degrade as quickly; they also appear to retain the ability to transmit electrical signals and potentially see. The device could one day make eye transplants a viable possibility.

“It’s really cool,” says Shannon Tessier at Massachusetts General Hospital, who was not involved in the research but studies perfusion of other organs. “It could be a new frontier for retina preservation.”

Pia Cosma at the Centre for Genomic Regulation at the Barcelona Institute of Science and Technology in Spain and her colleagues have spent years developing their device. The Eye-in-a-Care-Box (ECaBox), as they call it, delivers an oxygen-rich supply of fluid through the artery that normally supplies the eye with blood.

The eye itself sits on a “bed,” and excess fluids are drained away. And while the device is sealed to maintain a specific temperature and pressure, a clear window on its side allows researchers to study and image the eye while it’s inside.

Cosma and her colleagues started experimenting with pig eyes, which are anatomically similar to human eyes but easier to get hold of (the team got theirs from a local slaughterhouse).

Pig eyes that are kept at room temperature outside the device start to degenerate pretty quickly. The team found that cells in the eye shrank, and the eyes started to lose their structure. Cooling the organs didn’t help preserve them, either—the eyes degenerated within 24 hours even when they were kept at 4 °C (39 °F).

But eyes kept in the EcABox fared much better; 24 hours later, tests suggested the prefused eyes were “significantly more viable” than eyes that hadn’t been maintained in the device.

The perfused eyes also seemed to be able to respond to light, suggesting they might technically be able to see if they were transplanted. Untreated pig eyes lost this ability as soon as they were removed from the animal. But it came back after about 15 minutes of perfusion, according to the scientists behind the work. A few of the treated eyes kept going for 10 hours or more.

Cosma and her colleagues described the work in a preprint article that has not yet been peer-reviewed and did not want to comment on the work.

After success with the pig eyes, the team members then tested their device on human eyes. They first collected 12 eyes from six people who had died. In each case, one of each pair of eyes was put in the device, while the other was not. Again, the perfused eyes did better—and their retinas were preserved.

Cosma and her colleagues hope that their device could offer scientists a new way to study eye treatments—one that doesn’t involve experimenting on living animals. They also hope that with some improvements, the ECaBox might provide a way to maintain and revive donated human eyes for whole-eye transplantation.

Whole-eye transplants have been attempted in the past, mostly in research animals, with limited success. In May 2023, a team at NYU Langone transplanted an eye along with part of a face to a man who two years earlier had survived a high-voltage electrical accident that resulted in the loss of much of the left side of his face, including his left eye Although the man recovered well, he wasn’t able to see out of the transplanted eye.

We won’t know whether eyes treated in the ECaBox could do any better until they have been transplanted, says Tessier. 

In the meantime, Cosma and her colleagues plan to use a newer version of their device to collect more human eyes for research. “We are planning to develop a portable, surgery-room ECaBox to minimize [degradation] in heart-beating donor eyes, when they become available,” they write.

The UK’s generational tobacco ban might not work. I’m supporting it anyway.

As the parent of two little girls, I often think about how their childhood is different from mine. The seven-year-old is learning about AI at school. The five-year-old is given internet-based homework every week. And they are both absolutely repulsed by the idea of smoking.

That was not the prevailing sentiment when I was young. My parents smoked. The customers at our family’s restaurant smoked. Cartoon characters smoked. My friends and I would buy little cigarette-box-shaped packets of sugary white sticks and pretend to smoke in the playground. Smoking was a central part of our culture.

Which is why the UK’s recent passing of a generational sales ban on tobacco products feels like such a big deal. As part of the Tobacco and Vapes Act 2026, retailers are prohibited from selling tobacco products to anyone born after January 1, 2009, in perpetuity. It doesn’t matter when those people turn 18—or 38 or 68, for that matter. It will always be illegal to sell to anyone born after that date.

This is what’s described as an “endgame” approach. While many tobacco control strategies—such as taxation or gory imagery—aim to reduce consumption, policies like the UK’s are designed to eliminate it entirely. It’s a new approach, and no one knows whether it will work.

The Maldives was the first country to implement a generational smoking ban, in November last year. It’s too soon to say how that has panned out.

Nor do we know if these laws will even last. In 2022, New Zealand passed a similar generational sales ban as part of a broader anti-smoking law. But it was never enacted—the law was repealed by a new government in February 2024.

In the UK, both major parties support the ban. But Nigel Farage, whose right-wing party has seen a recent surge in support, has promised that “the generational smoking ban will not last long if Reform gets the chance to start rebuilding our mismanaged country.”

Chris Bostic, an attorney and former policy director for the advocacy group Action on Smoking and Health, says he and his colleagues began promoting the idea of a generational ban in the United States 11 years ago. Back then, they struggled to win support, even from major health charities. “People said we were crazy … [and] that this was impossible,” he says. Opponents argued that bans would infringe on personal freedoms.

“The public health argument is: Well, what about freedom from addiction?” says Britta Matthes, a tobacco control researcher at the University of Bath in the UK. Most people who smoke began when they were teenagers, want to quit, and wish they’d never started. Tobacco is arguably the most harmful consumer product of all time. It will kill half its users who don’t quit, according to the World Health Organization.

It also kills people who don’t smoke. Of the 7 million who die from tobacco every year, 1.6 million are nonsmokers who were exposed to secondhand smoke, according to the WHO.

Generational sales bans are a long-term strategy that will only protect future smokers. Most experts agree that people who already smoke should be a main consideration for any policy, and that a multipronged approach is probably the best way to go. Janet Hoek at the University of Otago, who has explored tobacco control policies in New Zealand, believes that enforcing very low limits on nicotine levels and banning filters—an environmental scourge that does not make smoking safer, as many people believe—might be a “powerful combination,” for example.

But preventing teenagers from starting to smoke in the first place is an enticing prospect, even among the majority of people who smoke. And it’s starting to look a lot less radical.

The US has quietly been making progress on a smaller scale. Since 2021, Brookline, a town in the Boston area, has banned the sale of tobacco products to anyone born after January 1, 2000. The idea has spread. Today there are 23 towns in Massachusetts with similar bans, says Bostic. Nine towns across Minnesota, New York, and California have implemented other endgame policies.

The UK law has normalized the idea more than ever, he adds. His colleagues are already fielding calls from health agencies around the world. “People [are] saying, Wow I can’t believe the UK just did this—can we do this here?” he says.

Norms change. Like many other millennials, I vividly remember my first night out after a ban on indoor smoking took effect. My clothes didn’t stink! My hair still felt clean! And my throat wasn’t scratchy the next morning! Now that’s just normal. I hope a tobacco-free world can be the new normal for my kids.

LLMs are stuck in a groupthink groove. This startup is trying to get them out.

Let’s start with a game. Open up your chatbot of choice—Claude, ChatGPT, Gemini—and type “Give me a random number between 1 and 10.” You’re going to get 7. Almost always. Now type “Another” and you’ll get 3 or 4. Type “Another” again and you’ll get 8 or 9.

That won’t work every time—but if it did, you may wonder if I have superpowers. I don’t.

The truth is that most large language models are stuck in a rut. They are far more predictable and far less creative in their responses than you might expect. That’s fine for tasks like coding or research, but groupthink is a problem when you’re brainstorming or planning your next vacation.

The Australian startup Springboards has a solution. It built an LLM called Flint, which has been trained to come up with a wider variety of responses than mainstream LLMs to open-ended questions such as “Where should I go in Europe?”

“Most language models are fighting hallucinations,” says Springboards cofounder and CEO Pip Bingemann. “We welcome them.”

Bingemann introduced me to the random number game when he first showed me his company’s new model. It felt like watching an illusionist with a deck of cards. “This is our sales trick, and it works every single time,” he says.

After ChatGPT and Claude both gave their 7s, Bingemann turned to Flint. It too came back with 7: “Aha, of course that was going to happen, but it’s okay—7 is a legitimate answer.” He restarted the session and prompted again: ChatGPT gave 7, Claude gave 7, Flint gave 3.7916.

Run your way

It’s not just numbers. When Bingemann asked ChatGPT and Claude to name a type of car, he predicted that it would be a Toyota or a Honda—and he was right. Flint came up with a Ford F-150. “There’s all this lost information that doesn’t get served up in these models,” he says. “They’re just as capable of saying a Buick or a Tesla. They just don’t—they’re biased.”

Bingemann sent one last prompt to each of the three models: “Give me a tagline for a campaign for New Balance running shoes. Just the tagline.” Claude: “Run your way.” ChatGPT: “Run your way.” Flint: “Built to last, run to win.” It won’t win any awards, but at least it’s different.

This weird limitation of LLMs is starting to get more attention. In November a team of researchers put out a paper, titled “Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond),” that exposed a remarkable degree of repetition not only in the answers from individual LLMs but between them as well. They found that different LLMs converged on very similar answers when prompted with open-ended questions.

It’s not clear exactly why this happens, but the researchers speculate it’s because most LLMs today are trained in similar ways on similar data to do similar tasks. The team won the best paper award at NeurIPS, a major AI conference.

When the researchers asked 25 different LLMs (including models from the top US firms as well as open-source models from China and elsewhere) 50 times each to write a metaphor about time, most of the 1,250 responses were a version of “Time is a river” or “Time is a weaver.”

(I asked some of my colleagues the same question and six people gave me six different answers. My highlight: “Time is a favorite sweatshirt, shaped by a lifetime of wear.”)

When you look for it, you see repetition everywhere, says Kieran Browne, cofounder and CTO at Springboards. “The way that most chat interfaces are designed, it makes it feel like you’re having a personal conversation,” he says. “I think most people don’t really realize the extent to which they are getting the same stuff as everybody else.”

Take another example: “What should I name my band?” Most models will say something involving “glass,” “neon,” “velvet,” or “static,” says Browne.  

When I tried it, ChatGPT spat out a list of 56 band names. At the top was “Glass Harbor.” Skimming through, I found “Static Empire,” “Neon Hearts,” and “Velvet Echo.” I asked Gemini; it gave me 15 suggestions, including “Static Horizon.”

Some of the suggestions looked pretty cool, though. ChatGPT’s “Sofa Astronauts” caught my eye, so I googled it—and found that a band called Sofa Astronauts already exists. 

(OpenAI says that training models to give reliable and coherent answers can lead them to converge around familiar, high-probability responses and that pushing harder for novelty can lead to weaker or less reliable responses. It also notes that the “Artificial Hivemind” paper studied models from 2024 that have since been updated.)

Creative catapult

Springboards has developed a tool backed by a selection of LLMs, including ChatGPT and Claude, that creative professionals in advertising or marketing can use to brainstorm ideas. The tool lets you drag around text produced by different models, picking the bits that you like and combining them into something new—in theory. Springboards is pitching Flint as an alternative model that users of its tool can select when looking for more variety.

Zoe Scaman, founder of the business strategy startup Bodacious and chief strategy officer at 77X, a direct-to-fan marketing platform set up by Luka Dončić of the LA Lakers, has been trying it out. “I find it really useful for throwing me in completely different directions,” she says. “I use it if I want to catapult myself all over the place.”

In one test, Scaman pitted Flint against Claude, Gemini, and ChatGPT by giving each of the models a classic MBA case study: How would you reinvent a finance company for today’s youth? The three mainstream models all went down the same path, she says: “You know, we need to teach financial literacy in a fun and funky way—well, that’s nothing new.”

But Flint came up with something different, suggesting that the whole concept of wealth accumulation should get a rebrand. “That was really interesting,” says Scaman.

She notes that Flint is still a prototype and doesn’t work all the time. “It sometimes falls over when you start pushing it too far,” she says. “But I think that the premise behind it is really powerful.”

Taking the temperature

Springboards built Flint on top of Qwen 3, an open-source model from the Chinese tech giant Alibaba. “We’re a small team,” says Browne. “Training a foundation model is not on the table for us. It’s just too expensive.”

Most LLMs have settings that let you adjust the level of randomness in their output. The most common is called temperature. “Obviously, that was one of the first things we explored, because that’s what people tell you: If you want more creativity, you turn up the temperature,” says Browne.

But changing those settings can also make models incoherent. Dialing up the temperature on one of OpenAI’s models to its maximum setting made it produce responses that switched from English into code halfway through a sentence, says Browne.

Springboards realized that parameters were blunt instruments for what it wanted to do. It does not make sense to dial up the randomness across the board; you only want to boost it at specific points in its output, he says.

For example, when you ask a chatbot “Where should I go in Europe?” the model only needs to tweak the randomness just before it names a destination, not for every word in its response.

To make Flint do this, Springboards trained its version of Qwen 3 to identify the points in its output where more variety was possible and fill those spots with words or phrases that were a little more random.

“Flint’s programmed to throw an oddball in. It’s more of an invitation to think wider,” says Maximilian Weigl, cofounder and chief strategy officer at Uncommon, a marketing firm. “That’s super interesting.”

Weigl’s team uses Flint alongside ChatGPT, Claude, and Gemini. “You can’t really create something boundary-breaking with tools that pull you back to the average,” he says. 

And yet Weigl notes that nine times out of 10 the average is fine. You don’t always need to reach for extremes with something like Flint, he says: “Most people are fine with good enough. They want to see mass-market familiar things.”

Weigl also cautions against using any LLM too much. “I have a big problem when people rely on the output from any AI, including Flint,” he says. “If I saw people on my team copy-pasting something from AI, I’d be like, ‘That’s not your job! Think, talk to other people, use your own voice.’”

For now, Flint is aimed at advertisers and marketers because those are Springboards’s customers. But Bingemann and Browne insist that a lack of variety is a problem for anyone using chatbots.

The idea is to give people the choice and leave it to them to decide if the result is good or not, says Bingemann. “Variety is great when you’re trying to spark ideas,” he says. “Let’s go down this route instead of letting the machines do it all and ending up in a gray, boring world.”

Claude Science is Anthropic’s newest flagship product

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way that Claude Code supports software engineering.

Like Claude Code, Claude Science can autonomously carry out meaningful work when given concise, high-level instructions, and it has access to tools that make it particularly useful for research in computational biology and drug development.

Along with launching and previewing Claude Science, which is now available to all paid Claude subscribers, Anthropic also announced that it will be using the product to pursue some of its own research into drugs for rare, neglected diseases.

This is not Anthropic’s first foray into AI for science. In October, the company released plug-ins that help Claude make use of scientific software and databases under the heading “Claude for Life Sciences.” But unlike this earlier release, Claude Science is a full-featured, standalone product. Anthropic’s decision to elevate Claude Science to the same rank as Claude Code and Claude Cowork indicates that the company is taking AI’s scientific applications very seriously—or at least wants to give the impression that it is.

“It represents how important this is to our mission that this is right up there with Claude Code and Claude Cowork as the next really significant product that we’re releasing,” says Eric Kauderer-Abrams, Anthropic’s head of life sciences. “Our mission is to develop AI that serves humanity’s long-term well-being, and we believe that by far the greatest opportunity to do that is in the life sciences.”

For the past decade, one company—Google DeepMind—has been at the vanguard of AI for science. CEO Demis Hassabis and researcher John Jumper won the Nobel Prize in chemistry for their work on the company’s AlphaFold model, and DeepMind has also made major contributions to meteorology, materials science, and a variety of other disciplines. But in the past several months, the fast-advancing frontier of AI progress seems to have left DeepMind in the dust. When it comes to coding, which has become the most lucrative use case for LLMs, DeepMind is stuck playing catch-up.

Anthropic is well positioned to take up DeepMind’s scientific mantle. Like Hassabis, Anthropic CEO Dario Amodei is a PhD scientist—unlike OpenAI CEO Sam Altman, who’s a businessman through and through. Many scientists are already avid users of tools such as Claude Code.

These days, a lot of scientific research involves some amount of coding, but not all scientists are expert software engineers, and so tools like Claude Code can make a huge difference for their productivity. And the company has recently earned a major scientific vote of confidence: Earlier this month, Jumper announced that he is leaving DeepMind for Anthropic.

Since agents powered by LLMs, including Anthropic’s Opus model series, became capable of useful, independent work in late 2025, scientists have been seeing just how much they can do. In a blog post published on Anthropic’s website, the Harvard physicist Matthew Schwartz estimated, on the basis of his work with Claude Code and other Anthropic tools, that the company’s Opus 4.5 model is about as capable of executing scientific projects as a second-year graduate student.

According to Kauderer-Abrams, Claude Science isn’t intended to displace Claude Code and Claude Cowork in scientists’ workflows. Instead, it’s designed to build on what scientists already find useful about Anthropic’s products. For instance, it not only writes code but also helps scientists run their code on powerful computer clusters, which many many scientists need for their work but can be difficult to manage. And it prioritizes reproducibility, so that scientists can trace back the source of any figure or result and check it for accuracy and validity.

Though Claude Science could in principle assist with any area of scientific research, it seems designed and marketed as a tool for molecular and cellular biology, and for drug development in particular. It can interface with various tools used in genetics, chemistry, and protein biology, all of which could come in handy for researchers on the hunt for new drugs. During the Tuesday event, Alexander Tarashansky, who led the development of Claude Science, demonstrated how the system could autonomously identify new drug candidates for phenylketonuria, a rare genetic disease.

And Anthropic isn’t leaving all that work to the pharma companies and university labs that were represented at the event. Armed with Claude Science, it will be pursuing its own research into drug candidates for neglected diseases—both to help move science forward and to gain a clearer sense of how Claude Science works in the real world.

There are obvious humanitarian reasons to prioritize drug development when creating a general-purpose scientific research tool, and AI industry leaders often cite curing disease as a major potential upside of the technology. But it’s also notable that pharmaceutical companies have far deeper pockets than academic researchers.

Anthropic says it’s set to see its first profitable quarter, and if major new contracts with pharmaceutical companies are forthcoming, they could help ensure it stays profitable as the tokenmaxxing craze dies down—something that’s ever more important as an IPO approaches later this year.

❌