Reading view

There are new articles available, click to refresh the page.

How AI helps scientists design the next generation of medicines

Designing and developing a new medicine is an expensive, failure-prone scientific challenge. A new drug can take many years to develop, at the cost of a significant investment. And even then, most possible candidates never reach the patient. For biologic medicines, therapies made from engineered proteins rather than synthetic chemistry (which are often used to treat conditions across most major acute and chronic diseases), the complexity is even greater.

Scientists explore vast quantities of possible molecules, looking for the rare few that will bind to the right target, remain stable in the human body, and be manufacturable at scale. Today, AI is speeding up these processes and has quickly become a core part of the infrastructure in pharmaceutical R&D.

AI-assisted design is a growing part of how biologic drug candidates are developed, and companies like AstraZeneca are actively building its engineering teams to push this further. “Everything we do, whether it’s design, make, test, or analyze, is now computationally enhanced,” says Puja Sapra, senior vice president and head of R&D biologics engineering and oncology targeted discovery at AstraZeneca. “The cycle times are getting shorter while productivity and innovation increase.”

Sapra explains that AstraZeneca’s approach follows a build-measure-learn loop. AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed. Scientists then focus lab resources only on the top-ranked candidates. This leads to a tighter feedback cycle with fewer dead ends, faster iteration, and the ability to go after disease targets that were previously considered untreatable by medicine. Because the number of possible molecular combinations far exceeds what any human team can systematically explore, using AI to narrow and refine the options for testing has become a major focus in biologics drug design.

Navigating complex drug design problems

Beyond accelerating timelines, AI is also being applied to the discovery of entirely new classes of medicines. Traditional biologics typically target one disease pathway. The next generation of drugs can hit multiple targets simultaneously or precisely deliver therapeutic payloads to specific cells. Achieving this requires optimization across many variables at once. Looking ahead AI-driven models could help design these increasingly complex, multi-specific biologics, explains Puja Sapra. “For example,” she continues, “such models could help identify which two or three targets to prioritize based on the underlying biology, then optimize across multiple parameters to balance a molecule’s potency, stability, manufacturability, and safety.” “Drugging the undruggable is becoming a reality,” Sapra says. “These technologies will eventually enable us to develop medicines against targets once thought impossible to reach. The potential for benefit to patients is remarkable.”

The data moat

McKinsey estimates that generative AI, combined with other computational tools, could cut drug discovery timelines by as much as 50%. But every AI model is only as good as its training data. In drug discovery, that means ample quantities of high-quality biological data. Experiments can provide a rich source of such data. Whether they succeed or fail, each experiment generates a signal about what does and does not work.

“Data is our differentiator,” says Sapra, explaining how the company’s datasets are proprietary and multimodal and include molecular structures, binding measurements, safety profiles, and manufacturing outcomes. “We’ve built an intentionally diverse portfolio across multiple disease areas and drug types. All of that data empowers us to fine-tune frontier AI models with richer, more representative training sets.” She continues, “Further, we have invested in deep screening technologies to generate additional datasets required in volume to constantly refine and validate our models.”

Building an autonomous discovery engine

To bring all of that data together in one place, AstraZeneca is building what it calls a “lab of the future” facility in Kendall Square, Cambridge, Massachusetts where AI and robotic automation will be able to form a continuous, closed-loop discovery system. “Where a self-driving car uses sensors and models to navigate its environment, this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data,” explains Sapra. That data feeds directly back into the models, accelerating each subsequent cycle.

“Throughout, scientists will remain central to the process, providing the oversight, judgement, and strategic direction that ensure outputs are explainable, tolerable, and directed toward potential patient benefit,” she adds.

Eventually, automated high-throughput systems will be able to make and evaluate thousands of molecular interactions on a weekly basis. “This will generate AI-ready data at a scale that traditional workflows cannot match,” Sapra says. “Robotic sample handling, automated quality checks, and integrated data pipelines also have the potential to help accelerate early drug development timelines significantly.”

The next frontier: Generating medicines from scratch

Ultimately, Sapra says, the end-state vision for AI in biologic drug discovery is what the field calls “de novo” design. For this, the goal is for AI to generate entirely new protein sequences that precisely fit the desired drug properties. This includes designing the structure, predicting safety, how it will behave in the body and how to make it manufacturable.

“The field is making great progress toward a completely AI-generated biologic, designed from scratch all the way to a clinical candidate,” Sapra says. “As we continue to leverage frontier models and fine-tune them with the right datasets, we bring ourselves closer to this reality. I believe it will come. It’s a matter of time.”

Several key elements are needed to reach this point, however. First is richer and more standardized training data across the industry. Second, robust evaluation benchmarks for AI-generated candidates. And third, teams that know how to work at the intersection of machine learning and biology. Of all the prerequisites, however, safety prediction may be the most consequential, and perhaps the least discussed, Sapra says.

“One of the hardest problems in de novo design is predicting whether a computationally generated molecule will be safe in the human body,” Sapra explains. AstraZeneca is tackling this with what amounts to virtual clinical trials. These are advanced cell systems and micro-scale organ models that function as physical testbeds, paired with AI that learns from their outputs.

 “These systems have the potential to generate enhanced biological signals without traditional testing bottlenecks, and they’re a critical missing piece in closing the loop between AI-generated designs and clinical-ready candidates,” Sapra adds.

A shift currently underway is the move toward agentic AI systems that can simultaneously generate molecule candidates and predict how efficacious and safe they are likely to be. These autonomous workflows can connect disease-level insights directly to molecule design, bridging what were previously separate data silos. “The complexity of the biology goes hand-in-hand with the design of the molecule,” summarizes Sapra.

Human talent unlocks AI potential

The transformation underway in biologics is not just about technology. “With more autonomous systems, human oversight remains at the heart of this approach—ensuring explainable and ethical AI for the benefit of patients,” says Sapra.

For scientists, working with AI is a collaborative process. “Scientists will work hand-in-hand with these model systems,” she says. “There will be a world where models will design molecules, then scientists will work with the systems to test those molecules and put all that data together.” Through this process of human checks, balances, and judgement calls, the models will evolve and constantly improve, ultimately with potential to benefit patients.

For engineers, designing and building effective systems ready for human-AI collaboration will mean ensuring high levels of model transparency and explainability. According to Sapra, AstraZeneca’s engineering teams include data scientists, automation specialists, and AI engineers, who are developing systems that act as “thinking partners” rather than black boxes. “Engineers are designing systems that generate, validate, and learn at speed. And the problems are genuinely hard: Multimodal data fusion, closed-loop optimization, uncertainty quantification, and interpretability at the point of clinical decision-making,” she adds.

In taking on such technically demanding challenges, engineers and scientists have the opportunity to contribute to the research and development of potentially life-changing treatments for many diseases, says Sapra. “The biologic medicines we can develop today, and those we’ll design tomorrow, depend on combining world-class AI and engineering talent with deep scientific expertise.”

This article has been initiated and funded by AstraZeneca.  Z4-85058, July 2026.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Voice Control Toolkit Comes to a Pico Near You

A Raspberry Pi Pico 2 W connected to a speaker

Voice-controlled appliances are nothing new. What might be new, however, is [Moonshine AI] running it all locally on a Raspberry Pi Pico 2 W!

The voice interface is roughly divided into three parts: voice activity detection, SpellingCNN speech-to-text and a neural text to speech. The speech to text supports up to 50 tokens, and can be re-trained to support any specific words you want. It runs a simple loop: detect voice activity, listen for (command) tokens, process them in C++, use the TTS to reply, and repeat.

Now, to be fair, it is a bit of a squeeze: 3.6 MiB of the available 4 MiB FLASH and 468 KiB SRAM on a stock Pi Pico 2 board. It leaves you with just about enough space to write a small amount of extra software, but it’ll be a challenge to fit anything substantial. Still, fitting three different types of AI model needed to make this possible in such a space is quite impressive.

Microsoft 2.5: A new series on the people shaping the company’s future

Nearly 20 years ago (!), in 2007, I published my first and only book: Microsoft 2.0. It focused on changes I expected at the company in the “Post-Gates” era. What would remain the same and what likely would be different once co-founder and CEO Bill Gates had left the building?

CEO Satya Nadella has not exited the company (yet). But there’s no question that Microsoft and its mission have morphed considerably in the past year or two. I’m not quite ready to christen this the Microsoft 3.0 era, even though Nadella handed the reins of Microsoft’s dominant commercial business to Judson Althoff nearly a year ago.

That decision resulted in Nadella moving into more of a “founder mode” role, allowing him to focus less on the day-to-day work of running the business. (Microsoft historians may recall that Gates made a somewhat similar move back in 2000 when he became Microsoft’s chief software architect.)

While it might not yet be time for Microsoft 3.0, we arguably could be in the “Microsoft 2.5” era. Windows and Office are still around and still play a big role. Microsoft still builds and sells developer tools and databases. But there’s no question that the cloud and all things AI are at the top of the pecking order now.

I’m embarking on a series here at GeekWire that will focus on what matters to Microsoft and, by extension, to its customers, partners, investors, and employees these days. Who are some of the people shaping and leading the company? What are their opportunities and challenges right now?

Over the next few weeks, I will be profiling various Microsoft execs working on plans for Microsoft’s ongoing evolution. Some are company veterans; some are newcomers. I’ll be talking with top execs from Microsoft’s Security, Copilot, Windows + Devices, Xbox, GitHub, and more.

I’m interested in their strategies for Microsoft’s key products and technologies and how they plan to try to turn Microsoft’s ambitious vision into reality. What are their teams building? What do they see as their biggest challenges and opportunities? And where do they see the technologies in their respective areas heading?

I feel like many of us who’ve been keeping track of the biggest tech companies (myself included) have fallen into the trap of blaming or attributing everything a company does to AI. Layoffs? AI is the culprit. Price increases? It’s all thanks to AI. Changing sales strategies? Chalk it up to AI …

But upon further reflection, I believe Microsoft’s strategy is more nuanced than “AI or bust.” There’s no question that Microsoft’s AI ambitions are shaping its goals and tactics. But Microsoft, as a heavily enterprise-focused entity, can’t simply stop supporting products that aren’t built from the ground up with AI (as much as it might like to do so). Nor can it just leave behind customers who aren’t 100% onboard with its AI moves.

Couple those enterprise hurdles with some not-so-popular consumer decisions, like axing 3,200 people in the gaming unit, and Microsoft’s approach to turning the ship looks a lot trickier.

Our Microsoft 2.5 series kicks off Thursday. Stay tuned.

New Markdown rival: Open-source DGML format aims to turn docs into data that AI (and humans) can trust

L-R: Mantra CEO John Patrick Mullin, Docugami CEO Jean Paoli, and Inveniam CEO Patrick O’Meara. The companies are partnering to make DGML a standard for AI, with Docugami turning documents into data, Inveniam verifying it on a blockchain, and Mantra providing the chain.

Jean Paoli has spent his career making documents readable by machines — first as a co-creator of XML, then helping build the file formats behind Microsoft Office. Now his Kirkland, Wash.-based startup, Docugami, is open-sourcing the technology at the heart of its business, betting it can become a standard way to turn documents into data that people and AI agents can trust. 

The company is releasing its technology, called DGML (short for Document Graph Markup Language), under Apache 2.0, a widely used open-source license, so other developers and companies can adopt it.

The idea is to turn it into a shared standard that no single company owns, much as XML became a common foundation across the tech industry. 

The move reflects a shift in where the value is created in AI. Docugami until now has made its money selling software that turns unstructured documents into usable data. It’s betting now that there’s more value in proving that data is trustworthy instead. 

How it works: Docugami is teaming up with Inveniam, a Detroit company whose software helps big investors keep tabs on the mountains of paperwork behind real estate and other hard-to-value assets. Inveniam will record a kind of digital fingerprint of each piece of DGML data on NVNM Chain, its blockchain built with Mantra, a crypto firm that Inveniam is acquiring.

That means, for example, that a single fact buried in a 200-page lease — such as the rental rate, a renewal option, or a default clause — can be verified on its own, without exposing the whole document. An investor, auditor, or AI agent can trace it to the page it came from. 

To work with documents, AI systems usually convert them into a simpler format first. DGML enters a growing field of contenders in that regard, competing with the popular Markdown format and DocLang, a new open standard for AI-ready documents backed by IBM, Nvidia and Red Hat.

The business model: This is a big move for a company of Docugami’s size, taking the 30-person startup in a new direction. Paoli is handing the industry the technology his team spent years building, and pinning the company’s future on a larger idea.

The plan is to make money not from the format itself but from the value of the trusted data. Once a company converts its leases or loans into DGML and anchors the key numbers on the blockchain, investors, lenders and auditors can pay to draw on that verified data.

Docugami will share in the revenue through its partnership with Inveniam. The company also stands to collect a small fee each time a piece of data is recorded on the chain. 

The company is giving away the DGML format and a working version of the software, but not everything. Paoli said the company is keeping some of its own technology private, including AI models it has fine-tuned to read documents, and could sell those or other tools to enterprises. 

“The business model of everybody is changing. And if you know any company where it’s not true, you need to tell me, because I haven’t met them yet,” Paoli said in an interview. 

Docugami has raised about $13 million to date, including a $10 million seed round in 2020 that drew the first investment in Grammarly’s history.

The partnership: Paoli met Patrick O’Meara, Inveniam’s CEO, a few months ago, through a former Microsoft colleague who had become one of O’Meara’s advisers. They quickly realized they had been working toward the same idea from different directions.

Inveniam, founded in 2017, helps big investors keep track of assets that are hard to value, like office towers, private loans and infrastructure. It monitors the documents behind those assets and flags changes as they happen, and its clients include some of the world’s largest sovereign wealth funds, according to O’Meara.

What it lacked was a consistent way to break those documents into verifiable pieces. That is what Docugami provides.

“We’re not putting the data itself on-chain, just a fingerprint of the document. Change one bit, one byte, one pixel, and the hash won’t match,” O’Meara said.

The blockchain comes from Mantra, a crypto company run by John Patrick Mullin. Inveniam invested $20 million in Mantra last year and has since agreed to acquire it outright. Mantra’s OM token collapsed in April 2025, erasing several billion dollars in value. 

Paoli said the project uses the underlying blockchain, not the token.

“Crypto as an industry has gone through a lot of changes in the last 18 to 24 months, and it’s growing up in a lot of ways. This is a real use case with fundamental value, not just pure speculation,” Mantra’s Mullin said in an interview. 

The result is a division of labor: Docugami turns documents into data, Inveniam verifies it and brings the customers, and Mantra provides the chain where the proof is recorded.

The DGML specification, sample documents and reference code are at dgml.io and on GitHub

Editor’s note: This story was updated after publication to correct the name of a competing document format, DocLang, and to note that Inveniam’s blockchain is called NVNM Chain.

❌