In This Article
- Executive Summary
- What a Self-Driving Lab Actually Is
- The Four Suitability Criteria
- Where Closed-Loop Experimentation Genuinely Works
- Where the Claims Outrun the Evidence
- The Measurement Bottleneck
- The Underrated Payoff: A Complete Experimental Record
- Where This Sits Against the Regulated World
- What Leaders Should Do With This
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
A self-driving lab closes the loop between hypothesis, experiment, measurement, and the next hypothesis. A model chooses what to run next, robotics execute it, an instrument measures the result, and the model updates. No human sits in each cycle. The idea is old. What changed recently is that a handful of groups have published results where the loop genuinely ran unattended and produced something a human program would have taken far longer to find. Those results are real. They are also concentrated in a narrow set of problems, and the marketing around the field has moved much faster than the science.
The useful question is not whether self-driving labs work. It is which problems they work on. Four conditions decide it: the objective has to be measurable directly by the system, the measurement cycle has to be fast relative to the number of experiments needed, the search space has to be navigable by an algorithm, and the experiment has to be something a robot can execute reliably and repeatably. Reaction condition optimization, formulation screening, and materials synthesis score well on all four. Much of biology fails on at least two, usually the measurement ones. That framework explains almost every success and almost every disappointment in the published record.
This article gives the four criteria in detail, walks through the published work that actually supports the claims and the published work that does not, explains why measurement time is the binding constraint rather than model quality, makes the case that the complete machine-readable experimental record may be worth more than the autonomy itself, and sets out what would have to change before a closed-loop system could inform a decision that a regulator reviews.
What a Self-Driving Lab Actually Is
The plain definition is simple. A self-driving lab is a system where a decision algorithm selects the next experiment, automated hardware runs it, an instrument measures the outcome, the result feeds back into the algorithm, and the cycle repeats without a person approving each step. Every part of that loop existed separately for decades. High-throughput screening gave us the robotics. Design of experiments gave us the planning. What is new is the closure: the algorithm gets to choose, and the choice is informed by results the system generated itself minutes or hours earlier.
Two things follow from that definition, and both matter more than they first appear.
First, autonomy is not binary. The literature review by Volk and Abolhasani proposes a scale that runs from piecewise systems, where the platform and the algorithm are completely separate and a human moves data between them, through semi-closed-loop and closed-loop systems, up to self-motivated systems that set their own objectives.6 Almost everything called a self-driving lab in a vendor deck sits in the first two categories. Very little published work sits in the fourth, and nothing in drug development does.
Second, the loop has a scope. A closed-loop system optimizes within a space someone defined. Somebody chose the reaction, the reagent set, the concentration ranges, the objective function, and the analytical method. The system searches that space efficiently. It does not decide the space was the wrong one. Andrew Cooper’s group made this point in practice with their mobile robotic chemist: the robot used the same instruments a human chemist uses, in the same lab, and searched a ten-variable space that the chemists had specified.1
The term has been diluted. “Self-driving lab” now gets applied to a Bayesian optimization loop running on a single flow reactor and to a facility with fifty coordinated instruments. Both may be legitimate work. They are not the same claim, and they do not carry the same implications for an operating budget. When a vendor or an internal team uses the phrase, ask which level of the autonomy scale they mean and how many consecutive experiments have actually run without human intervention.
Why the field grew where it did
Self-driving labs emerged first in chemistry and materials science, not in biology, and that is not an accident of funding. Abolhasani and Kumacheva traced the rise of these platforms through chemical and materials sciences specifically, and the common thread across the successful platforms is a set of shared physical characteristics: liquid handling that is well understood, reactions that complete in minutes to hours, and analytical methods that produce a number rather than an image requiring interpretation.13
Those characteristics are not universal. They are the reason the field looks the way it does, and they are the basis of the framework in the next section.
The Four Suitability Criteria
If you want to know whether a problem in your organization is a candidate for closed-loop experimentation, four questions answer it. They are not weighted equally, and a problem that fails on measurement time cannot be rescued by a better model.
A measurable objective the system can read directly
The system needs a number it can obtain by itself at the end of each experiment. Yield by HPLC area percent qualifies. Transfection efficiency by luminescence qualifies. “Developability” does not, because it is a composite judgment. If a scientist has to interpret the output before the model can use it, the loop is not closed.
A measurement cycle fast enough to matter
Optimization campaigns typically need tens of experiments, sometimes hundreds. If one measurement takes a week, the campaign takes a year and autonomy buys nothing. The binding constraint is the slowest step in the loop, and in almost every real platform that step is analysis, not synthesis.
A search space an algorithm can navigate
Continuous or ordinal variables with known bounds work well. Temperature, residence time, stoichiometry, solvent ratio, lipid tail length. Categorical spaces with no useful similarity structure work badly, because the model has no basis for generalizing from one point to the next.
An experiment a robot can execute reliably
Reliability here means low variance across replicates, not just physical feasibility. A step that a robot can perform but performs inconsistently injects noise the model interprets as signal. Anything requiring tactile judgment, live cell handling under changing conditions, or manual troubleshooting fails this test.
How to apply the criteria
Score a candidate problem honestly against all four. A problem that clears all four is a genuine candidate. A problem that fails one can sometimes be reworked: a slow analytical method can occasionally be replaced with a faster surrogate, and a categorical space can sometimes be reparameterized into descriptors the model can use. A problem that fails two is not a self-driving lab problem, and pursuing it anyway produces an expensive automation project with a machine learning wrapper.
| Problem domain | Measurable objective | Fast cycle | Navigable space | Robot-executable | Verdict |
|---|---|---|---|---|---|
| Reaction condition optimization (flow) | Strong | Strong | Strong | Strong | Proven |
| Inorganic materials synthesis | Moderate | Moderate | Strong | Strong | Proven, with caveats on characterization |
| Nanoparticle and colloid synthesis | Strong | Strong | Strong | Strong | Proven |
| Lipid nanoparticle formulation screening | Strong | Moderate | Moderate | Strong | Demonstrated |
| Enzyme thermostability engineering | Strong | Moderate | Moderate | Moderate | Demonstrated in narrow scope |
| Crystallization and polymorph screening | Moderate | Moderate | Moderate | Moderate | Partial, characterization limits it |
| Cell-based potency assays for biologics | Moderate | Weak | Weak | Weak | Not a closed-loop problem today |
| In vivo efficacy and tolerability | Weak | Weak | Weak | Weak | Not a closed-loop problem |
The pattern in that table is worth stating plainly. The domains that work are the ones where a physical measurement produces a clean number quickly. The domains that do not work are the ones where the biology takes time to respond and the response is variable. That is a property of the science, not of the technology, and no amount of model improvement changes it.
Where Closed-Loop Experimentation Genuinely Works
Here is the published evidence, with the details that make it credible rather than the headline numbers that make it exciting.
Reaction condition optimization is the strongest case
This is the area where the results are least ambiguous, because the objective is unambiguous and the measurement is fast. A group working with an automated flow reactor platform combining Vapourtec modules, a Gilson liquid handler, and inline LC-MS applied multi-task Bayesian optimization to four palladium-catalyzed reactions producing pharmaceutically relevant oxindoles, including intermediates for a serine palmitoyl transferase inhibitor, for linezolid, and for an NK1 receptor antagonist.10
The efficiency numbers are the point. The first campaign used 23 experiments in total, 16 of them training runs plus 7 optimization runs, against a conventional design of experiments approach the authors estimated would have needed more than 750. Later campaigns, drawing on the accumulated data from earlier ones, needed 11, 5, and 10 experiments respectively, consuming 980 mg, 250 mg, and 450 mg of starting material.10 In early development, when material is genuinely scarce, that difference is not a marginal improvement. It is the difference between running the campaign and not running it.
A separate and equally instructive result came from self-optimizing a telescoped three-step sequence, a Heck cyclization followed by an acid-catalyzed deprotection. The Bayesian optimization loop found the optimum in 13 experiments across 14 hours of run time, reaching 81 percent overall yield.11 What makes it instructive is the engineering behind the feedback: rather than buying multiple analytical instruments, the team fitted sampling valves at each reactor outlet and sequenced them so that a single HPLC could quantify both steps. That is a measurement-bottleneck solution, and it is the sort of detail that decides whether a platform works.
Materials and nanoparticle synthesis
Cooper’s mobile robotic chemist ran 688 experiments over eight days inside a ten-variable space, driven by a batched Bayesian search, and found photocatalyst mixtures roughly six times more active than the starting formulations.1 The design choice that made it work was deliberately conservative: instead of building a bespoke automated instrument, the team used a mobile robot that operates the standard laboratory instruments a human would use. That keeps the analytical methods, and therefore the measurement characteristics, unchanged.
Flow-based nanomaterial platforms have produced comparable results. AlphaFlow used reinforcement learning to discover and optimize multi-step colloidal nanoparticle chemistry in a self-driven fluidic system, exploring a space that would have been impractical to cover manually.16 More recently, a self-driving lab for photochemical synthesis of plasmonic nanoparticles targeted specific structural and optical properties directly, which is possible because optical measurement is essentially instantaneous.17
Formulation screening, where biology starts to enter
The most relevant recent result for drug delivery is LUMI-lab, a platform built around a transformer-based molecular model pretrained on tens of millions of structures and coupled to an automated closed-loop workflow that synthesizes ionizable lipids, formulates them into lipid nanoparticles, and screens them. Over ten rounds of active learning it synthesized and evaluated more than 1,700 lipid nanoparticles, and the top candidates outperformed clinically used benchmarks in preclinical testing. The system also surfaced a design feature nobody had directed it toward: brominated lipid tails made up roughly 8 percent of the chemical library but accounted for more than half of the top performers. The work was published in Cell.12
This clears the four criteria because the readout is a transfection measurement in cultured cells, which is quantitative, reasonably fast, and available to the system without human interpretation. It is biology, but it is the part of biology that behaves like chemistry.
Protein engineering in a narrow scope
The SAMPLE platform is the clearest published case of closed-loop autonomy applied to proteins. Intelligent agents learned sequence-to-function relationships for glycoside hydrolase enzymes, designed new variants, and sent them to a fully automated system that built and tested them. Four independent agents, each with different search behavior, converged on thermostable enzymes.5
The reason it works is worth naming explicitly: thermal tolerance is a single, quantitative, physically measured property. It does not require a cell assay, a multi-day incubation, or an interpretive judgment. Change the objective to something like immunogenicity risk or in vivo half-life and the loop cannot close, because the system cannot measure the objective.
The common structure across every success. In each case above, the system measures a physical quantity, the measurement is fast relative to the experiment, and the objective is a single number the algorithm can act on. When a published result looks impressive, check whether it has that structure. If it does, the result probably generalizes to similar problems. If it does not, be careful.
Where the Claims Outrun the Evidence
The field’s most publicized result is also its most instructive cautionary case, and it has nothing to do with the robotics.
The A-Lab and what happened next
In late 2023, a team at Berkeley published an autonomous laboratory for solid-state synthesis of inorganic powders that combined computation, literature data, machine learning, and active learning. Over 17 days of operation it produced 41 novel compounds out of 58 targets.2 The engineering was genuine, the throughput was real, and the paper appeared in Nature.
Within weeks, materials chemists Robert Palgrave at University College London and Leslie Schoop at Princeton posted an analysis arguing that no new materials had in fact been discovered. Their objections were specific. Roughly two thirds of the reported compounds were ordered versions of materials already known to be disordered. The automated Rietveld refinement of the X-ray diffraction data was, in their assessment, at a novice level. The system did not account for substitution and site mixing, so it treated variants of known materials as new discoveries, and several supposedly distinct compounds had effectively identical diffraction patterns.3
The response from the A-Lab team acknowledged that a human would perform a higher-quality refinement and argued that the objective had been to demonstrate autonomous laboratory operation rather than to replace expert analysis.3 That is a fair position for the authors to take. It is also exactly the gap a leader needs to understand.
The failure was in the characterization layer, not the robotics. The robots did what they were told. The synthesis worked. What failed was the automated interpretation of the analytical result, which is the step that converts a physical outcome into the number the loop depends on. Every closed-loop system has this step, and it is almost always the weakest link. When you evaluate a platform, spend your scrutiny there rather than on the robot arm.
The reporting problem across the field
It is difficult to compare self-driving lab results because most papers do not report enough to compare. Volk and Abolhasani surveyed seventeen publications and found that only 23 percent included real-world benchmarking of the optimization algorithm and only 12 percent included simulated benchmarking, leaving 65 percent with no algorithm comparison at all. Seventy-one percent reported no quantitative precision data, meaning no measure of how reproducible a single condition was. None reported the accessible parameter space in a way that would let a reader judge the difficulty of the search.6
Those are not small omissions. Without a comparison against random sampling, a reported optimization success does not establish that the model contributed anything. Without replicate precision, the reader cannot tell whether an improvement exceeds experimental noise. The authors proposed eight metrics to fix this, covering degree of autonomy, operational lifetime, throughput, experimental precision, material usage, accessible parameter space, and optimization efficiency.6 Any organization evaluating a platform can use that same list as a request for information.
What large language model agents did and did not demonstrate
The Coscientist work is frequently cited as evidence that language models can run experiments. What it actually showed is narrower and still notable: a GPT-4-driven system that used web search, documentation search, code execution, and laboratory automation to design and execute experiments across six tasks, including successfully optimizing palladium-catalyzed cross-couplings.4 The system’s contribution was orchestration: reading documentation, writing the code to drive the instrument, and planning the sequence. The optimization itself still rested on established methods, and the chemistry chosen was chemistry that already met the four criteria.
That is a real capability and it lowers the engineering burden of building these systems. It is not evidence that a language model can conduct science in domains where the measurement problem is unsolved.
The Measurement Bottleneck
This is the section that decides most real projects, and it comes down to arithmetic.
A closed-loop campaign needs some number of experiments to converge. Call it N. The wall-clock time of the campaign is N multiplied by the cycle time, and the cycle time is dominated by the slowest step in the loop. In the platforms described above, that step is almost never the synthesis or the liquid handling. It is the analysis.
Working the numbers
Take a realistic optimization needing 40 experiments. With a 20-minute inline HPLC method and parallel reactors, the campaign runs in a couple of days and autonomy is transformative because nobody has to be present overnight. With a 24-hour cell-based assay run in batches of 20, the same campaign takes several days and autonomy still helps, though the benefit is now scheduling rather than speed. With a two-week in vivo readout, the campaign takes more than a year, and the model’s ability to pick the next experiment is irrelevant because the organization will have made the decision by other means long before the loop finishes.
The rule to carry into any evaluation
Estimate the number of experiments the campaign needs, multiply by the cycle time of the slowest measurement, and compare that to the decision deadline. If the answer exceeds the deadline, the problem is not a closed-loop candidate no matter how good the model is. This single calculation eliminates most of the proposals that reach an R&D leadership team under the self-driving lab heading.
Variability is the second half of the measurement problem
Speed is not the only measurement property that matters. Optimization algorithms treat the measured value as information about the underlying system. When the assay has high well-to-well or run-to-run variability, the algorithm spends its budget chasing noise. In chemistry, replicate precision on a yield measurement is often within a couple of percent. In cell-based biology, coefficients of variation of 15 to 30 percent are common and are treated as acceptable. A model receiving those numbers cannot distinguish a genuine 10 percent improvement from assay drift.
There are two responses. One is to build replication into the loop, which multiplies the experiment count and therefore the campaign time. The other is to accept that the loop can only resolve large effects and to set the objective accordingly. Both are legitimate. Neither is what a vendor pitch usually implies.
Why the interoperability problem makes this worse
The measurement step is also where the software breaks. The Nature Communications perspective on self-driving lab accessibility is blunt about instrument interfaces: few application programming interfaces are provided or supported by manufacturers, and many that exist are poorly documented or come with restrictive licensing. The result is that research groups build redundant, non-transferable integrations for the same instruments.7 The same paper notes that complete experimental metadata, including conditions such as temperature and humidity as well as the model parameters used to select the experiment, is often not captured by commercial units at all, which undermines reproducibility across platforms.7
For a pharma or biotech organization, this is the practical entry barrier. Before any question of autonomy arises, someone has to get structured, timestamped, contextualized data out of the analytical instruments in a form a model can consume. That work is unglamorous, it is where the budget goes, and it has value whether or not the loop ever closes.
The Underrated Payoff: A Complete Experimental Record
Here is the argument that gets the least attention and may matter most.
A closed-loop system cannot function without recording every condition it tried and every result it obtained, including the failures. It has no choice: the model needs the negative results to update. So the output of a closed-loop campaign is not just the optimum. It is a complete, structured, machine-readable record of the entire search, with conditions, timestamps, instrument settings, raw measurements, and the model state that produced each choice.
Compare that to how the same campaign is normally documented. A scientist runs conditions, records the promising ones carefully, records the failures briefly or not at all, and writes up the successful route. The failures, which contain most of the information about the boundaries of the chemistry, largely disappear.
The negative data problem is well documented
This is not a theoretical concern. Machine learning models for reaction prediction are trained overwhelmingly on published and patented reactions, which are biased toward successes. Work published in Science Advances examined this directly, using a dataset of 748 negative reactions against as few as 22 positive examples, and showed that reinforcement learning on the negative data improved reaction outcome prediction over fine-tuning alone. The authors are explicit that where negative datasets exist at all, they are often inaccessible or not in machine-readable form.14
An organization running closed-loop campaigns generates exactly the asset the field is short of, as a by-product, in machine-readable form, with full context. Over a few years of operation across a development portfolio, that record is a genuine proprietary advantage, and it does not depend on any of the autonomy claims being true.
The practical reframe. If the autonomy case is uncertain but the data case is strong, buy the data case. Instrument the workflow, capture every run including the failures, and standardize the format. If closed-loop optimization later becomes viable for that workflow, the hardest prerequisite is already in place. If it does not, you still have a structured record of what your scientists tried and what happened, which is more than most organizations have.
Reproducibility follows from the same property
The Cronin group’s work on chemical programming languages makes this concrete. Rather than describing a synthesis in prose, the procedure is encoded in a machine-executable format that captures parameters a written method typically omits, such as stirring rates. Procedures encoded this way were executed and repeated across multiple robotic platforms in different laboratories, which is a far stronger reproducibility claim than a published method section supports.9 The earlier work on digitizing and automatically executing published synthesis procedures established the same principle: a synthesis expressed as code can be versioned, transferred, and re-run without reinterpretation.18
For regulated organizations, that framing should sound familiar. A procedure that is executable, versioned, and produces its own complete record is closer to what a validated process looks like than a prose method ever was.
Where This Sits Against the Regulated World
Almost everything described in this article sits in discovery and early development. That placement is deliberate on the part of the researchers and it is also where the regulatory framework currently draws its line.
The current boundary
The FDA’s draft guidance on the use of artificial intelligence to support regulatory decision-making for drug and biological products covers AI models used in nonclinical, clinical, post-marketing, and manufacturing phases where the model produces information supporting a regulatory decision about safety, effectiveness, or quality. It explicitly does not cover AI use in drug discovery, or operational efficiencies that do not affect patient safety, drug quality, or study reliability.15 CDER has said publicly that it will keep developing a risk-based regulatory framework for AI across the drug development life cycle.19
That exclusion is why closed-loop experimentation has been able to develop as fast as it has. A model that picks the next reaction condition during route scouting is not making a regulatory decision, and nobody expects a validation package for it.
What changes when the output crosses the line
The line gets crossed sooner than people expect. Consider a few cases:
- A closed-loop campaign defines the operating ranges that become the basis of a proposed design space in a filing.
- Autonomous crystallization screening is offered as the evidence that a comprehensive polymorph search was performed.
- A formulation selected by an autonomous platform proceeds to a clinical batch, and the selection rationale becomes part of the development history.
- Data generated by an autonomous system supports a specification justification or a stability claim.
In each of those, the output is now supporting a regulatory decision, and the guidance’s risk-based credibility assessment framework applies. That framework asks for the context of use to be defined precisely, for model risk to be assessed against that context, and for credibility evidence to be gathered proportionate to the risk.15 The relevant point for a self-driving lab is that credibility evidence for a decision model is not the same as credibility evidence for a prediction model. The question is not only whether the model predicted well, but whether the search it conducted was adequate to support the conclusion drawn from it.
The three questions that would have to be answered
Can you show where every data point came from?
Each measurement has to be traceable to a qualified instrument in a known state at a known time, with the raw data retained and the processing applied to it recorded. Autonomous platforms are capable of doing this better than manual workflows, but only if the integration was built with that requirement in mind. Retrofitting provenance onto a research platform after the fact is difficult and often incomplete.
Can you explain what the system chose not to do?
A model that selects experiments also declines experiments. If the conclusion is that a search was comprehensive, the basis for the regions the algorithm did not sample has to be defensible. This is the question that a polymorph screening claim turns on, and it is a different question from model accuracy.
Can you demonstrate the system was in a known state throughout?
Autonomous platforms update models between cycles. If the model changed during a campaign whose output supports a decision, the change history is part of the record. That means version control on models and on the code that drives the instruments, with the same discipline applied to changes as to any other system whose output matters.
None of these are unreachable. All three are engineering and governance problems rather than scientific ones. The honest position today is that very few platforms have been built with them in mind, because the platforms were built to do research and the research did not require them. An organization that wants closed-loop work to eventually touch a regulated decision should specify these requirements at procurement, not after the first useful result.
What Leaders Should Do With This
The practical path is narrower and less exciting than the coverage suggests, and it is also more achievable.
Start from a problem, not from the technology
The most common failure mode is an organization deciding it needs a self-driving lab and then searching for something to point it at. Run the four criteria across the problems you already have. Reaction condition optimization in process chemistry, excipient and formulation screening, and analytical method optimization are the places where the criteria usually clear. If nothing in your portfolio clears all four, that is a legitimate answer and it saves a large amount of money.
Fix the measurement layer first
Instrument integration and structured data capture are prerequisites for closed-loop work and are valuable independently. Getting timestamped, contextualized, machine-readable results out of your analytical instruments improves every downstream use of that data whether or not a model ever selects an experiment. This is also where the interoperability problems described earlier will surface, and it is better to discover them during an integration project than during an autonomy project.
Ask platform vendors the right questions
Questions worth asking before any commitment
- Where does your system sit on the autonomy scale, and what is the longest run of consecutive experiments it has executed without human intervention in a customer environment?
- What is the demonstrated replicate precision for a single condition on this platform, expressed as a standard deviation?
- Has the optimization algorithm been benchmarked against random sampling on this problem class, and can we see the comparison?
- What is the cycle time of the slowest step in the loop, measured end to end, including analysis?
- What data does the system capture for a failed or aborted experiment, and in what format?
- Which instrument interfaces are documented and supported by the manufacturer, and which did you build yourself?
- If we later need this data to support a regulatory submission, what would have to change?
Set the expectation with the science organization
Closed-loop systems do not reduce the need for scientific judgment. They relocate it. Defining the search space, choosing the objective function, deciding what counts as an acceptable measurement, and interpreting what the search found are all still human work, and they are harder work than running the experiments was. Teams that understand this adopt the technology well. Teams that were told the system would do the science are the ones that abandon it after a year.
The A-Lab episode is the clearest illustration. The autonomy worked. The judgment about what the diffraction data meant did not, and no amount of additional autonomy would have fixed it. Expert review of the characterization was the missing element, and it will remain the missing element in any system where the interpretation step is harder than the measurement step.
Conclusion
Closed-loop experimentation is real and it is genuinely valuable in a narrow band of problems that share a specific structure: a directly measurable objective, a fast measurement, a search space an algorithm can navigate, and an experiment a robot performs consistently. Reaction optimization, nanoparticle and materials synthesis, and a growing portion of formulation screening sit inside that band, and the published results there are strong enough to act on. Most of biology sits outside it, not because the models are inadequate but because the measurements are slow and variable, and that is a property of the biology rather than a gap that better software closes. The most valuable thing a leader can do with this technology is apply those four criteria honestly to their own portfolio rather than to the field as a whole.
The second point is the one we would press harder. The autonomy is the headline, but the complete machine-readable record of every experiment attempted, including the failures, is probably the more durable asset. It is the thing the wider field is short of, it compounds across campaigns, and it can be built without betting on any autonomy claim. Organizations that instrument their measurement layer properly get that asset regardless of what happens next, and they also happen to be the ones positioned to answer the provenance and model-state questions if this work ever needs to support a regulatory decision.
Sakara Digital works with pharma and biotech organizations building the data and validation foundations that autonomous and AI-assisted research depends on. If you are evaluating closed-loop experimentation, or trying to work out whether the instrument integration should come before the autonomy project, we are happy to have that conversation.
For Further Reading
For Further Reading
- Building AI as Scientific Infrastructure: Platform Strategies for Drug Discovery
- The Biotech Digital Twin: When Reality Catches Up With the Pitch
- Selecting a LIMS in 2026: What to Look For Beyond Features
- The Data Lineage Tools Comparison for Pharma R&D in 2026
- The AI Model Risk Assessment for Pharma: A Structured Checklist
References & Sources
- Burger, B. et al. “A mobile robotic chemist.” Nature 583, 237-241, July 2020. https://www.nature.com/articles/s41586-020-2442-2
- Szymanski, N. J. et al. “An autonomous laboratory for the accelerated synthesis of inorganic materials.” Nature 624, 86-91, November 2023. https://www.nature.com/articles/s41586-023-06734-w
- Chemistry World. “New analysis raises doubts over autonomous lab’s materials discoveries.” December 2023. https://www.chemistryworld.com/news/new-analysis-raises-doubts-over-autonomous-labs-materials-discoveries/4018791.article
- Boiko, D. A., MacKnight, R., Kline, B., Gomes, G. “Autonomous chemical research with large language models.” Nature 624, 570-578, December 2023. https://www.nature.com/articles/s41586-023-06792-0
- Rapp, J. T., Bremer, B. J., Romero, P. A. “Self-driving laboratories to autonomously navigate the protein fitness landscape.” Nature Chemical Engineering, 2024. https://www.nature.com/articles/s44286-023-00002-4
- Volk, A. A., Abolhasani, M. “Performance metrics to unleash the power of self-driving labs in chemistry and materials science.” Nature Communications 15, 1378, February 2024. https://www.nature.com/articles/s41467-024-45569-5
- “Science acceleration and accessibility with self-driving labs.” Nature Communications, April 2025. https://www.nature.com/articles/s41467-025-59231-1
- Seifrid, M. et al. “Autonomous Chemical Experiments: Challenges and Perspectives on Establishing a Self-Driving Lab.” Accounts of Chemical Research 55, 2454-2466, 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC9454899/
- Rauschen, R. et al. “Universal chemical programming language for robotic synthesis repeatability.” Nature Synthesis, 2024. https://www.nature.com/articles/s44160-023-00473-6
- Taylor, C. J. et al. “Accelerated Chemical Reaction Optimization Using Multi-Task Learning.” ACS Central Science, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC10214532/
- Clayton, A. D. et al. “Bayesian Self-Optimization for Telescoped Continuous Flow Synthesis.” Angewandte Chemie International Edition, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC10108149/
- Lab Manager. “AI-Powered Self-Driving Lab Accelerates Discovery of mRNA Delivery Materials” (reporting the LUMI-lab study published in Cell, University of Toronto). 2026. https://www.labmanager.com/ai-powered-self-driving-lab-accelerates-discovery-of-mrna-delivery-materials-35042
- Abolhasani, M., Kumacheva, E. “The rise of self-driving labs in chemical and materials sciences.” Nature Synthesis 2, 483-492, 2023. https://www.nature.com/articles/s44160-022-00231-0
- “Negative chemical data boosts language models in reaction outcome prediction.” Science Advances, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12164950/
- Goodwin Procter LLP. “FDA Publishes Draft Guidance on Use of Artificial Intelligence in the Development of Drugs and Biological Products.” Client alert on the January 2025 FDA draft guidance, January 2025. https://www.goodwinlaw.com/en/insights/publications/2025/01/alerts-lifesciences-aiml-fda-publishes-its-first-draft-guidance
- Volk, A. A. et al. “AlphaFlow: autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by reinforcement learning.” Nature Communications, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC10015005/
- “Self-driving lab for the photochemical synthesis of plasmonic nanoparticles with targeted structural and optical properties.” Nature Communications, 2025. https://www.nature.com/articles/s41467-025-56788-9
- Mehr, S. H. M., Craven, M., Leonov, A. I., Keenan, G., Cronin, L. “A universal system for digitization and automatic execution of the chemical synthesis literature.” Science 370, 101-108, 2020. https://pubmed.ncbi.nlm.nih.gov/32913000/
- U.S. Food and Drug Administration, Center for Drug Evaluation and Research. “Artificial Intelligence for Drug Development.” Program page, accessed August 2026. https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development








Your perspective matters—join the conversation.