The Record, Checked Against the Primary Sources

Claims about AI drug discovery travel fast and get checked rarely, so the first job is to establish what is actually true as of the day this is written. Three primary sources settle it: the regulators’ own approval lists, the ASCO abstract that is being cited everywhere, and the sponsor’s own disclosures about the most advanced candidate.

What FDA and EMA have approved

FDA’s Center for Drug Evaluation and Research keeps a running list of novel drug approvals by year. As of its most recent entry, dated September 3, 2026, the 2026 list holds 37 approvals.2 FDA does not classify approvals by the method used to discover the molecule, so the list cannot tell you directly whether a drug was AI-designed. What it can do is let you check names. None of the 36 comes from the cohort of AI-native discovery companies discussed in this article, and none is claimed by any of them.

The picture in Europe is the same. EMA’s annual report, published in June 2026, records 104 human medicines recommended for marketing authorization in 2025, of which 38 contained a new active substance never before authorized in the EU.3 Again, no AI-native company’s molecule appears among them.

One caveat that honest analysis has to state: baricitinib. In January 2020, BenevolentAI used its knowledge graph to search already approved drugs for a COVID-19 treatment and proposed baricitinib, an existing Lilly JAK inhibitor. FDA converted its emergency authorization to a full approval for hospitalized COVID-19 patients in May 2022.4 That is a real, approved use that AI helped identify. It is also a repurposed molecule that had already been through discovery, development, and approval for another indication years earlier. The zero-approvals claim is about medicines discovered or designed with AI. Baricitinib is the strongest evidence that AI can shorten a specific, narrow task (screening known drugs for a new use). It is not evidence that AI can design a new one.

What the ASCO abstract actually says

The number that gets repeated is 117. It comes from an abstract presented at the 2026 ASCO Annual Meeting by Waqas Haque and colleagues at the University of Chicago and Gustave Roussy, published in the Journal of Clinical Oncology supplement.1 The methodology matters more than the headline. The authors manually curated AI-enabled therapeutic assets that had entered at least one Phase 1 to 3 interventional trial through July 1, 2025, using industry databases, public disclosures, and trial registries. AI enablement was defined at the asset level and required evidence that AI contributed to discovery or design.

The results, with status as of December 1, 2025: 117 assets across 63 companies. Oncology accounted for 69 of them (59.0 percent). Small molecules made up 96 of 117 (82.1 percent). Sixty assets (51.3 percent) had completed Phase 1. Eight (6.8 percent) had completed Phase 2. About 35.9 percent targeted a novel biological target. At the company level, the median time from founding to first Phase 1 entry was 6.5 years, median total funding was $186.7 million, and the median company had 73 employees when it reached the clinic.1

117AI-enabled therapeutic assets across 63 companies that had entered interventional trials by July 20251
8Of those 117 had completed Phase 2 as of December 1, 2025 (6.8 percent)1
0AI-discovered or AI-designed new molecules among FDA’s 37 novel approvals in 2026 or EMA’s 38 new active substances in 20252,3

Two things about this abstract deserve attention before anyone quotes it. First, “AI-enabled” is a broad category. It includes programs where AI proposed both the target and the molecule, and programs where AI contributed to one step of an otherwise conventional process. The authors were explicit that they required evidence of AI contribution, but the category still mixes very different levels of involvement. Second, the authors describe their own work as a baseline for evaluating clinical impact as the cohort matures. That framing is right. Most of these assets entered the clinic recently, and a Phase 2 completion rate of 6.8 percent partly reflects how young the cohort is rather than how it will perform.

Why “eight years”

The first molecule created with AI to enter human trials was DSP-1181, a serotonin 5-HT1A agonist for obsessive-compulsive disorder, discovered by Sumitomo Dainippon Pharma using Exscientia’s technology. Its Phase 1 study in Japan was announced on January 30, 2020, with the discovery phase completed in under 12 months.5 Count back from there through the 6.5-year median founding-to-clinic time in the ASCO cohort and you reach 2018 to 2019, when the current generation of AI-native discovery companies was being founded and funded.1 Eight years is a reasonable round number for how long the field has had serious money and serious intent.

The most advanced candidate

Rentosertib, Insilico Medicine’s TNIK inhibitor for idiopathic pulmonary fibrosis, is the candidate most often cited as the first drug where both the target and the molecule came from generative AI. Its Phase 2a results were published in Nature Medicine in June 2025: 71 patients across 21 sites in China, four arms, 12 weeks of treatment. The 60 mg once-daily arm showed a mean forced vital capacity change of +98.4 mL against a decline of 20.3 mL on placebo, with a manageable safety profile.6 On July 7, 2026, Insilico announced the start of a Phase III study: 320 patients across 47 centers in China, randomized, double-blind, placebo-controlled, with a primary endpoint of the annual rate of FVC decline over 52 weeks.7

The press release gives no readout date, but the arithmetic is straightforward. A trial with a 52-week primary endpoint that begins enrolling in the second half of 2026 cannot produce topline data before late 2027, and a realistic expectation is 2028. Even then, a positive result in a China-only study is the start of a regulatory conversation in the US and EU, not the end of one. This is the single most advanced AI-designed molecule in the world, and its earliest plausible approval in a major market is 2028 or later.

Why This Is Not an Argument Against AI

It would be easy to read the record above as a verdict on the technology. It is not, and the same primary sources explain why.

The clearest evidence comes from a 2024 analysis in Drug Discovery Today by Madura Jayatunga and colleagues, who examined the clinical pipelines of AI-native biotech companies. They found that AI-discovered molecules had an 80 to 90 percent success rate in Phase I, substantially above historic industry averages. In Phase II, the success rate was about 40 percent on a limited sample, comparable to historic averages.9 Read those two numbers together and the shape of the field becomes clear. AI is very good at producing molecules with drug-like properties: molecules that are safe enough, absorbed well enough, and tolerable enough to pass a first-in-human study. AI has not yet shown any advantage at the thing Phase II tests, which is whether the biological hypothesis was right and the drug actually helps patients.

That distinction lines up with a broader critical review published in Nature Reviews Drug Discovery in August 2026 by Andreas Bender and fifteen co-authors from academia, pharma, and AI-native companies. Their assessment is that a wide variety of AI methods have been developed and benchmarked, yet “evidence of their clinically relevant impact is, so far, disappointingly limited.” They attribute this to insufficient focus on clinical translation during model development, difficulty applying algorithms to the conditional, context-dependent data that life science produces, and underspecified problem definitions, along with a pattern they describe as technology push rather than science pull.8

None of that says AI cannot design medicines. It says the field is roughly where you would expect a genuinely new capability to be after eight years: strong on the parts of the problem that reduce to chemistry and physics, unproven on the parts that depend on human biology in a heterogeneous patient population, and still waiting for the cohort of clinical assets to mature. The Phase I numbers are a real achievement. They are also, for a budget owner, the wrong numbers to plan around, because Phase I success does not produce revenue. Approval does.

So the argument in this article is not “AI does not work.” It is “AI works at different speeds in different places, and a budget should reflect that.” The rest of the article is about where those places are.

Where AI Returns Value in Life Sciences Today

While the discovery pipeline matures, there is a set of workflows in pharma and biotech where AI is already returning value, with evidence that can be inspected rather than promised. They share a profile: the work is document-heavy or record-heavy, the correct answer is usually knowable and checkable by a qualified person, the volume is high, and the improvement can be measured inside a single budget year. Six areas stand out.

RETURNS NOW

Data quality remediation

Profiling, deduplication, reference data reconciliation, and cleanup of master data and historical records across LIMS, QMS, ERP, and clinical systems. The output is measurable: fewer defects, fewer reconciliations, fewer investigation triggers.

RETURNS NOW

Documentation and audit trail work

Audit trail review, record completeness checks, and drafting of controlled documents from structured inputs. Documentation and records deficiencies appear in almost every FDA warning letter, so the improvement is visible to inspectors too.

RETURNS NOW

Validation authoring

Drafting risk assessments, test scripts, traceability matrices, and validation summaries under GAMP 5 and FDA’s risk-based Computer Software Assurance approach. The reviewer stays; the blank page goes.

RETURNS NOW

Regulatory operations

First drafts of submission sections, response-to-query packages, and label change tracking, grounded in the source documents. Human authorship and sign-off remain, and the cycle time drops.

RETURNS NOW

Investigation triage

Classifying deviations by severity and likely impact, retrieving precedent investigations and CAPAs, and turning investigator notes into structured draft reports. Root cause analysis is where quality teams say the effort goes.

RETURNS NOW

Pharmacovigilance case processing

Intake, translation, duplicate detection, MedDRA coding, and causality support for adverse event cases. Case volumes have grown far faster than headcount, and the per-case work is repetitive and well defined.

Data quality remediation

Poor data quality has a price that predates AI. Gartner’s research puts it at an average of at least $12.9 million a year per organization.23 In pharma the price shows up as regulatory exposure as well. A full-enumeration study of 1,766 FDA warning letters issued between 2016 and 2023, published in Therapeutic Innovation and Regulatory Science in 2026, reclassified data integrity violations against an ALCOA+ rubric and found that violations related to endurance, availability, and completeness of records rose year over year after 2020, with the average number of data integrity violations per cited company increasing in 2023.22 A separate analysis of 85 warning letters issued to drug manufacturers between January 1 and December 9, 2025 found that 15 percent explicitly cited data integrity, that 21 CFR 211.192 (production record review and investigation of discrepancies) appeared in 22 letters, and that FDA recommended engaging a GMP consultant in 87 percent of them.21

AI’s contribution here is unglamorous and effective: profiling large record sets for anomalies, matching duplicate entities across systems, flagging incomplete or inconsistent records for human correction, and reconciling reference data such as product codes and terminology versions. These are pattern-matching tasks on data the organization already owns, and the result is a defect count that goes down. The payback is inside the year, and it compounds, because every downstream AI use case depends on the same data being right.

Documentation and audit trail work

Documentation is the most consistently cited category in FDA enforcement, and audit trail review is one of the most consistently underperformed activities in GxP operations, because it is tedious, high-volume, and rarely finds anything on any given day. That profile is exactly where AI helps. Models can review audit trail entries for patterns that a human reviewer would miss at scale (out-of-sequence changes, edits after approval, unusual user activity), draft controlled documents from structured inputs, and check records for completeness against a template before a human signs. The human review does not go away. Its focus moves from reading everything to reading what the system flagged.

Validation authoring

Validation is a writing problem as much as a testing problem. FDA’s final Computer Software Assurance guidance, announced in the Federal Register on September 24, 2025 and reissued on February 3, 2026 as Computer Software Assurance for Production and Quality Management System Software, describes a risk-based approach that establishes confidence in automation through assurance activities proportionate to risk rather than through documentation volume.24 The guidance is written for device production and quality software, and pharma teams read it alongside GAMP 5, but the direction is the same across both: less paper, more targeted evidence. AI fits that direction well. Drafting a risk assessment, generating test scripts from requirements, building a traceability matrix, and writing a summary report are all tasks where a model produces a usable first draft and a qualified person edits and approves. The risk-based judgment stays human. The blank page does not.

Regulatory operations

Regulatory writing is source-grounded by nature: every claim in a submission traces to a study report, a dataset, or a prior submission. That makes it a strong fit for retrieval-based generation, where a model drafts from an approved corpus and every sentence can be checked back. First drafts of submission sections, response-to-query packages, periodic reports, and label change tracking are the common starting points. The January 2026 FDA and EMA guiding principles of good AI practice in drug development, discussed below, apply directly here: clear context of use, data governance and documentation, and clear essential information about what the model did.16 When those are in place, regulatory operations is one of the lowest-risk, highest-volume places to put AI to work.

Investigation triage

An ISPE Pharmaceutical Engineering piece from November 2025 described the integration of a generative AI assistant, built on Microsoft Copilot inside an organization’s own Microsoft 365 tenant, into deviation management at a contract testing laboratory. In an ISPE webinar on the topic, more than 60 percent of participants identified root cause analysis as the most resource-intensive and challenging step. The assistant guides investigators through structured prompts covering problem statement, scope, and impact, retrieves historical deviations and CAPAs, and assembles a draft investigation narrative for human review. The reported results were faster deviation closure, better first-pass documentation quality, and improved CAPA effectiveness, with the explicit condition that all AI-generated content is reviewed, approved, and archived by qualified personnel and that the implementation runs under change control aligned with GAMP 5.18

That last condition is the point. Investigation triage with AI is not autonomous decision-making. It is severity classification, precedent retrieval, and draft assembly for a human investigator who was previously doing all three from scratch. The same FDA analysis cited above found 211.192 in 22 of 85 warning letters in 2025.21 Better investigations are a compliance outcome as well as an efficiency one.

Pharmacovigilance case processing

Pharmacovigilance is the most mature of the six. The volume problem is documented: Deloitte noted that adverse event reports submitted to FDA’s Adverse Event Reporting System grew from about 500,000 in 2009 to more than 2.2 million in 2021, and that PV spending goes predominantly to case processing. In a case study they cite, natural language processing doubled auto-coding of adverse events from 30 percent to more than 60 percent and cut the manual time required by half.20 A 2025 narrative review in the International Journal of Clinical Pharmacy reported a deep learning model applied to free-text hospital adverse reaction reports that improved allergic reaction detection by 24 percent and reduced the need for manual review by 64 percent, and described the authors’ own expert-defined Bayesian network tool at a regional pharmacovigilance center, which reached a positive predictive value of 87.3 percent at the “probable” causality level and reduced causality assessment from days to hours.19

Across all six areas, the pattern is the same: high-volume, well-defined work, a human who remains accountable, and a measurable result inside the budget year. That profile is what “returns value today” means in this article. It is a different profile from the discovery bet.

Where AI Is Still a Bet: Targets and Molecules

Target identification and molecule design are where the money and the attention have gone, and the money has been large. Isomorphic Labs, the Alphabet spinout built around AlphaFold, raised $600 million in its first external round in March 2025, led by Thrive Capital, to develop its drug design engine and move programs into the clinic.12 Xaira Therapeutics launched in April 2024 with more than $1 billion in committed capital from ARCH Venture Partners and Foresite Labs.13 The ASCO cohort’s median total funding of $186.7 million per company, across 57 companies with disclosed figures, gives a sense of the scale below those headline raises.1

What has that capital produced so far? Molecules that pass Phase I at unusually high rates, and then face the same Phase II odds as everyone else.9 The attrition is now visible in company disclosures. Recursion, which combined with Exscientia in November 2024 to form the largest AI-native drug discovery company, announced in its first quarter 2025 results that it was discontinuing REC-2282 in neurofibromatosis type 2 and REC-994 in cerebral cavernous malformation after the totality of data did not support continuing, and would consider out-licensing REC-3964 in C. difficile infection. The company reported $509 million in cash and a runway into mid-2027.10 Before the merger, Exscientia had already pruned its pipeline to four active candidates and ended internal development of EXS-21546, its A2A receptor antagonist, after concluding the program could not reach the prolonged high level of target coverage it would need.11

None of these discontinuations is unusual for drug development. That is the point. AI-designed molecules are failing in Phase II for the reasons drugs have always failed in Phase II: the target was wrong, the exposure was insufficient, the effect was too small. The Bender review’s diagnosis, insufficient focus on clinical translation during model development, describes this directly.8 Designing a molecule that binds is a solved-enough problem. Choosing the right thing to bind, in the right patients, is not, and the field’s own leaders say so.

The economics of the bet

Deloitte’s 16th annual analysis of pharmaceutical R&D returns, covering 20 large companies, puts the average cost to develop an asset from discovery to launch at $2,671 million in 2025, up from $2,229 million in 2024, with that figure including the cost of failure. Projected internal rate of return on late-stage assets reached 7.0 percent in 2025, the third consecutive year of improvement, but the median company was at 4.0 percent, and the headline number was driven by a small number of GLP-1 and GLP-1/GIP assets in obesity. Deloitte’s own recommendation to its cohort includes demanding measurable, end-to-end productivity returns from AI investments.14

Put those numbers next to the discovery pitch and the shape of the bet is clear. The promise is that AI reduces the time and money spent before a candidate is nominated. Even if that promise is fully kept, discovery is the front of a process that then runs through Phase I, II, and III at industry-standard attrition, and the money is overwhelmingly spent after discovery. A faster start is worth something. It is worth much less than a higher Phase II success rate, and a higher Phase II success rate is the thing no one has yet demonstrated. Until it is demonstrated, discovery AI is a call option on the pipeline, and it should be priced like one.

The number that matters for a discovery vendor is not how fast they reached a candidate. It is how many of their candidates have completed Phase 2, and what the Phase 2 success rate is across the whole platform, including the programs that were discontinued. Jayatunga and colleagues found about 40 percent, on a small sample, matching the industry baseline.9 Any vendor claiming better should be able to show the asset-level data.

The Mistake of Funding Both as One Line Item

In most 2027 budget cycles, AI arrives as a single line, or as a single steering committee, or as a single “AI strategy” that a board asked for. The operational use cases and the discovery use cases get discussed in the same meeting, defended with the same language, and measured against the same expectations. That is the mistake this article is about, and it fails in one of two directions.

In the first direction, the operational work is judged by the standard of the discovery pitch. Data quality remediation, audit trail review, and pharmacovigilance automation are not exciting, and next to a slide about designing new medicines they look like maintenance. They get funded thinly, staffed with people who also have day jobs, and asked to justify themselves in terms of “transformation” rather than in terms of defects closed, hours returned, and findings avoided. The result is that work with a documented payback inside the year is starved because it does not tell a good enough story.

In the second direction, the discovery bet is judged by the standard of the operational work. A leadership team that has seen AI return real value in case processing assumes the same technology will return real value in target selection on the same timeline, and treats a multi-year, venture-scale, high-attrition investment as an operating expense with a twelve-month payback. When the first program is discontinued in Phase II, as programs are, the reaction is not “that is what a portfolio bet looks like” but “AI did not work,” and the whole budget, including the operational work, gets cut.

Both failures come from the same cause: treating two investments with different risk profiles, different time horizons, and different evidence bases as if they were the same thing because they share a technology. Deloitte’s 2024 analysis of AI value in life sciences estimated that R&D accounts for 30 to 40 percent of the potential value, commercial for 25 to 35 percent, with manufacturing, supply chain, and enabling functions providing the rest, and put the timeline for an enterprise to capture peak value at about five years.15 Even that estimate, which is generous to R&D, puts most of the value outside discovery and most of the timeline beyond a single budget year. A budget that reflects that has to separate the buckets.

A Portfolio Approach to the 2027 Budget

The alternative is to run AI spending the way a portfolio manager runs capital: by risk bucket, with different sizing rules, different evidence requirements, and different review cadences for each. Three buckets are enough for most pharma and biotech organizations.

1

Bucket one: operational returns

The six areas above. Funded from operating budgets, owned by the functions that do the work (quality, regulatory, safety, IT quality), and measured on outputs that already exist in those functions: deviation cycle time, case processing time per case, validation deliverable turnaround, audit trail review coverage, data defect counts. Expected payback inside 12 months. This bucket should be the largest in 2027 for any organization that has not already worked through it.

2

Bucket two: foundations

The work that every other AI use depends on and that no single use case will fund on its own: data quality and governance, model inventory and risk classification, validation approach for AI-enabled systems, and the policies and training that let people use the tools without creating findings. Funded as a capability investment with a two- to three-year horizon and measured on readiness, not return. This is where the FDA and EMA guiding principles translate into internal standards.

3

Bucket three: discovery bets

Target identification and molecule design, whether through an internal platform, a partnership with an AI-native company, or an equity position. Funded and governed as R&D portfolio investments with the same stage gates, probability-of-success adjustments, and kill criteria as any other early program. Expected readout 2027 to 2028 at the earliest. Sized so that a total loss does not threaten buckets one and two.

How to size each bucket

There is no universal percentage split, and any consultant who offers one is guessing. What there is instead is a set of sizing rules that follow from the evidence.

Size bucket one to the backlog, not to a percentage. The operational use cases have a natural ceiling: the volume of deviations, cases, validation deliverables, and records the organization actually has. Count them. A mid-sized biotech with a few hundred deviations a year and a modest safety caseload has a smaller bucket one than a multi-site manufacturer, and that is correct. The sizing question is “what would it take to work through the documented backlog and the recurring volume with AI assistance and qualified review,” and the answer is a number of tools, integrations, and people, not a share of the AI budget.

Size bucket two to what bucket one and bucket three both need. Foundations are shared infrastructure. If the discovery partnership will require the organization to supply clean, well-governed internal data, that requirement belongs in bucket two and should be funded before the partnership starts, not discovered after. If the operational use cases will need a validation approach for AI-enabled systems that the quality unit will accept, that belongs here too. The test for every item in bucket two is whether two or more use cases depend on it.

Size bucket three to the loss the organization can absorb. This is the venture rule, and it is the right one. Deloitte’s cost-per-asset figure of $2,671 million and the ASCO cohort’s median funding of $186.7 million per company describe what serious discovery efforts spend.14,1 An organization that cannot lose its bucket-three money without cutting bucket one has sized bucket three wrong. For most mid-sized companies that means bucket three is a partnership or a co-development deal with milestone-based payments rather than a platform build, and the milestone structure is what protects the downside.

Review cadence

Bucket one gets reviewed quarterly on the operational metrics it was funded against. Bucket two gets reviewed twice a year on readiness milestones. Bucket three gets reviewed at the same stage gates as any other early-stage program, which means it is reviewed when data arrives, not when the calendar says so. Mixing those cadences is how discovery bets end up being asked for quarterly ROI and operational tools end up being asked for a five-year vision.

A practical test for any AI line item in the 2027 budget: can the owner say, in one sentence, which bucket it belongs to, what metric it will be judged on, and when the first data point arrives? If the answer includes “transformation” or “the future of the industry” and no metric, it is a bucket-three bet that has been mislabeled as bucket one. If the answer includes a cycle time or a defect count and a date in 2027, it is bucket one and should be funded on that basis.

What Evidence to Demand From Discovery Vendors

If bucket three is going to be funded, the diligence should be as rigorous as for any other R&D investment, and the record above gives a clear list of what to ask for. The FDA and EMA guiding principles of good AI practice in drug development, published jointly on January 14, 2026, provide a useful frame. The ten principles are: human-centric by design, a risk-based approach, adherence to standards, clear context of use, multidisciplinary expertise, data governance and documentation, model design and development practices, risk-based performance assessment, life cycle management, and clear essential information. They are written to cover AI use across the whole lifecycle of a medicine, from early research through manufacturing and safety monitoring.16,17 A discovery vendor that cannot describe its platform in those terms is not ready for a regulated customer.

Vendor claimEvidence to requireWhy it matters
“Our platform discovered this candidate”A written account of exactly which steps AI performed (target hypothesis, hit generation, lead optimization) and which were conventional, with named human decision points“AI-enabled” covers a wide range. The ASCO analysis required asset-level evidence of AI contribution; you should too1
“Faster time to candidate”Time from program start to candidate nomination for every program on the platform, not the best one, with discontinued programs includedSpeed to candidate is real but is the least valuable part of the pipeline economically14
“High clinical success rate”Asset-level Phase I, II, and III outcomes for all platform assets, including partnered and discontinued programs, with datesPhase I success of 80 to 90 percent is documented across the field; Phase II is about 40 percent. A vendor claiming better must show it9
“Novel target”The biological evidence for the target independent of the model output, and the patient selection strategyRoughly 36 percent of the ASCO cohort targeted a novel target; novelty raises Phase II risk, it does not lower it1
“Validated model”Data governance documentation, training data provenance, performance assessment against the stated context of use, and a lifecycle management planThese map directly to the FDA/EMA principles and will be asked about if the work supports a submission16
“Partnership with a top pharma”What the partner has actually paid, what milestones have been reached, and whether any partnered program has reached Phase IIPartnership announcements are not outcomes. Look for milestone payments and clinical progression

Two further questions belong in every diligence conversation. First, what happens to the organization’s own data? Many discovery partnerships require the customer to supply proprietary assay, screening, or clinical data, and the terms for how that data trains the vendor’s model and who owns the resulting improvements are often the most valuable part of the contract. Second, what is the vendor’s financial runway relative to the readout timeline? Recursion disclosed a runway into mid-2027 against programs that will not read out until later.10 A partner that runs out of money before the data arrives leaves the customer holding a program with no platform behind it.

None of this is hostile to discovery vendors. The best of them will welcome these questions because they have the answers. The questions exist to separate the platforms with asset-level evidence from the platforms with a good demo, and that separation is the whole job of bucket-three diligence.

Milestones That Would Change the Picture

A portfolio approach only works if the allocations move when the evidence moves. The record described in this article is a snapshot as of September 2026. These are the specific events that would justify shifting money from bucket one toward bucket three, and the events that would justify the reverse.

Milestones that would increase the discovery allocation

  • A positive rentosertib Phase III readout. The 320-patient, 52-week study started in July 2026.7 A clear win on the primary endpoint, followed by regulatory engagement in the US or EU, would be the first demonstration that a molecule with an AI-proposed target and an AI-generated structure can carry a late-stage program. That is a 2028 event at the earliest.
  • A Phase II success rate that separates from the baseline. Jayatunga’s 40 percent was on a small sample.9 As the ASCO cohort matures, the number of assets completing Phase 2 will rise from 8, and a follow-up analysis with a larger denominator is the evidence to watch. If AI-enabled assets start passing Phase II at a rate meaningfully above the industry norm, the discovery bet’s risk profile changes and so should its budget.
  • A first approval, on any AI-designed molecule, in a major market. Whichever molecule it is, an FDA or EMA approval of a new active substance that its sponsor documents as AI-discovered ends the zero-approvals record. It will not, by itself, prove the platform economics, but it will change the conversation with boards.
  • Regulators classifying AI involvement at approval. FDA does not currently record how a molecule was discovered.2 If FDA or EMA begins to capture that information, the field will finally have an authoritative count rather than a set of competing trackers.

Milestones that would decrease it

  • Further consolidation or runway failures among AI-native companies. The Recursion and Exscientia combination, the pipeline cuts that followed it, and disclosed runways that end before readouts are signals that the capital cycle is turning.10,11 A partner’s financial health is part of the bet.
  • A negative rentosertib readout. A failed Phase III on the most advanced candidate would not disprove the approach, but it would push the first-approval timeline out to 2029 or beyond and would justify holding bucket three flat.
  • Phase II attrition in the ASCO cohort that matches or exceeds the baseline as the sample grows. If the follow-up data show 40 percent or worse on a larger denominator, the honest conclusion is that AI accelerates the front of the pipeline without improving its odds, and the budget should treat it that way.

The value of writing these down now, before the 2027 budget is set, is that the allocation decisions in 2027 and 2028 become responses to evidence rather than reactions to headlines. A board that agreed in advance what a positive rentosertib readout would mean for the budget will make a better decision when it arrives than a board that reads about it in the news.

Conclusion

The record is what it is: eight years, billions of dollars, 117 assets in the clinic, 8 past Phase 2, and no approved medicine designed by AI in a major market as of September 2026. The record is also younger than it looks, and the same sources that document it document the reasons for patience. Phase I success rates are genuinely high. The cohort is maturing. The most advanced candidate is in Phase III. The people who know the field best are writing candid reviews about what needs to change. None of that is a reason to stop funding discovery AI. All of it is a reason to fund it as what it is: a bet with a 2027 or 2028 readout, sized to the loss the organization can absorb, and governed at stage gates like any other early program. Meanwhile, the work that returns value this year (data quality remediation, documentation and audit trail work, validation authoring, regulatory operations, investigation triage, and pharmacovigilance case processing) deserves to be funded on its own evidence, measured on its own metrics, and protected from the fortunes of the discovery pipeline. The leader who separates those buckets in the 2027 budget will be right whichever way the discovery bet turns out. The leader who does not will be wrong in one direction or the other.

Sakara Digital works with pharma and biotech organizations building this kind of evidence-based AI portfolio: identifying the operational use cases with a documented payback, putting the data quality and governance foundations in place that every use case depends on, and structuring the diligence for discovery partnerships so the bet is sized and governed properly. If you are working through a 2027 AI budget and want an independent perspective on where the money should go first, we are happy to have that conversation.

For Further Reading