Where the Guidance Stands on September 3, 2026

Start with what can be checked. The FDA guidance page for Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products lists the document as a Draft Guidance for Industry and Other Interested Parties, issued January 2025, under docket number FDA-2024-D-4689. The status banner reads, in full: “Draft Level 1 Guidance. Not for implementation. Contains non-binding recommendations.”1 We checked that page on the day this article was written. Nothing on it has changed since publication.

The Federal Register notice announcing the draft was published on January 7, 2025, with a 90-day comment period that closed on April 7, 2025. That notice also records what informed the draft: more than 800 comments FDA received on its 2023 discussion papers on AI in drug development and manufacturing, and the agency’s own experience reviewing submissions with AI and machine learning components.3 CDER’s page on artificial intelligence for drug development puts that experience at more than 500 submissions with AI components between 2016 and 2023.5 The draft itself cites a 2023 landscape analysis by CDER staff in Clinical Pharmacology & Therapeutics, covering AI and machine learning in submissions from 2016 to 2021, as the evidence that AI use in submissions has increased.2, 11

The public docket on regulations.gov still accepts comments, since FDA takes comments on any guidance at any time under 21 CFR 10.115(g)(5). As of September 3, 2026 the docket shows more than 120 public comments posted against the draft guidance document.4 Five of those were posted in the last 90 days. People are still writing to FDA about a draft that is now in its second year.

20 monthsTime the guidance has been in draft, from the January 7, 2025 Federal Register notice to September 3, 20263
123Public comments posted on docket FDA-2024-D-4689 against the draft guidance document as of September 3, 20264
500+Drug submissions with AI components received by CDER from 2016 to 2023, per FDA’s own count5

What did and did not happen in 2026

Two things worth separating. First, the expectation of a final version by mid-2026 was an industry expectation, not an FDA commitment. We could find no FDA statement, in the Federal Register or on any agency page, that promised a finalization date. Second, CDER’s guidance agenda for calendar year 2026, updated in July 2026, lists two new AI drafts planned for publication: one on computer software assurance for AI-based systems in drug manufacturing and clinical investigations, which is the single entry under the agenda’s Artificial Intelligence category, and one on AI and machine learning quality considerations in pharmaceutical manufacturing, which is listed under Pharmaceutical Quality and CMC.9 The agenda covers new and revised draft guidances only. It does not list finalizations, so its silence on the January 2025 draft says nothing either way about when a final version will appear.

What did happen is more useful than a date. In January 2026, FDA and EMA jointly published Guiding Principles of Good AI Practice in Drug Development, ten short principles that describe how both agencies think about AI used to generate evidence across the nonclinical, clinical, post-marketing, and manufacturing phases.6 Law firm commentary at the time noted that adoption of the principles is voluntary and that they are not formal guidance, but that sponsors should assess current practices against them because they signal what later guidance on both sides of the Atlantic will look like.15 On January 29, 2026, the ICH Assembly adopted ICH M15, General Principles for Model-Informed Drug Development, at Step 4.7 FDA then issued M15 as final Level 1 guidance in June 2026 under docket FDA-2024-D-5580.8 M15 explicitly names artificial intelligence and machine learning among the modeling and simulation methods it covers. That matters because M15 is adopted and final rather than draft, and shares its assessment vocabulary with the FDA draft. It is not binding: FDA’s own M15 guidance carries the standard nonbinding recommendations header. We come back to that in the section on what will survive.

A note on scope. This article is about the draft guidance for sponsors and the records a sponsor keeps. It does not cover FDA’s internal use of AI in its own review work, which is a separate subject.

What the Draft Actually Asks For

The draft is 23 pages, and much of its length is worked examples and a table of engagement options. The substance can be stated compactly. It applies when an AI model is used to produce information or data that supports a regulatory decision about the safety, effectiveness, or quality of a drug. It does not apply to AI used in drug discovery, or to AI used for operational efficiency (the draft’s own examples are internal workflows, resource allocation, and drafting or writing a regulatory submission) where the use does not affect patient safety, drug quality, or the reliability of study results.2

Three definitions carry the whole document. Credibility is trust in the performance of an AI model for a particular context of use, established by collecting credibility evidence. The context of use, abbreviated COU, is the specific role and scope of the model in answering a question of interest. And model risk is the possibility that the model’s output leads to an incorrect decision that results in an adverse outcome. The draft is careful on this last point: model risk is about the decision, not about some risk intrinsic to the model.2

The two-factor model risk assessment

Model risk combines two independently rated factors. Model influence is the contribution of the AI model’s evidence relative to other evidence used to answer the question. If the model is the sole determinant, influence is high. If independent testing or clinical data also bears on the decision, influence is lower. Decision consequence is the significance of an adverse outcome if the decision is wrong, considering both severity and probability. The draft plots these on a matrix, and the risk rating rises as either factor rises.2

The two worked examples show how this plays out. In the clinical example, a model is the only thing deciding which trial participants go home after dosing rather than staying for 24-hour inpatient monitoring for a life-threatening reaction. Influence is high, consequence is high, model risk is high. In the manufacturing example, an AI-based visual system checks fill volume in every vial, but release testing independently measures fill volume on a representative sample from each batch. Consequence is still high because volume is a critical quality attribute, but influence is low because of the release test, so model risk is medium.2 Notice what the manufacturing example teaches: a sponsor can lower model risk by design, by keeping an independent check in the process, without touching the model at all.

FactorWhat it measuresHow to lower it
Model influenceHow much weight the model’s output carries relative to other evidence answering the same questionAdd independent evidence: release testing, clinical data, a second method, human adjudication that is actually performed and recorded
Decision consequenceSeverity and probability of harm if the overall decision is wrong, assessed on the question itself and irrespective of how the model is usedGenerally cannot be lowered by the sponsor; it is a property of the decision. Controls that catch the harm before it reaches a patient can affect probability
Model riskThe combination of the two, rated on a matrix from low to highDetermines how stringent the credibility activities, acceptance criteria, documentation, and FDA oversight should be

The seven steps

The framework itself is a seven-step process. Steps 1 through 3 define the problem. Steps 4 through 7 plan, run, document, and judge the credibility work.2

1

Define the question of interest

The specific question, decision, or concern the AI model addresses. The draft’s clinical example: “Which participants can be considered low risk and do not need inpatient monitoring after dosing?”

2

Define the context of use

What will be modeled, how the outputs will be used, and whether other evidence (animal studies, clinical data, release tests) will be used alongside the model to answer the question.

3

Assess the AI model risk

Rate model influence and decision consequence independently, then combine them. The draft says this takes subject matter expertise and judgment shared among the sponsor, interested parties, and FDA.

4

Develop a credibility assessment plan

Describe the model, the development data, the training, and the evaluation process, at a depth commensurate with model risk. This is the step FDA most wants to discuss early.

5

Execute the plan

Run the activities. The draft notes that discussing the plan with FDA before execution helps set expectations and surface problems early.

6

Document the results and discuss deviations

Produce a credibility assessment report covering the results of steps 1 through 4 and any deviations from the plan. Agree with FDA whether, when, and where the report goes.

7

Determine adequacy for the context of use

Decide whether credibility is established for the model risk. If not, the draft lists five options: lower model influence by adding evidence, increase rigor, add controls, change the modeling approach, or reject or revise the context of use.

What the plan has to contain

Step 4 is where the draft gets specific, and it is worth reading the list closely because it is the closest thing to a submission checklist the document offers. For each model, the plan should describe inputs and outputs, architecture, features, the feature selection process and any loss functions, model parameters, and a rationale for the modeling approach. For the development data, it should describe how the data were split into training and tuning sets, how the data were collected, processed, annotated, stored, and controlled, how labels were established, why the data are relevant and reliable for the COU, and whether development was centralized or federated. For training, it should describe the learning methodology, performance metrics with confidence intervals, techniques used to prevent over- or under-fitting, hyperparameters, any pre-trained model and its provenance, ensemble methods, calibration, and the quality assurance and version control of the software and packages used.2

For evaluation, the test data must be independent of the development data and never shown to the algorithm during training. The plan should explain how independence was achieved (different trial, different health system, different batches), justify any overlap, describe the reference method used to create the test data and how well that reference performs, address the applicability of the test data to the deployed environment (the draft names data drift here), give performance metrics with confidence intervals, specify how uncertainty and confidence in predictions were estimated, describe the limitations of the approach including potential biases, and describe the quality assurance of code verification.2

One sentence in that evaluation list deserves its own paragraph. If the COU involves a human in the loop, the draft says the evaluation methods should consider the performance of the human-AI team rather than the model in isolation.2 Many sponsors describe a human reviewer as a risk control and then test only the model. The draft closes that gap explicitly, and as we will see, the January 2026 joint principles close it again.

Life cycle maintenance and early engagement

The draft’s Section IV.B addresses models whose use extends over time, with manufacturing as the worked example. Performance metrics should be monitored on an ongoing basis, the level of oversight should be commensurate with model risk, and changes to the model or to the manufacturing process that could affect the model should flow through the change management system of the pharmaceutical quality system. Depending on the impact, steps of the credibility plan may need to be re-executed, including retraining and retesting, and changes that affect performance should be reported in accordance with postapproval change regulations. Detailed maintenance plans belong in the site’s pharmaceutical quality system, with a summary in the marketing application for product- or process-specific models. The draft also points sponsors to ICH Q12 tools: model-related elements can be proposed as established conditions, with a plan for managing changes to them.2

Section IV.C lists the ways to engage. Formal meetings such as INTERACT and pre-IND are the route for a specific development program. Beyond those, the draft names the Center for Clinical Trial Innovation, the Complex Innovative Trial Design program, the Drug Development Tools and ISTAND qualification programs, the Digital Health Technologies program, the Emerging Drug Safety Technology Program for pharmacovigilance, CDER’s Emerging Technology Program and CBER’s Advanced Technologies Team for manufacturing, the Model-Informed Drug Development paired meeting program, and the Real-World Evidence programs.2 The draft states that FDA “strongly encourages” early engagement, and it says the plan a sponsor brings to that engagement should, at minimum, contain steps 1 through 3 and the proposed credibility activities.2 Regulatory Focus, reporting on the draft the day the notice published, picked out that same encouragement and the ongoing monitoring expectation as the two points sponsors most needed to hear.16

Why the Framework Will Survive the Final Version

The practical question for a sponsor is not “what does the draft say” but “which parts of the draft can I build against without regret.” The answer comes from tracing where the framework’s pieces came from and where they have since been adopted.

The framework is older than the draft

Footnote 13 of the draft states that the high-level concepts of the credibility assessment framework, specifically the question of interest, the context of use, and the assessment of model risk, were informed by ASME V&V 40, an FDA-recognized consensus standard for assessing the credibility of computational modeling for medical devices.2 Footnote 22 sends readers to FDA’s November 2023 guidance on assessing the credibility of computational modeling and simulation in medical device submissions for more on decision consequence.2, 10 In other words, the two-factor model risk idea was already FDA’s position in a final guidance before the drug draft was written. The drug draft borrowed it. A final version is not going to un-borrow it.

ICH M15 adopted the same skeleton in January 2026

ICH M15 is the stronger anchor because it is now adopted text, not a draft. M15 defines six key assessment elements: question of interest, context of use, model influence, consequence of wrong decision, model risk, and model impact. Its definitions match the FDA draft nearly word for word. Model influence is “the intended weight of the model outcomes in decision-making considering the contribution of additional data or evidence.” Consequence of wrong decision is “the potential negative effect (e.g., on patient safety and/or lack of efficacy) resulting from an incorrect decision.” Model risk “is derived by combining model influence and consequence of wrong decision.” And M15 makes the same point the FDA draft makes: model risk “is not to be perceived as a risk intrinsic to MIDD or M&S.”7 M15’s own footnote credits the same source, ASME V&V 40.

M15 goes further in ways that are directly usable. Each element is rated low, medium, or high, and “justification is always expected.” It adds a sixth element, model impact, meaning how far the proposed modeling strategy departs from current regulatory standards or expectations. It provides an assessment table in Appendix 1 that is meant to travel with the program from planning through submission. And it describes model evaluation as three activities: verification (the code and calculations are correct), validation (the model agrees with data, prior information, and knowledge), and applicability assessment (the data and model are adequate for the intended use). Among method-specific issues to consider, M15 names “overfitting for an artificial intelligence/machine learning model.”7

ANCHORED

Question of interest and context of use

Defined identically in the FDA draft and in adopted ICH M15. Also GAIP principle 4, “Clear context of use.” Build these now.

ANCHORED

Two-factor model risk

Model influence plus decision consequence, from ASME V&V 40 through FDA’s 2023 device guidance into the drug draft and M15. Not going to change.

ANCHORED

Plan, execute, report, judge

The FDA draft’s credibility plan and report map onto M15’s Model Analysis Plan and Model Analysis Report. Pre-specification and documented deviations are common to both.

ANCHORED

Independent test data, metrics with confidence intervals, drift monitoring

Present in the FDA draft, GAIP principles 8 and 9, the EMA reflection paper, and draft Annex 22. Every regulator has written this down.

The joint principles say the same things at a higher altitude

The ten FDA and EMA principles of January 2026 are short, and reading them against the draft shows that the draft is an instance of them, not a departure. Principle 2, “Risk-based approach,” calls for “proportionate validation, risk mitigation, and oversight based on the context of use and determined model risk.” Principle 4 is “Clear context of use,” defined as the role and scope for why the technology is being used, the same two words the draft uses. Principle 8, “Risk-based performance assessment,” asks that assessments “evaluate the complete system including human-AI interactions, using fit-for-use data and metrics appropriate for the intended context of use.” Principle 9, “Life cycle management,” calls for scheduled monitoring and periodic re-evaluation, and names data drift as the example.6 A final guidance that contradicted these would contradict FDA’s own signature on a joint document eight months old. That is not how agencies behave.

The practical conclusion. Question of interest, context of use, the two-factor risk rating with written justification, a pre-specified plan, independent test data, performance metrics with confidence intervals, a written report with deviations, and a monitoring plan for drift: every one of these is in adopted ICH text, in the joint principles, or both. Work done on these items in 2026 will be reusable when the final guidance appears, whatever its exact wording.

What Is Most Likely to Change

If the skeleton is fixed, what is exposed? Reading the public comments, the joint principles, M15, and the peer-reviewed critique together gives a fairly consistent list.

Generative and foundation models

The draft was written in 2024 and its vocabulary shows it. Its example architecture is a convolutional neural network; its performance metrics are sensitivity, specificity, predictive values, and F1; its examples are a risk classifier and a vision system. Nothing in it is wrong for a large language model, but nothing in it tells a sponsor how to define a test set, a reference standard, or a confidence interval for a model that produces free text. Commentary from Troutman Pepper Locke summarizing the docket notes that commenters asked for “more examples, clarity on generative and foundation models, and risks with third-party AI.”12 The peer-reviewed critique published in the Journal of Chemistry in January 2026 recommends, among other things, establishing tiered explainability requirements and strengthening bias mitigation strategies.14 Expect the final to say something about generative models, even if only to route them to the same seven steps with a note about how the COU should be scoped.

Third-party and vendor models

The draft asks sponsors to describe pre-trained models and their provenance, but it says nothing specific about a model a sponsor licenses from a vendor and cannot fully inspect. The joint principles hint at where this goes: principle 6 asks that “data source provenance, processing steps, and analytical decisions are documented in a detailed, traceable, and verifiable manner.”6 The EMA reflection paper is blunter and already in force as EMA’s position: where a third-party model is used in a high-impact or high-patient-risk setting, EMA expects the manufacturer of the system to have provided details through a methodology qualification process covering the specific context of use.17 A final FDA guidance that stays silent here would leave a gap that EMA has already filled, so some language on vendor responsibility is likely.

Human-AI team performance and plain-language disclosure

The draft mentions the human-AI team once, inside the evaluation list. The joint principles elevate it to principle 8 and add principle 10, “Clear, essential information,” which asks for plain language about the technology’s context of use, performance, limitations, underlying data, updates, and explainability, addressed to users and patients.6 The National Health Council’s April 2025 comment made the same point from the patient side, asking that patients be clearly informed whenever AI tools significantly affect regulatory or clinical decisions related to their care, and that clinicians and regulators retain the ability to verify or override AI-generated recommendations.13 A final guidance is likely to say more about both.

Vocabulary and the scope boundary

Two smaller points. First, the draft’s footnote 29 says FDA does not use the word “validation” for tuning data, the way the machine learning community does.2 The EMA reflection paper flags the same collision and keeps the AI usage.17 M15 uses “validation” in the modeling sense.7 Expect the final FDA text to add a glossary or to align with M15. Second, the scope carve-out for operational efficiency, including drafting a regulatory submission, is the line most sponsors ask about. The Journal of Chemistry review argues for expanding scope to include discovery and operational phases.14 We think the carve-out survives, because FDA has other instruments (CGMP, GCP, data integrity expectations) for operational uses, but the final may sharpen the language on what “does not impact the reliability of results” actually means.

Where to hold off. Do not build a company-specific taxonomy of generative AI risk tiers and present it to FDA as if it were the guidance. Do not treat a vendor’s model card as a substitute for a credibility assessment. And do not write internal SOPs that quote draft line numbers; the numbers will move. Write SOPs to the seven steps and the six M15 elements by name, and reference the guidance by title.

The No-Regrets Work List

Here is the work that will not be thrown away. It is ordered so that each item feeds the next, and each item is tied to text that is already adopted or jointly signed.

  1. Inventory every AI model whose output could reach a regulatory decision. Use the draft’s scope test: does the output support a decision on safety, effectiveness, or quality? Sort models into in-scope and operational. Keep the operational list too, because the boundary will be questioned at inspection and you want to show you drew it deliberately.
  2. Write one question of interest and one context of use per model, in M15’s words. Role, scope, data used to build the model, and what other evidence answers the same question. M15 recommends a separate assessment table for each question of interest.7 If you cannot write the question in one sentence, the model is not ready for a credibility plan.
  3. Rate model influence, decision consequence, and model risk as low, medium, or high, and write the justification. M15 says justification is always expected. Do this with a multidisciplinary team, which is what both the draft and GAIP principle 5 call for. Record who rated it and when.
  4. Design down model influence where you can. The draft’s manufacturing example is the template: keep an independent check in the process and the model risk drops from high to medium. This is the single highest-return design decision available and it is invisible to the model itself.
  5. Build the credibility assessment plan as a controlled document, using M15’s Model Analysis Plan structure. Introduction, objectives, data, methods, planned evaluation activities, and pre-defined technical criteria. Pre-define means documented before the data are accessed or the analysis is run.7
  6. Establish test data independence with technical controls, not just a policy. Every regulator asks for this. Draft Annex 22 spells out the mechanics that will satisfy the strictest reading: access control and audit trail on the test set, a record of which data were used for testing and how many times, and staff who touched test data kept out of training.19 Building to that standard now satisfies FDA’s independence requirement with room to spare.
  7. Report performance with confidence intervals and a quantified statement of uncertainty. The draft says this twice, once for training and once for evaluation.2 Include the reference method’s own performance, because a model cannot be more accurate than the labels it was tested against.
  8. Test the human-AI team, not the model alone, wherever a human reviewer is part of the control strategy. Required by the draft’s evaluation list and by GAIP principle 8. Define the reviewer’s task, measure the reviewer’s performance, and keep the records the way draft Annex 22 describes for human review.19
  9. Write the credibility assessment report to M15’s Model Analysis Report structure, with deviations from the plan described and justified. Executive summary, introduction, objectives, data and methods, results, discussion, conclusions, and appendices including the code. Cross-reference the assessment table.7
  10. Put deployed models under change control and configuration control, and define drift metrics before go-live. The draft’s life cycle section, GAIP principle 9, the EMA reflection paper, and draft Annex 22 all require it. Define the monitoring frequency, the retest triggers, and the threshold at which the model flags an output as undecided rather than guessing.
  11. Request an early engagement on one program. The draft says the minimum package is steps 1 through 3 plus proposed credibility activities. Pick the program with the highest model risk, or the one where the COU is genuinely ambiguous, and use the meeting to learn how a review division reads the framework. Record the feedback; M15 encourages including a summary of previous regulatory feedback in later submissions.7
  12. Assign a named owner for the guidance itself. Someone checks the FDA guidance page and the docket monthly and reports to the governance body. When the final appears, that person runs a gap check of the SOP set against the final text within 30 days.
What this list deliberately leaves out. It does not include writing a generative AI credibility policy ahead of the final text, buying a governance platform to “comply with” a draft, or requalifying operational tools that are outside scope. Those are the activities most likely to be redone.

How to Document AI Use in a Submission Today

The question we hear most from regulatory affairs teams is not “what do we test” but “where does this go.” The draft answers it in three places, and M15 fills in the rest.

Where the plan goes

The draft says the credibility assessment plan can be described in a formal meeting package or through another engagement option, and that whether, when, and where it is submitted depends on the engagement route, the model, and the COU. In early discussions the plan can be high-level, with a detailed version drafted after FDA feedback.2 For a clinical program, that means the plan goes into the pre-IND, INTERACT, or MIDD paired-meeting package. For manufacturing, it goes to the Emerging Technology Program or CBER’s Advanced Technologies Team before the application is filed.

Where the report goes

The credibility assessment report may be a self-contained document included in a regulatory submission or a meeting package, or it may be held and made available to FDA on request, for example during an inspection. The draft says submission of the report “should be discussed with FDA.”2 Footnote 25 addresses the case where no meeting option fits, naming postmarketing pharmacovigilance as the example: certain documentation is not generally submitted but is maintained under the sponsor’s SOPs and made available on request, and in those cases a sponsor may complete all seven steps without seeking early engagement.2

Where manufacturing model records go

For product- or process-specific models in manufacturing, detailed life cycle maintenance plans belong in the site’s pharmaceutical quality system, with a summary in the marketing application. Model-related elements can be proposed as established conditions under ICH Q12, with a change management plan, so that the sponsor obtains FDA’s view in advance on which changes need prior submission.2

What M15 adds about placement

M15 is specific in a way the FDA draft is not. It says the assessment table should be included in the most appropriate sections of the regulatory documentation, naming regulatory interaction background materials and Common Technical Document sections, in line with the question of interest. Individual Model Analysis Plans and Reports should be cross-referenced from the table. And all documents and files supporting submitted evidence, including the data used, relevant code, and definition files, should be submitted or available for review.7 For any AI model that is functioning as a modeling and simulation method within a development program, M15’s placement rules apply now, because M15 is final.

RecordWhere it goes todaySource
Question of interest, context of use, model risk ratings and justifications (the assessment table)Meeting background package at planning; CTD sections aligned to the question of interest at submissionICH M15 Section 4.3 and Appendix 17
Credibility assessment plan (Model Analysis Plan)Formal meeting package (pre-IND, INTERACT, MIDD paired meeting) or the relevant program contact; appended to the report laterFDA draft Step 4; M15 Section 4.12, 7
Credibility assessment report (Model Analysis Report)In the submission or meeting package, or held for inspection, as agreed with FDAFDA draft Step 6; M15 Section 4.22, 7
Life cycle maintenance plan for a manufacturing modelPharmaceutical quality system at the site, with a summary in the application; established conditions under Q12 if proposedFDA draft Section IV.B2
Pharmacovigilance AI proceduresMaintained under SOPs and available on request or at inspectionFDA draft footnote 252
Code, datasets, definition filesSubmitted or available for regulatory reviewM15 Section 4.37

One more practical point. A sponsor filing in both regions should note that the EMA reflection paper expects, for a high-impact or high-patient-risk use in a clinical trial that has not been qualified, that the full model architecture, development logs, validation and testing logs, training data, and a description of the data processing pipeline may be requested at marketing authorization, at clinical trial application, or at GCP inspection.17 That is a longer list than the FDA draft’s. Build the record to the longer list once and both filings are covered.

How the Draft Relates to EMA’s Reflection Paper and Draft Annex 22

Three documents, three different jobs. Confusing them is the most common error we see in global AI governance frameworks, so it is worth being precise.

The EMA reflection paper: adopted, lifecycle-wide, and more prescriptive on clinical trials

EMA’s Reflection paper on the use of Artificial Intelligence (AI) in the medicinal product lifecycle was adopted by CHMP on September 9, 2024 and by CVMP on September 11, 2024, under reference EMA/CHMP/CVMP/83833/2023.17 It is not a guideline, but it is final, and it states EMA’s expectations. It covers the same territory as the FDA draft, from nonclinical through post-authorization, and it adds drug discovery, precision medicine, and product information, which the FDA draft excludes or does not mention.

Where the FDA draft uses one risk axis (model risk from influence and consequence), EMA uses two labels: “high patient risk” for systems affecting patient safety, and “high regulatory impact” for cases where the effect on regulatory decision-making is substantial, with the primary endpoint of a late-stage trial as the glossary’s example.17 The two schemes are compatible. High regulatory impact roughly corresponds to high model influence on a high-consequence decision.

The EMA paper is far more specific about pivotal clinical trials, and this is where a sponsor building to the FDA draft alone would fall short. In late-stage trials, EMA expects that before a model is deployed in a high-impact setting such as the primary endpoint, performance is tested with prospectively generated data from a later calendar time, in a setting or population representative of the intended use. Incremental learning approaches are not accepted. Any modification of the model during the trial requires a regulatory interaction to amend the statistical analysis plan. Before database lock and unblinding, the data pre-processing pipeline and all models must be frozen and documented in the statistical analysis plan, and any non-prespecified change afterward makes the results post hoc.17 EMA also prefers transparent models, accepts black-box models only where the sponsor shows transparent ones perform inadequately, expects explainability methods such as SHAP or LIME where black-box models are used, and recommends performance metrics insensitive to class imbalance, naming the Matthews correlation coefficient.17

EMA’s engagement routes parallel FDA’s: the Innovation Task Force for early interaction on experimental technology, and the Scientific Advice Working Party for scientific advice and for qualification of novel methodologies.17 EMA has already used the qualification route for an AI tool: in March 2025 CHMP issued a qualification opinion for AIM-NASH, an AI system that assists pathologists in reading liver biopsy scans in clinical trials, which EMA describes as the first time it will consider data generated with the assistance of an AI-based tool to be scientifically valid.21 There is no equivalent qualified AI tool on the FDA side under the drug draft, though the draft’s ISTAND route exists for exactly that purpose.

Draft Annex 22: GMP only, static models only, and still open

Draft Annex 22 is a different kind of document. It is a proposed new annex to EudraLex Volume 4, the EU GMP guide, released for public consultation by the European Commission alongside revised Chapter 4 and revised Annex 11 on July 7, 2025, with comments due October 7, 2025.18 It has no final text and no implementation date. Its scope is narrow and explicit: computerized systems used in manufacturing where AI models are used in critical applications with direct impact on patient safety, product quality, or data integrity. It applies to models that learned from data rather than being explicitly programmed, to static models that do not adapt during use, and to models with deterministic output. Dynamic models that learn continuously “should not be used in critical GMP applications.” Probabilistic models are likewise out of scope and should not be used in critical applications. And the draft states plainly that it “does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications.”19

Within that narrow scope, Annex 22 is more prescriptive than either the FDA draft or the EMA reflection paper. Its ten sections cover intended use with a full characterization of the input sample space and its subgroups, acceptance criteria that must be at least as high as the performance of the process the model replaces, test data that is representative and large enough for statistical confidence, test data independence with access control, audit trail, and staff separation, a pre-approved test plan, feature attribution during testing, confidence score logging with a threshold below which the model should flag an outcome as undecided, and operation under change control, configuration control, performance monitoring, and input drift monitoring.19

The generative AI exclusion may not hold. EMA’s Annex 22 drafting group, working with its Quality Innovation Group, convened a two-day multistakeholder expert workshop on June 30 and July 1, 2026 in Amsterdam to gather expert opinion and evidence for the Annex 22 text. The workshop page says the 2025 consultation “suggested support for potentially enabling the use of technologies such as generative AI (GenAI) or large language models (LLMs) in medicines manufacturing,” and that EMA is now seeking expert input on control and mitigation measures such as guardrails as part of a risk-based approach. The workshop is to produce a report. No timeline for the final Annex 22 is given.20

DimensionFDA draft guidance (Jan 2025)EMA reflection paper (Sept 2024)Draft EU GMP Annex 22 (July 2025)
Status on Sept 3, 2026Draft, not for implementationAdopted by CHMP and CVMP; a reflection paper, not a guidelineConsultation draft; no final text, no implementation date
ScopeAI producing information or data supporting safety, effectiveness, or quality decisions; excludes discovery and operational usesEntire lifecycle including discovery, precision medicine, product information, manufacturing, post-authorizationManufacturing of medicinal products and active substances only; critical GMP applications only
Risk conceptModel risk = model influence x decision consequenceHigh patient risk and high regulatory impactQuality risk management based on risk to patient safety, product quality, and data integrity
Model typesAll AI; ML noted as most commonAll models developed through ML; incremental learning not accepted in pivotal trialsStatic, deterministic ML only; dynamic, probabilistic, generative, and LLM excluded from critical use
Test dataIndependent of development data, fit for use, applicability to COU describedEarly train-test split; prospective testing with future calendar-time data for high-impact useRepresentative, stratified, sufficient in size, independent with access control and staff separation
ExplainabilityMethodological transparency; limitations and biases describedTransparent models preferred; SHAP or LIME for black-box modelsFeature attribution captured during testing; feature review part of test approval
Human in the loopEvaluate the human-AI team, not the model aloneSystem risk management plan especially where no human in the loopOperator responsibility documented; operator performance monitored; records of human review kept
EngagementFormal meetings plus nine named programsInnovation Task Force; scientific advice and qualification via SAWPNot applicable (GMP inspection framework)
Change managementPQS change management; Q12 established conditions; postapproval change rulesRe-evaluation for non-trivial stack changes in high-impact useChange control and configuration control before deployment; retest decision documented

What this means for one set of records

A sponsor with U.S. and EU filings does not need three governance frameworks. The FDA draft supplies the risk vocabulary and the seven-step process. The EMA reflection paper supplies the stricter clinical-trial rules on freezing, prospective testing, and explainability. Draft Annex 22 supplies the most detailed mechanics for test data independence and operational monitoring in manufacturing. The joint principles of January 2026 are the statement, signed by both agencies, that these are meant to be read together.6 Build the record to the most demanding requirement in each category and the other two are satisfied by construction. The one place this does not work is generative AI in critical GMP use, where the EU position is currently “do not,” the FDA draft is silent, and the June 2026 workshop suggests the EU position is being reconsidered. That is a hold, not a build.

Conclusion

A draft guidance that is twenty months old and still not final creates an odd kind of paralysis. Teams read it, agree that it is sensible, and then decide to wait for the final before changing anything, on the theory that building to a draft is wasted effort. The evidence says otherwise. The framework’s core came from a consensus standard FDA already used in a final device guidance, it was adopted in ICH M15 in January 2026 and issued by FDA as final in June 2026, and it was restated in the ten joint principles FDA and EMA signed in the same month. The final version of the drug guidance will add examples, will address generative and vendor models, and will tidy the vocabulary. It will not replace the skeleton. Meanwhile EMA’s reflection paper is final and stricter than the FDA draft on clinical trials, and draft Annex 22, though not in force, already describes the test data and operational controls an inspector will eventually expect. A sponsor who builds the question of interest, the context of use, the justified risk rating, the pre-specified plan, the independent test set, and the monitored deployment now is not guessing at the future. That sponsor is doing what three regulators have already written down.

Sakara Digital works with pharma and biotech organizations building this kind of AI credibility record: the model inventory, the context-of-use and risk rating discipline, the plan and report structure that satisfies FDA, EMA, and ICH M15 at once, and the SOPs that will survive the final text. If you are deciding what to build before the guidance is final and want an independent perspective on where to start, we are happy to have that conversation.

For Further Reading