Where Annex 22 Stands at the End of September 2026

Before deciding what to do, it helps to be exact about what exists. Annex 22 is a proposed new annex to the EU Guidelines for Good Manufacturing Practice, published as a draft on 7 July 2025.1 The European Commission’s Directorate-General for Health and Food Safety ran the consultation from 7 July to 7 October 2025, alongside a revised Chapter 4 (Documentation) and a revised Annex 11 (Computerised Systems).2 PIC/S, the Pharmaceutical Inspection Co-operation Scheme, ran the same consultation jointly, because the three texts were drafted by the EMA inspectors’ group together with PIC/S for both the EU and PIC/S GMP guides.3

The consultation closed almost a year ago. Since then, several things have happened that matter for planning.

The Work Plan Target Is a Handoff, Not a Go-Live

EMA’s current three-year work plan for the GMP/GDP Inspectors Working Group, covering January 2026 to December 2028, lists Annex 22 with a target date of Q4 2026. The comment next to that date reads: “To provide the European Commission with a final text for the amended annex in order to assure use of artificial intelligence in the context of GMP.”4 The same plan sets the same Q4 2026 target for Chapter 4 and Annex 11, and says the three are being handled in parallel.

This is where the common phrase “final text by year end” comes from, and it is easy to over-read. A final text delivered to the Commission is a drafting milestone. After that, the Commission must adopt and publish the annex in EudraLex Volume 4, and the published version will carry its own date for coming into operation. The work plan itself anticipates this gap: it lists, as a separate activity, training for EU GMP inspectors to support implementation of the revised Chapter 4, Annex 11, and Annex 22 during 2026 to 2028.4

Nothing Has Appeared in EudraLex Volume 4

The Commission’s EudraLex Volume 4 page, checked on 29 September 2026, does not list an Annex 22. It still shows Chapter 4 and Annex 11 in their January 2011 versions.5 In other words, the binding EU GMP text for computerized systems today is the 2011 Annex 11, and there is no AI annex in force.

Target Dates in This Program Have Already Moved Once

Work plan targets are plans, not commitments. The previous work plan, covering 2025 to 2027, set Q1 2026 as the target for giving the Commission final texts of the revised Chapter 4 and Annex 11.6 The current plan moved both to Q4 2026.4 That is not a criticism of EMA. Three linked documents that drew wide comment take time to settle. It is a reason not to build a compliance program around a single quarter.

EMA Is Reconsidering the Scope

The largest source of uncertainty is scope. On 30 June and 1 July 2026, EMA held a two-day workshop to gather expert evidence for Annex 22. EMA’s event page explains why: the 2025 consultation suggested support for potentially enabling generative AI and large language models in medicines manufacturing, while the draft had said dynamic, adaptive, and probabilistic models should not be used in critical GMP applications. EMA states that it “is still considering the implications of the stakeholder consultation results,” and that the Annex 22 drafting group is seeking input on “possible control and mitigation measures such as guardrails.”7 The first day was an open session with speakers from EFPIA, PDA, and ISPE, among others; the second was a closed session for the drafting group.8 EMA said it expects the workshop to produce a report. At the time of writing, the event page lists only the Day 1 agenda.

7 Oct 2025Consultation on draft Annex 22 closed. No final text has appeared in EudraLex Volume 4 since.25
Q4 2026EMA inspectors’ target to give the European Commission a final text. Not a date for coming into force.4
12 monthsGap between publication and coming into operation for the last major annex rewrite, Annex 1 (2022 to 2023).9

How Long After Publication? No One Has Said

No source we could find states an implementation period for Annex 22, and any figure you see quoted is an estimate. The closest precedent is the 2022 revision of Annex 1 (Manufacture of Sterile Medicinal Products). It was published in EudraLex Volume 4 on 25 August 2022, with a deadline for coming into operation one year later, on 25 August 2023, and two years later for one specific point (8.123).9 Annex 22 may follow a similar pattern, or it may not. Treat the precedent as a reasonable planning assumption and nothing more.

Correcting the Premise

The accurate statement is: EMA’s inspectors aim to give the European Commission a final Annex 22 text in Q4 2026. Publication, the date it comes into operation, and PIC/S adoption into its own GMP Guide are later steps with no announced dates. The final scope may differ from the July 2025 draft, most likely on generative AI and probabilistic models.

What the Draft Says, Read Closely

The draft is short: six pages, ten numbered sections, and a glossary.1 Because it is short, small words carry a lot of weight. Much of the confusion we see in industry summaries comes from reading it quickly. Here is what each part says.

Scope: Critical Manufacturing Uses, Static and Deterministic Models

Section 1 applies the annex to computerized systems used in the manufacture of medicinal products and active substances, where AI models are used in critical applications with direct impact on patient safety, product quality, or data integrity. It gives predicting or classifying data as examples. It describes itself as additional guidance to Annex 11 for systems in which AI models are embedded, and it covers machine learning models that gained their function through training on data rather than explicit programming.1

Then come three limits. The annex applies to static models, meaning models that do not adapt their performance during use by taking in new data. It applies to models with deterministic output, meaning identical inputs give identical outputs. And it does not apply to generative AI or large language models. For each excluded category, the draft adds that such models “should not be used in critical GMP applications.”1

There is one more sentence that matters a great deal and is often skipped. If generative AI or LLMs are used in non-critical GMP applications, the draft says personnel with adequate qualification and training should always be responsible for making sure the outputs are suitable for the intended use (a human-in-the-loop, or HITL), and that the principles of the annex may be considered where applicable.1 So the draft does not ban generative AI from GMP work. It keeps it out of critical applications and puts a qualified person in charge of it elsewhere.

Principles, Intended Use, and Acceptance Criteria

Section 2 asks for close cooperation among process subject matter experts (SMEs), QA, data scientists, IT, and consultants; says documentation must be available and reviewed by the regulated user even when a supplier trained and tested the model; and ties all activities to quality risk management.1

Section 3 requires a detailed intended use, including a full description of the input data the model will see, with common and rare variations (the “input sample space”), known limitations, and possible erroneous or biased inputs. A process SME is responsible for it, and it must be documented and approved before acceptance testing starts. Where relevant, the sample space is divided into subgroups, such as decision output, site or equipment, material characteristics, or defect type.1

Section 4 asks for test metrics suited to the use (it lists a confusion matrix, sensitivity, specificity, accuracy, precision, and F1 score as examples for a classifier) and acceptance criteria set and approved by a process SME before testing. It also sets a floor: acceptance criteria should be “at least as high as the performance of the process it replaces.”1 That floor means you need to know how well the current process performs, which is a measurement task many sites have never done.

Test Data and Test Data Independence

Section 5 covers test data: it should represent the full sample space, be stratified across subgroups, be large enough for adequate statistical confidence, and have labels verified to a very high degree of correctness. Pre-processing must be pre-specified, exclusions justified, and generating test data or labels (for example with generative AI) is not recommended and must be fully justified if used.1

Section 6 is the one most often misquoted. Its independence requirement is about test data. It asks for technical or procedural controls to ensure that “data which will be used to test a model, is not used during development, training or validation of the model.”1 It offers two ways to achieve this: capture test data only after training and validation are complete, or split it from the full data pool before training starts. If you split, “employees involved in the development and training of the model have never had access to the test data,” the test data is kept behind access control with an audit trail, and “There should be no copies of test data outside this repository.” You must record which data was used for testing, when, and how many times. Where test data comes from physical objects, those objects must not have been used to train or validate the model unless features are independent. Staff who have seen test data should not work on training or validation of the same model; where that is impossible, they may do so only in pairs with a colleague who has not had access (a four-eyes arrangement).1

Common misreading. The draft does not require training data to be “independent,” and it does not set a data independence rule for the whole data pipeline. Section 6 protects the hold-out test set from the people and processes that built the model. Training data is addressed through the intended use and sample space requirements, not through Section 6. Getting this right matters because the controls are different: test data independence is mostly about access control, audit trail, and staffing, not about sourcing.

Test Execution, Explainability, and Confidence

Section 7 requires an approved test plan before testing (with the intended use summary, metrics, acceptance criteria, test data reference, test script, and calculation method), documented and justified deviations, and retention of all test records alongside the actual test data and access records.1

Sections 8 and 9 are worded more narrowly than many summaries suggest. Section 8 says that during testing of models used in critical GMP applications, systems should capture and record the features that contributed to a classification or decision, with techniques such as SHAP values, LIME, or heat maps used “where applicable,” and that a review of those features should be part of approving test results. Section 9 says that when testing a model that predicts or classifies, the system should log confidence scores where applicable, and that models should have an appropriate threshold so that very low-confidence cases can be flagged as “undecided.”1

Operation

Section 10 covers life in production: change control over the model, the system, and the whole process (including changes to physical inputs), with any decision not to retest fully justified; configuration control with measures to detect unauthorized change; regular performance monitoring; monitoring of input data drift against the sample space; and, where a human operator makes the decision with model input and testing effort was reduced, records of the human review, which may mean review of every output depending on criticality.1

Draft SectionWhat It Asks ForWho Owns It in the Draft
1. ScopeCritical GMP uses of static, deterministic ML models; GenAI and LLMs excluded from critical use; HITL for non-critical useRegulated user
2. PrinciplesCross-functional cooperation, supplier documentation reviewed, risk-based effortRegulated user
3. Intended useDetailed use, input sample space, subgroups, limitations, operator responsibility for HITLProcess SME, approved before testing
4. Acceptance criteriaMetrics and criteria no lower than the process replacedProcess SME, approved before testing
5. Test dataRepresentative, stratified, sized, correctly labeled; generated data discouragedNot assigned to a role
6. Test data independenceTest data kept from development; access control, audit trail, no copies, usage log, staff separationNot assigned to a role
7. Test executionApproved test plan, deviations justified, records retainedProcess SME involved in plan
8. ExplainabilityFeature attribution during testing of critical models; feature review at approvalNot assigned to a role
9. ConfidenceConfidence logging where applicable; threshold and “undecided” handlingNot assigned to a role
10. OperationChange and configuration control, performance and drift monitoring, human review recordsNot assigned to a role

One small detail tells you how far from final this text is. The glossary definition of “AI system” is copied word for word from Article 3(1) of the EU AI Act, including the phrase that such a system “may exhibit adaptiveness after deployment.”110 The scope section then excludes adaptive models from critical use. Both can be true at once (the glossary defines the broad category and the scope narrows it), but it is the kind of inconsistency a final edit tends to remove.

A Decision Rule for Settle Now Versus Wait

With a draft this likely to change in places, the practical question is not “should we comply with Annex 22 yet?” It cannot be complied with, because it does not exist in final form. The question is which pieces of work are worth doing now because they will hold, and which are worth holding back because they may be rewritten.

We use three tests. An item is a settle-now item if it passes at least two of them.

Test 1

It Already Follows From Something Final

The requirement is already expected by a binding or final text: current EU GMP (Chapter 1, Chapter 4, the 2011 Annex 11), ICH Q9(R1), or EMA’s final 2024 reflection paper on AI. If the draft only restates it for AI, the final Annex 22 is unlikely to remove it.

Test 2

It Is Not What EMA Is Reconsidering

EMA has said publicly what it is looking at: scope for adaptive, probabilistic, and generative models, guardrails, human oversight under guardrails, and lifecycle and supplier questions. Items outside that list are less likely to move.

Test 3

It Cannot Be Fixed Later

Some work can be done after the fact. Some cannot. A test set that developers have already seen cannot be made independent again, and a baseline for the manual process you replaced cannot be measured after it is gone.

Wait Signal

It Depends on Wording Still in Play

Heavy investment tied to a specific clause number, a specific technique the draft names only “where applicable,” or a scope boundary under review is a candidate to hold until the final text, with a light placeholder in the meantime.

The first test deserves one more line of support. EMA’s reflection paper on AI in the medicinal product lifecycle is final, issued in September 2024 by its human and veterinary medicines committees (CHMP and CVMP). Its manufacturing section says model development, performance assessment, and lifecycle management should follow quality risk management principles, with ICH Q8, Q9, and Q10 considered “awaiting revision of current regulatory requirements and GMP standards.”11 Annex 22 is that revision. Anything the reflection paper already expects is a safe base.

Applied to the draft, the three tests sort cleanly. Four areas come out as settle-now: inventory and criticality, intended use with acceptance criteria and a measured baseline, test data independence, and designed human review. The next three sections cover each. The wait-for list follows after that.

Settle Now: Inventory, Criticality, and Intended Use

Why the Inventory Comes First

Every other decision depends on knowing which models you have and which of them are in critical GMP uses. The draft’s entire scope turns on the phrase “critical applications with direct impact on patient safety, product quality or data integrity.”1 Whatever EMA decides about generative AI, that phrase is very likely to remain the test for scope. If the final text opens a pathway for probabilistic models under guardrails, you will still need to know which of your uses are critical, because that is where the guardrails would apply.

An inventory also passes the first test. The Annex 11 in force today (the 2011 version) already says an up-to-date listing of all relevant systems and their GMP functionality should be available, which it calls an inventory.12 Extending that inventory to AI components is not new work in principle, only in detail.

What to Record for Each AI Use

Keep this practical. For each model or AI feature touching GMP work, record:

  • The process step and decision it supports, in the words of the process owner, not the vendor.
  • A criticality call: does the output have direct impact on patient safety, product quality, or data integrity? Record the reasoning, not just yes or no.
  • Static or adaptive: are the parameters frozen in production, or does the model update itself from new data?
  • Deterministic or probabilistic: do identical inputs always give identical outputs? For vendor tools, ask; do not assume.
  • Generative or not: is it a generative model or LLM, including AI features embedded in document management, QMS, or LIMS products?
  • Human role: does a person make the final decision, and did anyone reduce testing on the strength of that review?
  • Supplier: who trained and tested it, and do you hold their documentation? Section 2.2 of the draft says you must have it and review it either way.1

The generative line is where many inventories fall short. Teams list the machine vision model on the packaging line and forget the drafting assistant inside the deviation module. Under the draft, the drafting assistant is not banned, but it needs a qualified person accountable for its outputs, and you cannot assign one to a tool you have not listed. Our earlier pieces on building an AI model registry and on shadow AI discovery go deeper on the mechanics.

Align Your Definitions With the Ones Regulators Use

Because the draft glossary uses the EU AI Act’s definition of an AI system word for word, you can use one definition in your inventory SOP for both purposes.10 Use the draft’s own terms for the scope labels (static, dynamic, deterministic, probabilistic, critical, non-critical), and write your SOP so that the labels can be re-mapped if the final text changes the boundaries. A label that says “probabilistic: excluded from critical use per draft Annex 22 section 1” will need editing. A label that says “probabilistic” plus a separate field for “current regulatory position” will not.

Write the Intended Use Now, and Have the Right Person Own It

Intended use is the second settle-now item, and it passes all three tests. It is already a basic expectation of computerized system validation, it is not on EMA’s reconsideration list, and a good one is hard to write after a model is in production because the people who knew the edge cases have moved on.

The draft is specific about ownership: a process SME is responsible for the adequacy of the intended use and for the acceptance criteria, and both are approved before acceptance testing.1 In many organizations, these documents are drafted by the data science team or the vendor and signed by QA. That is the pattern to change now. The data scientist can write the metrics section; the process expert must own the description of the input sample space, the subgroups, and the limitations, because only they know which rare variations matter on the line.

Measure the Baseline You Will Be Compared Against

The “no decrease” clause in Section 4.3 is the item most likely to find organizations unprepared. Acceptance criteria must be at least as high as the performance of the process the model replaces, which “implies, that the performance should be known for the process which is to be replaced by a model.”1 If a model replaces manual visual inspection, you need a defensible measure of how well human inspectors perform today, by subgroup, with the same metrics you will use for the model.

The principle is not new. The revised draft Annex 11 has its own “no risk increase” clause saying a system replacing a manual operation should cause no decrease in product quality, patient safety, or data integrity.13 What is new is the expectation of a number. This is the clearest case of Test 3: once the manual process is retired, the baseline can only be reconstructed from historical records, which are rarely good enough. If a manual process is on your automation roadmap for 2027, start measuring it this quarter.

Settle now, part one. A complete AI inventory with a reasoned criticality call and scope labels. An intended use for each critical model, owned and approved by a process SME, with the input sample space and subgroups described. A measured baseline for every manual process you plan to replace, using the metrics the model will be judged on.

Settle Now: Test Data Independence

If you do only one thing from this article before the final text arrives, make it this one. Test data independence is the clearest example of work that cannot be done after the fact. Once the people who build a model have seen the test set, the test set is compromised for that model, and no document can restore it.

It Is Already Expected by a Final EMA Text

This is not an Annex 22 novelty. EMA’s final reflection paper describes the same discipline in plain terms: after model selection and tuning, the final performance is evaluated once using the hold-out test set, and if the model needs further development, the current test set cannot be reused and “a completely new and independent test dataset is required” to repeat the test for the updated model.11 The reflection paper’s glossary also names data leakage from the test set into the development environment as a risk to manage.11 So Test 1 is met by a text that is already final, and Test 3 is met by the nature of the problem. Nothing EMA has published about its reconsideration touches this area.

What to Put in Place

1

Choose the Independence Method Per Model

The draft allows two routes: split test data from the full pool before training begins, or capture test data only after training and validation are complete. For a new line or a new product, capturing later is often simpler. For a model trained on years of historical images, an early split is usually the only option. Record the choice and the reason.

2

Put Test Data Behind Its Own Access Control and Audit Trail

A separate repository, access limited to named people who are not developing or training the model, and an audit trail that logs both access and changes. The draft also says there should be no copies outside that repository, so check for extracts in notebooks, shared drives, and vendor environments.

3

Log Every Use of the Test Set

The draft asks you to record which data was used for testing, when, and how many times. This log is what lets you show that a test set was used once for a given model version, which is the point the reflection paper makes about reuse.

4

Handle Physical Test Objects Deliberately

For inspection models, the test set may be physical units or samples. Keep a register so the units used in the final test were never used in training or validation, unless you can show the features are independent.

5

Separate Staff, or Pair Them

People with access to test data should not work on training or validation of the same model. In small teams where that is impossible, the draft allows a person with test data access to work on training or validation only in a pair with a colleague who has not had that access.

What About Supplier Models?

Most pharma and biotech companies will buy more AI than they build. The draft is clear that documentation for training, validation, and testing must be available to and reviewed by the regulated user whether the work was done in-house or by a supplier.1 For test data independence, that means your supplier assessments should ask how the vendor kept its test set away from its developers, whether there is an access log, and whether you can see the test records. It also means that when you test a vendor model on your own data before go-live, your own test set needs the same protections. A vendor’s engineers tuning a model on your site data during implementation can compromise your acceptance test just as easily as your own staff can.

Do Not Generate Test Data to Fill Gaps

Section 5.6 says generating test data or labels, for example with generative AI, is not recommended and any use must be fully justified.1 If your inventory work shows that rare subgroups have too few real examples for adequate statistical confidence, the answer the draft points to is collecting more real data or narrowing the intended use, not synthesizing examples. That is worth knowing now, because collecting real rare-event data takes months.

Why this cannot wait. Every week a development team works with open access to what will become the test set is a week of potential leakage you will later have to explain. Setting up the repository, access control, and log takes little effort when a project starts and cannot be done retroactively once a model is built.

Settle Now: Human Review That Is Designed and Recorded

Human review appears in three places in the draft, and in each it is more demanding than “a person looks at it.”

Three Places the Draft Relies on People

First, in scope: for generative AI and LLMs in non-critical GMP use, qualified and trained personnel should always be responsible for making sure outputs are suitable for the intended use.1

Second, in intended use (Section 3.3): where a model gives input to a decision made by a human operator, and the effort to test the model has been reduced for that reason, the intended use should include the operator’s responsibility, and the operator’s training and consistent performance should be monitored “like any other manual process.”1

Third, in operation (Section 10.5): in the same situation, records of the human review should be kept, and depending on criticality and the level of model testing, this may mean consistent review or testing of every model output, according to a procedure.1

The logic is consistent. If you rely on a person to catch model errors, and you tested the model less because of that reliance, then the person becomes part of the control, and a control needs to be defined, trained, measured, and recorded.

Why This Is Safe to Settle Even With Scope Under Review

Human oversight is on EMA’s reconsideration list, but in a specific way. One of the six workshop topics asks what level and form of human-in-the-loop oversight is still required when guardrails are in place, and whether that oversight is enough for accuracy, traceability, and accountability.7 The question is how much human oversight is needed alongside new controls, not whether it is needed. Any pathway that opens generative or probabilistic models to wider GMP use is likely to lean on human review more, not less. Commentators have made the same point: writing in Pharmaceutical Technology in August 2026, Brian Drapeau argued that if EMA opens a pathway, qualified personnel confirming output suitability will be the central control.14 That is one author’s view, but it matches the direction of EMA’s own questions.

So designing human review well now passes Test 2 in substance, and it passes Test 1 because training personnel for their duties is already a basic GMP expectation under Chapter 2 (Personnel).

What Good Looks Like

  • A written review procedure per use, saying what the reviewer checks, against what source, and what they do when they disagree with the model.
  • Reviewer qualification: the draft’s words are adequate qualification and training. Define the qualification for each use and train to it.
  • Records that show review happened, not just an approval signature. For higher-criticality uses, record what was reviewed and the outcome for each output.
  • Monitoring of the reviewer: agreement rates between reviewer and model, override rates, and periodic checks on a sample of accepted outputs. A reviewer who never overrides the model is a signal to investigate, not a sign that the model is perfect.
  • A link back to testing: if testing was reduced because of human review, the validation record should say so, so that a later change to the review step triggers a look at the testing.

We covered what FDA and EMA expect from human oversight more broadly in our earlier piece on human-in-the-loop requirements. The Annex 22 angle adds one specific point: the draft treats the human as a manual process to be monitored, which means the human’s performance is data you should already be collecting.

Settle now, part two. Test data independence controls on every model in development, including vendor models tested on your data. A human review design for every AI use where a person makes the decision, with qualification, records, and reviewer monitoring. These hold whatever EMA decides about generative AI.

What to Wait For: The Questions Still Open

The other half of the decision is restraint. Some parts of the draft are likely to change, and investing heavily in them now risks rework. Waiting does not mean doing nothing; it means doing a light version and keeping options open.

Whether Generative, Probabilistic, and Adaptive Models Can Be Used in Critical Applications

This is the largest open question, and it is open by EMA’s own account. The workshop questions asked how adaptive and probabilistic models could be accommodated in Annex 22, what the validation approach for adaptive models should be, how reliably guardrails can prevent or contain hallucinations and fabricated data, and whether there are classes of critical decisions where no combination of guardrails and oversight would be enough.7 Those are open questions, and the answers could fall anywhere from “keep the exclusion” to “allow with conditions.”

What to do meanwhile: do not plan a critical GMP use of an LLM on the assumption that the exclusion will be lifted. Do not abandon non-critical uses either, since the draft already permits them with a qualified person in charge. For any generative use you are piloting near the critical boundary, document the risk assessment and the controls in a way that could be reused as evidence if a guardrail pathway opens. Our earlier article on the June 2026 expert workshop covers the workshop itself in more detail.

The Final Wording on Explainability and Confidence

Sections 8 and 9 name specific techniques (SHAP, LIME, heat maps) and specific behaviors (confidence logging, an “undecided” outcome), but qualify them with “where applicable” and tie them to testing.1 If the scope widens to probabilistic or generative models, these sections will almost certainly need rewording, because feature attribution and a single confidence score do not map neatly onto those models. Even for classic classifiers, the final text may tighten or loosen the wording.

What to do meanwhile: for critical classifiers in development, capture feature attribution and confidence during testing, because it is good practice and inexpensive to add at that stage. Do not buy a dedicated explainability platform or retrofit attribution into production systems solely to meet a draft clause. We discussed what inspectors tend to accept as explainability evidence in a separate article.

Cross-References to Annex 11 and Chapter 4

Annex 22 is written as an add-on to Annex 11, and Annex 11 is itself being finalized. The drafts do not yet line up perfectly. Section 4.3 of draft Annex 22 points the reader to “Annex 11 2.7” for the principle that performance of the replaced process should be known.1 In the consultation draft of Annex 11, Section 2.7 is about security, and the “no risk increase” principle is Section 2.8.13 That is an ordinary drafting mismatch that the final texts will resolve, but it shows why you should not copy clause numbers from the drafts into SOPs, templates, or traceability matrices yet.

What to do meanwhile: map your procedures to the draft by topic, not by clause number. When all three final texts appear, update the references in one pass. Our Annex 11 gap assessment template and Annex 22 CSV deliverables mapping are built to be re-pointed this way.

Security, Tampering, and Outsourced AI

The draft’s operation section asks for configuration control and “effective measures” to detect unauthorized change, but says little about cloud-hosted models or guardrail services outside the manufacturer’s control.1 EMA’s sixth workshop topic asked directly whether protection of AI systems from tampering needs to be in Annex 22 or is already covered by Annex 11, and how supplier qualification, change control visibility, and audit rights can be kept in cloud-based AI supply chains.7 The answer could add text to Annex 22, or it could leave the topic to Annex 11.

What to do meanwhile: apply your existing Annex 11 supplier and security controls to AI vendors now, since those are in force. Hold off on writing AI-specific contract clauses for guardrail infrastructure until you see where the final text puts them.

The Date It Comes Into Operation, and PIC/S Adoption

No source states when Annex 22 will come into operation or whether there will be a transition period. The Annex 1 precedent of one year, with two years for one point, is a reasonable planning assumption, not a forecast.9 For companies manufacturing in or supplying to PIC/S member countries outside the EU, adoption into the PIC/S GMP Guide is a separate step, and we found no date announced by PIC/S.3

What to do meanwhile: plan for a transition window of about a year after publication, and do not promise leadership a compliance date until the published annex states one.

Open QuestionWhy It Is OpenDo NowHold Until Final Text
GenAI, LLM, probabilistic, adaptive models in critical useSubject of EMA’s 30 June to 1 July 2026 workshopKeep non-critical uses with a qualified person accountable; document risk assessmentsAny critical GMP use of an LLM
Explainability and confidence wording“Where applicable” language; likely rewording if scope widensCapture attribution and confidence during testing of new classifiersPlatform purchases and production retrofits
Clause numbering and cross-referencesAnnex 11 and Chapter 4 finalizing in parallel; draft references do not alignMap procedures by topicDraft clause numbers written into SOPs and matrices
Tampering, cloud, outsourced guardrailsWorkshop asked whether Annex 22 or Annex 11 should cover itApply current Annex 11 supplier and security controlsAI-specific contract language for guardrails
Date of coming into operationNot announced; handoff to the Commission is only the targetPlan on roughly a year after publicationCommitting to a compliance date

A caution on outside guidance. The ISPE GAMP Guide: Artificial Intelligence, published in July 2025, is a useful industry framework for the lifecycle of AI-enabled computerized systems.15 It is not a regulation and it is not Annex 22. Use it to structure your work, but do not assume that anything it recommends will match the final annex word for word, and do not cite it to an inspector as the requirement.

A Sequenced Plan for the Months Before Publication

Here is how we would sequence the work for a pharma or biotech manufacturer with a handful of AI uses in or near GMP, starting in the fourth quarter of 2026. The order follows the decision rule: the things that cannot be fixed later come first.

1

Protect Test Sets on Anything in Development

For every model currently being built or tuned, including vendor models on your data, confirm the independence method, move the test set into a controlled repository with an audit trail, remove stray copies, and start the usage log. This is first because every week of delay is potential leakage.

2

Complete the Inventory and Criticality Calls

Include embedded AI features in commercial GMP software. Record static or adaptive, deterministic or probabilistic, generative or not, and the human role. Have QA and the process owner agree on each criticality call.

3

Start Measuring Baselines for Processes on the Roadmap

For each manual process you expect to automate with a model in the next two years, agree the metrics with the process SME and start collecting performance data now, by subgroup.

4

Rewrite Intended Use Documents With SME Ownership

For critical models, make the process SME the owner of the intended use and acceptance criteria, with the input sample space and subgroups written out. Approve before any further acceptance testing.

5

Design Human Review Where People Make the Decision

Write review procedures, define reviewer qualification, set up records, and begin monitoring agreement and override rates. Link each review design to any testing that was reduced because of it.

6

Record the Open Items With Light Placeholders

Keep a short register of wait-for items (generative scope, explainability wording, cross-references, cloud and guardrail controls, effective date) with an owner and a trigger: the publication of the final annex. Review it when the Commission publishes.

7

When the Final Text Appears, Run One Gap Pass

Compare the published annex against your settle-now work and your parked register together, update clause references across procedures in one pass, and set a plan against the stated date of coming into operation.

If you want to test readiness before the final text arrives, a mock inspection built around the draft works well, provided the team understands it is rehearsing against a draft and focuses on the settle-now areas.

What This Plan Deliberately Leaves Out

It does not include a full Annex 22 compliance program, a new validation SOP written clause by clause against the draft, or a large tooling purchase. Each of those would be built on a text that has not been finalized, and each would need rework if the scope changes. It also does not assume that the generative AI exclusion will be lifted or kept. The four settle-now areas hold under either outcome, which is the point.

Questions to Bring to Your Leadership Team

Most of the decisions above can be made inside quality and IT. A few need leadership, and it is better to raise them now than when the final text appears.

  • Which manual processes are we planning to replace with models in the next two years? The answer sets the baseline measurement work, and that work needs operational time from the people doing the process today.
  • Are we prepared to separate model developers from testers? For small data science teams, this is a staffing question as much as a quality question. The draft’s four-eyes alternative helps, but someone has to plan for the second reviewer’s time.
  • Are any business cases built on using an LLM in a critical GMP decision? If so, leadership should know that the current draft excludes it and that EMA has not yet said whether that will change.
  • What date are we telling the board or the inspectorate? The honest answer today is that no date of coming into operation has been set. Plans should say so.

Conclusion

The Annex 22 final text is not in force, has not been published, and has no date for coming into operation. What exists is a July 2025 draft, an EMA target to give the Commission a final text in Q4 2026, and an open question about whether generative and probabilistic models will be brought inside the annex. Leaders who hear “final by year end” and plan a compliance program to a January deadline are likely to build on wording that changes. Leaders who wait for the final text before doing anything are likely to find that the most important controls, especially test data independence and a measured baseline, cannot be put in place after the fact.

The better path is to separate the two. Settle the inventory and criticality calls, intended use owned by the process expert, test data independence, and designed human review, because each holds under any plausible final text and several are already expected by final EMA guidance. Log the rest with an owner and a trigger. Sakara Digital works with pharma and biotech organizations on exactly this kind of sorting: working out which AI controls to build now and which to hold until the regulation is final. If you are weighing where to start before the Annex 22 final text appears, we are happy to have that conversation.

For Further Reading