In This Article
- Executive Summary
- What Is in Force Today, and What Is Still a Draft
- What the Draft Says About Determinism, Word for Word
- What Deterministic Does and Does Not Mean
- The Plain Test: Four Questions, in Order
- How to Run the Repeatability Check
- Why LLM Tools Fail the Test Even at Temperature Zero
- Practical Consequences for LLM-Based Tools in GMP Work
- How the Test Holds Up if the Final Scope Changes
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
The draft EU GMP Annex 22 on artificial intelligence draws its main line in one short passage. It covers AI models in critical GMP applications only when the model is static and gives identical outputs for identical inputs. It says probabilistic models, dynamic models, generative AI, and large language models should not be used in critical GMP applications at all. Annex 22 is still a consultation draft: nothing in it is in force, and EMA is reconsidering the generative AI question. Even so, the deterministic AI GMP scope line is already shaping how firms classify their tools, and many teams apply it loosely.
This article gives a plain test built from the draft’s own words: four questions, asked in order. Is the use critical? Is the tool generative AI or an LLM? Is the model static? Does it give the same output for the same input when you measure it? The last question is an experiment, not a vendor claim. We explain how to run it, and why “deterministic” does not mean “correct” and does not rule out a confidence score.
We then look at LLM-based tools. Vendor documentation from Anthropic, OpenAI, Microsoft, and Google says outputs are not guaranteed to repeat, even at temperature zero, and published research shows why. We set out what that means for LLM tools in GMP work today, what US rules expect instead, and how the test holds up if the final Annex 22 text widens its scope.
What Is in Force Today, and What Is Still a Draft
Before applying any test, it helps to be exact about which rules exist. A lot of industry commentary talks about Annex 22 as though it already binds manufacturers. It does not.
Annex 22 Is a Consultation Draft
Annex 22, titled “Artificial Intelligence,” is a proposed new annex to the EU Guidelines for Good Manufacturing Practice in EudraLex Volume 4. The European Commission published it for stakeholder consultation in 2025, alongside revised drafts of Chapter 4 (Documentation) and Annex 11 (Computerised Systems), and that consultation has closed.14 As of 29 September 2026, the Commission’s EudraLex Volume 4 page lists no Annex 22, and it still shows Annex 11 and Chapter 4 in their 2011 versions.2
EMA’s GMP/GDP Inspectors Working Group has set a target in its current three-year work plan: Q4 2026, “to provide the European Commission with a final text for the amended annex in order to assure use of artificial intelligence in the context of GMP.”3 That target is a handoff to the Commission. It is not publication, and it is not a date on which anything comes into force. The same plan lists training for EU GMP inspectors to support implementation of the revised Chapter 4, Annex 11, and Annex 22 during 2026 to 2028.3
EMA Is Reconsidering the Exact Line This Article Is About
On 30 June and 1 July 2026, EMA held a two-day multistakeholder workshop to gather expert input for Annex 22. EMA’s event page explains why. The 2025 consultation “suggested support for potentially enabling the use of technologies such as generative AI (GenAI) or large language models (LLMs) in medicines manufacturing,” while the draft had said that dynamic, adaptive, and probabilistic models should stay out of critical GMP applications. The page states that “EMA is still considering the implications of the stakeholder consultation results,” and that the drafting group is seeking input on “possible control and mitigation measures such as guardrails.”4
EMA said it expects the workshop to produce a report with expert input. At the time of writing, that report is not posted on the event page, and no revised text of Annex 22 has been published.4
What Binds You Today
Whatever category an AI tool falls into, the rules for computerized systems already apply to it. In the EU, that is the 2011 Annex 11, which covers any computerized system used as part of GMP-regulated activities. In the US, 21 CFR 211.68 requires that computers and related systems used in manufacturing be “routinely calibrated, inspected, or checked according to a written program designed to assure proper performance.” It also says that “Input to and output from the computer or related system of formulas or other records or data shall be checked for accuracy.”5
So the deterministic test in this article does not decide whether you need controls. You need them regardless. It decides which set of AI-specific expectations the draft would add on top, and whether the draft would allow the tool in a given use at all.
A note on scope. A companion article in this series covers what to settle now and what to wait for while the final text is pending. This article stays on one question: how to tell whether a given AI tool is deterministic or probabilistic in the sense the draft uses, and what follows from the answer.
What the Draft Says About Determinism, Word for Word
The scope section of the draft is four short paragraphs. Because the deterministic line carries so much weight, it is worth reading the exact words rather than a summary.
The Critical-Use Gate Comes First
The first paragraph sets the overall reach. The annex applies to computerized systems used in manufacturing medicinal products and active substances “where Artificial Intelligence models are used in critical applications with direct impact on patient safety, product quality or data integrity, e.g. to predict or classify data.” It describes itself as additional guidance to Annex 11, and it covers machine learning models “which have obtained their functionality through training with data, rather than being explicitly programmed.”1
Two things follow. First, a fixed rules engine or a conventional statistical calculation that no one trained on data is not what this annex is about. Annex 11 still covers it. Second, everything else in the scope section is framed around critical applications. The model-type limits only bite once the use is critical.
The Static and Deterministic Limits
The next two paragraphs set the model-type limits. On static models, the draft says:
“The document applies to static models, i.e. models that do not adapt their performance during use by incorporating new data. The use of dynamic models which continuously and automatically learn and adapt performance during use, is not covered by this document, and should not be used in critical GMP applications.”1
On deterministic output, it says:
“The document applies to models with a deterministic output which, when given identical inputs, provide identical outputs. Models with a probabilistic output which, when given identical inputs, might not provide identical outputs are not covered by this document and should not be used in critical GMP applications.”1
Notice how the draft defines the terms. It does not define “deterministic” by the type of algorithm, the architecture, or the math inside the model. It defines it by observed behavior: identical inputs, identical outputs. And it defines “probabilistic output” by the possibility of different outputs for identical inputs. The word “might” matters. A model does not have to vary often to fall on the probabilistic side. It only has to be capable of varying.
The glossary defines “Static” as “Frozen model: A model where all parameters have been finally set, not allowing further adaption to new data.” It gives no separate glossary entry for deterministic or probabilistic, so the scope sentences above are the only definitions the draft offers.1
Generative AI and LLMs Are Excluded by Name
The fourth paragraph names two technologies:
“Following the above, the document does not apply to Generative AI and Large Language Models (LLM), and such models should not be used in critical GMP applications.”1
The phrase “Following the above” tells you the drafters’ reasoning: they treated generative AI and LLMs as examples of models that are dynamic or probabilistic, or both. But the exclusion itself is written as a category. It does not say “to the extent that such models are probabilistic.” This detail comes back later, because it means an LLM that you have engineered to repeat its outputs exactly would still be outside the draft’s scope as written.
The Non-Critical Carve-Out
The same paragraph continues with the sentence that most summaries skip:
“If used in non-critical GMP applications, which do not have direct impact on patient safety, product quality or data integrity, personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use, i.e. a human-in-the-loop (HITL) and the principles described in this document may be considered where applicable.”1
So the draft does not ban generative AI or LLMs from GMP work. It keeps them out of critical applications and puts a qualified person in charge of their outputs everywhere else.
Two Other Places the Draft Touches Generative AI
The draft mentions generative AI once more outside the scope section. In the test data section, it says: “Generation of test data or labels, e.g. by means of generative AI, is not recommended and any use hereof should be fully justified.”1 That applies even when the model under test is a conventional, deterministic classifier. If a team uses an LLM to write synthetic test cases or to label images for a vision model’s test set, the draft expects a written justification.
The confidence section is also relevant, though it does not name generative AI. It asks that, when a model used to predict or classify data is tested, the system log a confidence score for each outcome where applicable, and that such models have a threshold setting so that predictions are made only when suitable. When the confidence score is very low, the draft suggests considering whether the model should flag the outcome as undecided rather than make an unreliable call.1 We return to this below, because a confidence score is often mistaken for a sign that a model is “probabilistic.”
What Deterministic Does and Does Not Mean
The draft’s definition is simple. The confusion comes from the fact that “probabilistic” has several everyday meanings in machine learning, and only one of them matches the draft. Here are the four mix-ups we see most often.
A Probability Score Is Not a Probabilistic Output
A classifier that reports “reject, confidence 0.94” is producing a probability estimate. If the same image always produces 0.94 and “reject,” the output is deterministic in the draft’s sense. The draft itself asks for confidence scores, so it plainly expects deterministic models to produce them.
Deterministic Does Not Mean Correct
A model can give the same wrong answer every time. Determinism is about repeatability, not accuracy. The draft handles accuracy separately, through acceptance criteria, test data, and a rule that performance be at least as high as the process it replaces.
Static and Deterministic Are Different Tests
Static asks whether the model changes over time by taking in new data. Deterministic asks whether a single, unchanged model gives the same answer twice. A frozen model can still vary run to run, and a model that repeats perfectly today can be retrained tomorrow.
“Identical Inputs” Means the Digital Input
Two photos of the same vial are never identical at the pixel level. The draft’s test is whether the model, given the exact same data, returns the exact same output. Variation in the physical world is handled by test data design and monitoring, not by the determinism test.
Where the Probability Lives in a Language Model
The reason LLMs are the standard example of probabilistic output is how they produce text. A language model does not output a sentence directly. At each step, it computes a probability for every possible next word piece (a token), and then a decoding method picks one. Patrick von Platen’s widely used Hugging Face explainer describes the two basic choices. Greedy search “selects the word with the highest probability as its next word.” Sampling, in contrast, means “randomly picking the next word” according to those probabilities, and the author notes that language generation using sampling “is not deterministic anymore.”8
The “temperature” setting controls how much randomness the sampling step uses. A temperature of zero, in principle, reduces the choice to greedy search: always take the most likely token. That is why many teams assume that setting temperature to zero makes an LLM deterministic. As the next sections show, the vendors themselves say it does not.
Where Nondeterminism Hides in Conventional Models
Conventional deep learning models, such as the image classifiers used in visual inspection, have no sampling step at inference time. They are usually deterministic in the draft’s sense. Usually is not always. PyTorch, one of the most widely used frameworks, states in its reproducibility notes: “Completely reproducible results are not guaranteed across PyTorch releases, individual commits, or different platforms. Furthermore, results may not be reproducible between CPU and GPU executions, even when using identical seeds.”15
The same notes describe settings that reduce the problem. torch.use_deterministic_algorithms() makes PyTorch use deterministic algorithms where available and raise an error if an operation is known to be nondeterministic with no deterministic alternative. Disabling the cuDNN benchmarking feature makes the library pick a convolution algorithm deterministically, possibly with lower performance.15 These are engineering choices your supplier or data science team makes. A quality reviewer should ask whether they were made, and see the evidence.
The practical point. “Deterministic” is a property of the whole deployed stack (model weights, software libraries, hardware, and settings), not just of the algorithm on paper. A deterministic classifier moved from one GPU type to another, or upgraded to a new framework version, may produce slightly different raw scores. Under the draft’s change control expectations, that kind of change should be documented and evaluated for retesting.1
The Plain Test: Four Questions, in Order
Here is the test we suggest for sorting AI tools against the draft. It uses only the draft’s own words, and the order matters. Each question is only worth asking if the one before it did not already settle the answer.
Is the Use Critical?
Does the AI output have direct impact on patient safety, product quality, or data integrity? The draft’s examples are predicting or classifying data. If the honest answer is no, the model-type limits do not apply. Any kind of model, including an LLM, may be used, with a qualified person responsible for its outputs. Stop here, and record the reasoning.
Is the Tool Generative AI or an LLM?
If the use is critical and the tool is built on a generative model or a large language model, the draft says it should not be used in that application. No repeatability result changes this, because the exclusion is written by category. Stop here and either redesign the use so it is not critical, or wait for the final text.
Is the Model Static?
Are all the model’s parameters fixed while it is in use? Does anything retrain, fine-tune, or update the model automatically from new production data? If the model learns during use, it is dynamic and the draft says it should not be used in critical applications. A model retrained under change control, then frozen again, is still static.
Does It Give Identical Outputs for Identical Inputs, When Measured?
Run the same inputs through the deployed system repeatedly, under realistic conditions, and compare the outputs. If they are identical, the model is deterministic in the draft’s sense and the full set of Annex 22 expectations would apply. If they can differ, the draft says the model should not be used in critical applications.
What Each Outcome Means
| Result of the Test | What the Draft Would Mean | What Applies Today |
|---|---|---|
| Non-critical use, any model type | Outside the model-type limits. A qualified, trained person is responsible for outputs (HITL). Annex 22 principles “may be considered where applicable.” | Annex 11 (EU) or 21 CFR 211.68 and Part 11 (US), in proportion to risk. |
| Critical use, generative AI or LLM | Should not be used in the critical application. Not in scope of the annex. | Same computerized system rules. No specific AI rule permits or forbids it yet. |
| Critical use, dynamic model | Should not be used in the critical application. | Same computerized system rules, which already expect change control over the system. |
| Critical use, static, repeatability not shown | Our reading: treat it as probabilistic output until repeatability is shown, since the draft covers only models with deterministic output. Should not be used in the critical application. | Same computerized system rules. |
| Critical use, static, repeatability shown | In scope. Sections 2 to 10 apply: intended use, acceptance criteria, independent test data, explainability, confidence, and operational controls. | Same computerized system rules. Annex 22 would add AI-specific detail on top. |
Why Criticality Goes First
Teams often start with the technology question (“is it an LLM?”) because it is easier to answer. Starting there leads to two mistakes. The first is rejecting useful tools for low-risk work, such as drafting a training module outline, because they are LLMs. The draft does not require that. The second is more serious: approving a conventional model for a critical use without noticing that the use is critical, because the model “is not generative AI anyway.”
The criticality question is where the most judgment is needed, and the draft gives three triggers: patient safety, product quality, and data integrity. Data integrity is the broadest of the three. A tool that does not decide anything about product, but writes, alters, or summarizes content that becomes part of a GMP record, can have a direct impact on data integrity. That is our reading of the draft rather than a statement in it, but it is the reading we would expect an inspector to test.
Why “Is It an LLM?” Comes Before the Measurement
It might seem more rigorous to measure every tool. But for generative AI and LLMs, the measurement does not change the answer under the current draft, because of the category exclusion discussed above. Running a repeatability study on an LLM used in a critical application would produce a useful engineering fact and no change in regulatory position. The time is better spent redesigning the use.
Where the Line Gets Blurry
Some tools do not sort cleanly. A few examples, with our view on each:
- A vision model built on a large pretrained “foundation” model, then fine-tuned to classify defects. If it outputs a class and a score, has no text generation step, and is frozen, it behaves like a conventional classifier. We would treat it under questions 3 and 4, and document why it is not “generative AI” in the draft’s sense.
- An LLM used only to turn free text into fixed categories (for example, deviation type). It is still an LLM. If the categorization drives a critical decision, question 2 applies. If a qualified person reviews and confirms each category before it takes effect, the use may be non-critical by design.
- A conventional model with a generative component elsewhere in the system, such as a chatbot front end that explains results. Assess each model separately. The draft notes that “Models may consist of several individual models, each automating specific process steps in GMP.”1
How to Run the Repeatability Check
Question 4 is the only one that calls for an experiment. The draft does not describe how to show determinism, so what follows is our suggested method, built to be proportionate and to leave a clear record.
Test the Deployed System, Not the Development Notebook
A result from a data scientist’s laptop says little about the production system. Run the check on the system as it will operate: the same model file, the same serving software, the same hardware class, and the same preprocessing steps. If the model runs behind an API, call the API. If it runs inside an instrument or an inspection machine, feed it stored inputs through that path.
Use a Fixed, Stored Input Set
Capture a set of inputs as files: images, spectra, time series, or records. Include the subgroups defined in the intended use, and include edge cases near decision thresholds. Borderline inputs are where tiny numeric differences are most likely to flip a decision, so they are the most informative part of the check. Store the set with a checksum so you can show later that each run used exactly the same data.
This set is not the independent test set the draft describes for performance testing. It can be drawn from data that the development team has seen, because its only purpose is repeatability. Keep the two clearly separated in your records, so there is no question about test data independence.
Repeat Under Realistic Variation
Run the input set many times, and vary the conditions that change in real operation:
- Different times of day and different system load, including runs alongside other work if the serving hardware is shared.
- After a restart of the service or the machine.
- Across each hardware unit, if the model is deployed on more than one.
- With inputs submitted one at a time and in batches, if the system supports both.
The last item matters more than it looks. As the next section explains, research on LLM serving found that the main cause of run-to-run variation was the batch size changing with server load. Conventional models served with batching can show the same effect on raw scores.
Compare at Two Levels
Compare the final output (the class, the decision, the flag) and the raw output (the score or value before thresholding). There are three possible results:
- Both identical across all runs. The clearest evidence of deterministic output.
- Decisions identical, raw scores differ in distant decimal places. This is common with GPU inference. The draft’s words are “identical outputs,” so do not treat this as a pass by default. Decide before testing which output is the one the process uses, document the rationale, and show that no input in your set, including the borderline ones, changed its decision. Then make deterministic settings part of the configuration under change control, and retest.
- Decisions differ. The model has probabilistic output in the draft’s sense. It should not be used in a critical application until the cause is found and removed.
Record the Environment
A repeatability result is only meaningful for the configuration that produced it. Record the model version and file checksum, framework and library versions, deterministic settings, hardware type, and serving configuration. The draft already expects configuration control of a tested model and measures to detect unauthorized change.1 The repeatability record fits naturally alongside that.
What good evidence looks like. A short protocol approved before the run, a checksummed input set that includes borderline cases, results from repeated runs under varied load and after restarts, a comparison at both decision and raw-score level, and a signed conclusion that states which configuration was tested. That package answers question 4 and supports the draft’s change control and configuration control expectations at the same time.
Repeat the Check When Things Change
Determinism can be lost through ordinary maintenance: a framework upgrade, a driver update, a new GPU model, a change in batching. Under the draft, any change to the model, the system, or the process should be documented and evaluated to decide whether the model needs retesting, and any decision not to retest should be fully justified.1 Add the repeatability check to the list of tests considered in that evaluation.
Why LLM Tools Fail the Test Even at Temperature Zero
The draft excludes LLMs by name, so for critical uses the repeatability result does not matter. For every other use, it matters a great deal, because it affects how much you can rely on an LLM’s output being the same when a reviewer checks it again, when an investigator reproduces it, or when an auditor asks how a record was produced.
The Vendors Say So Directly
The major model providers are candid about this in their own documentation:
- Anthropic. Its glossary says: “Even with temperature set to 0, the results will not be fully deterministic and identical inputs may produce different outputs across API calls.” It adds that this applies both to Anthropic’s own inference service and to inference through third-party cloud providers.9
- Microsoft (Azure OpenAI). Its guide to reproducible output says of the seed parameter: “Determinism isn’t guaranteed,” and that even when the seed and the backend fingerprint are the same, “it’s currently not uncommon to still observe a degree of variability in responses.”10
- OpenAI. Its cookbook on the seed parameter describes the result as “(mostly) consistent outputs” and states: “Determinism is not guaranteed, and you should refer to the system_fingerprint response parameter to monitor changes in the backend.”11
- Google (Vertex AI). Its parameter documentation says that at a temperature of 0, “responses for a given prompt are mostly deterministic, but a small amount of variation is still possible,” and that with a fixed seed, “Deterministic output isn’t guaranteed.”12
Under the draft’s definition, a model whose outputs “might not” be identical has probabilistic output. By the vendors’ own descriptions, hosted LLMs meet that definition even at their most conservative settings.
The Cause Is Not Just Sampling
Many people assume the variation comes from randomness in sampling, and that setting temperature to zero removes it. Research published in September 2025 by Horace He and colleagues at Thinking Machines Lab explains why that is not enough. They ran the Qwen3-235B model 1,000 times at temperature 0 with the same prompt, generating 1,000 tokens each, and got 80 unique completions. The completions were identical for the first 102 tokens and first diverged at the 103rd.13
Their explanation is about arithmetic, not dice. Computers add decimal numbers with small rounding errors, and the result can depend on the order of the additions. The software that runs these models changes the order of additions depending on how many requests are processed together (the batch size). The authors conclude that “the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies!”13 In other words, your result can depend on what other users sent to the same server at the same moment.
The Variation Changes Answers, Not Just Wording
It would be reassuring if the differences were only in phrasing. A study by Berk Atil and colleagues, first posted in August 2024 and revised in April 2025, tested five LLMs “configured to be deterministic” on eight common tasks over 10 runs. They report “accuracy variations up to 15% across naturally occurring runs,” a gap of up to 70% between best and worst possible performance, and that “none of the LLMs consistently delivers repeatable accuracy across all tasks, much less identical output strings.”14
For a GMP reviewer, that finding is the important one. An LLM that classifies the same deviation differently on two occasions is not just untidy. It means a validation result from last month is not a reliable prediction of this month’s behavior on the same input.
Deterministic LLM Serving Is Possible, but It Does Not Change the Draft
The Thinking Machines work also showed a fix. Using what they call batch-invariant kernels (software that adds numbers in the same order regardless of batch size), all 1,000 completions became identical. It carried a speed penalty in their test: 1,000 requests to a smaller model took 26 seconds with the default setup, 55 seconds with the unoptimized deterministic version, and 42 seconds with an improved version.13
So a self-hosted LLM can, with specialist engineering, pass question 4. That is useful for reproducibility in non-critical work, where being able to rerun an extraction and get the same answer helps investigations and audits. But under the current draft, it would not bring the LLM into scope for critical use, because the draft excludes LLMs by category. If the final text moves to a risk-based approach, demonstrated determinism may well become one of the controls that counts. That would be a change in the text, not something to assume now.
Static Is a Problem Too, for Hosted Models
Hosted LLMs raise a second issue under question 3. The model you call through an API is not under your change control. Providers update the serving infrastructure, and they retire models. Anthropic’s deprecation page says it “regularly retires older ones” and provides “at least 60 days’ notice before model retirement for publicly released models.”16 OpenAI’s backend fingerprint exists precisely because changes to the serving system can change outputs.11
A model that does not learn from your data during use meets the letter of the draft’s definition of static. But the draft’s operational expectations assume that you control when the model, the system, and the process change.1 With a hosted LLM, you control the model version you request and little else. Build that into the risk assessment, and plan for the retirement date as a known change event.
This concern predates the draft. A 2023 article in ISPE’s Pharmaceutical Engineering by Martin Heitmann, Stefan Muench, Brandi Stockton, and Frederick Blumenthal noted that “traceability of results and content is limited, and sensitivity to input is quite high,” and that retraining and updates “may include other sources of variation in the responses.”17
Practical Consequences for LLM-Based Tools in GMP Work
Put together, the draft’s text and the technical facts lead to a clear set of consequences for pharma and biotech teams using or planning LLM tools.
1. Classify by Use, Then Design the Use
Because the draft allows LLMs in non-critical applications, the most useful work is designing uses so that they are non-critical in fact, not just on paper. The pattern that holds up is simple: the LLM proposes, a qualified person decides, and the decision is recorded. The LLM drafts an investigation summary; the investigator edits and signs it. The LLM suggests a deviation category; the QA reviewer confirms or changes it before it takes effect. The LLM extracts values from a supplier certificate into a form; a person checks each value against the source before the record is approved.
The test for whether a use is non-critical in fact is whether an LLM error could reach the product, the patient, or the record without a person catching it. If it could, the human review is not doing the job the draft assigns to it.
2. Make the Human Review Real
The draft’s non-critical carve-out puts responsibility on “personnel with adequate qualification and training.” For models that give input to a human decision, where the effort to test the model has been reduced, it asks that the operator’s responsibility be written into the intended use, that operator training and performance be monitored “like any other manual process,” and that records be kept of the review.1 Those expectations are written for in-scope models, but they describe what makes human review defensible anywhere. Reviewers need the source material in front of them, enough time, and a way to record that they checked, not only that they clicked approve.
3. Keep Deterministic Steps Deterministic
Many useful tools combine an LLM with conventional software. Where a step can be done with a fixed rule or a calculation, do it that way. If an LLM extracts a numeric result from a document, let validated code, not the LLM, compare that number to the specification. That keeps the pass or fail decision in a deterministic component that can be validated in the usual way, and limits the LLM to a step a person can check.
4. Pin Versions and Log Everything That Shapes the Output
For any LLM output that becomes part of a GMP record, log the model version requested, any backend identifier the provider returns, the full prompt including system instructions and any retrieved documents, the parameters, and the output. Without that, you cannot explain later why the tool said what it said, and because outputs can vary, you cannot recreate it by rerunning. The log is your only reproduction.
5. Treat Retrieval Content as Part of the Input
Many LLM tools look up documents (SOPs, specifications, past deviations) and add them to the prompt. If that document store changes, the “same question” is a different input, and the answer can change even with a fully repeatable model. The draft’s change control section reads: “Any change to the model itself, the system, or the process in which it is used, including any change to physical objects the model is using as input, should be documented and evaluated to determine if the model needs to be retested.”1 For an LLM tool in GMP work, the document store is part of the system.
6. Do Not Use an LLM to Build the Test Evidence for Another Model
As noted earlier, the draft says generating test data or labels by means of generative AI “is not recommended and any use hereof should be fully justified.”1 Using an LLM to label a test set for a visual inspection model would need that justification, even though the inspection model itself is in scope.
7. Know What US Expectations Look Like
US regulators have not drawn a deterministic line. FDA’s January 2025 draft guidance on AI to support regulatory decision-making for drugs and biologics, still marked as a draft that is not for implementation, uses a risk-based credibility framework instead. It assesses model risk as a combination of “model influence” and “decision consequence,” and it mentions that “Repeatability and/or reproducibility studies may help quantify the uncertainty associated with model outputs.”6 Its manufacturing example is an AI visual system checking vial fill volumes, where the model is not the sole determinant of release because independent verification is done on a sample from each batch.6
The joint EMA and FDA “Guiding principles of good AI practice in drug development,” published in January 2026, point the same way: a risk-based approach with “proportionate validation,” a clear context of use, and performance assessments that evaluate “the complete system including human-AI interactions.”7 For firms supplying both markets, the four-question test and a credibility assessment are compatible. The test tells you whether a use would be excluded in the EU draft; the credibility framework tells you how much evidence the use needs.
A trap to avoid. Do not relabel an LLM tool as “decision support” to move it out of the critical category if, in practice, people accept its output without checking. The draft’s HITL expectations assume review that is designed and recorded. An inspector will look at what happens on the floor, not at the label in the system inventory.
How the Test Holds Up if the Final Scope Changes
The final text may well differ from the draft on exactly this point. EMA’s workshop page lists the questions it put to experts. One asks to what extent available guardrail mechanisms can “reliably prevent, detect, or contain hallucinations, incorrect recommendations, or fabricated data when dynamic or adaptive models, probabilistic models, and generative AI / LLMs are used within GMP workflows.” Another asks how guardrails can stay effective “when the underlying AI model evolves (in instances such as updates, retraining and drift).”4
Those questions describe a possible shift from a fixed exclusion to controls matched to risk. We do not know which way EMA will go, and we will not guess. But the four questions are built so that the work stays useful either way:
- Question 1 (criticality) is the foundation of any risk-based approach. A clear, recorded criticality call for each AI use will be needed under any final text.
- Question 2 (generative AI or LLM) may change from a stop sign to a flag that brings in additional controls. Knowing which of your tools fall in this group stays useful.
- Question 3 (static) points to change control and version control, which every version of the text so far expects.
- Question 4 (repeatability) produces evidence that is likely to count as a control, or at least as a measure of risk, if probabilistic models are allowed with guardrails.
The one thing we would not do is plan a critical-use LLM deployment on the assumption that the final text will allow it. The draft says it should not be used, the workshop report is not yet public, and there is no published date for the final annex to come into operation.
Summary of the test. Criticality first. Then the category exclusion for generative AI and LLMs. Then static, checked through change control. Then deterministic, checked by measurement on the deployed system. Record each answer and the reasoning, and repeat the measurement when the configuration changes.
Conclusion
The draft Annex 22 defines deterministic output in one sentence: identical inputs, identical outputs. That definition is plain, testable, and more useful than most of the commentary around it. It turns a debate about model types into a question you can answer with evidence. For conventional models, the evidence is a repeatability check on the deployed system. For LLMs, the vendors’ own documentation and published research already give the answer, and the draft’s category exclusion makes that answer final for critical uses until the text changes. What remains, and where most of the value is, is the first question: deciding honestly which uses are critical, and designing LLM uses so that a qualified person, not the model, owns every output that matters.
Sakara Digital works with pharma and biotech organizations that are sorting their AI tools against draft Annex 22, the 2011 Annex 11, and US expectations at the same time. If you are working through criticality calls, repeatability evidence, or the design of human review for LLM tools, and want an independent perspective on where to start, we are happy to have that conversation.
For Further Reading
For Further Reading
References & Sources
- European Commission. “Annex 22: Artificial Intelligence” (stakeholder consultation draft, EudraLex Volume 4). 2025. https://health.ec.europa.eu/document/download/5f38a92d-bb8e-4264-8898-ea076e926db6_en?filename=mp_vol4_chap4_annex22_consultation_guideline_en.pdf
- European Commission. “EudraLex Volume 4: Good Manufacturing Practice (GMP) guidelines.” Checked 29 September 2026. https://health.ec.europa.eu/medicinal-products/eudralex/eudralex-volume-4_en
- European Medicines Agency, GMP/GDP Inspectors Working Group. “3-year work plan for the Inspectors Working Group.” 2026. https://www.ema.europa.eu/en/documents/other/3-year-work-plan-inspectors-working-group_en.pdf
- European Medicines Agency. “Good manufacturing practice: Multistakeholder workshop on expert contributions to artificial intelligence guidance development (Annex 22).” Event page, 30 June to 1 July 2026. https://www.ema.europa.eu/en/events/good-manufacturing-practice-multistakeholder-workshop-expert-contributions-artificial-intelligence-guidance-development-annex-22
- Legal Information Institute, Cornell Law School. “21 CFR § 211.68 Automatic, mechanical, and electronic equipment.” https://www.law.cornell.edu/cfr/text/21/211.68
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products” (draft guidance). January 2025. https://www.fda.gov/media/184830/download
- European Medicines Agency and U.S. Food and Drug Administration. “Guiding principles of good AI practice in drug development.” January 2026. https://www.ema.europa.eu/en/documents/other/guiding-principles-good-ai-practice-drug-development_en.pdf
- von Platen, Patrick. “How to generate text: using different decoding methods for language generation with Transformers.” Hugging Face, March 1, 2020 (edited July 2023). https://huggingface.co/blog/how-to-generate
- Anthropic. “Glossary” (entry: Temperature). Claude Platform documentation. https://platform.claude.com/docs/en/about-claude/glossary
- Microsoft. “How to generate reproducible output with Azure OpenAI in Microsoft Foundry Models (classic).” Microsoft Learn. https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/reproducible-output?view=foundry-classic
- Anadkat, Shyamal. “How to make your completions outputs consistent with the new seed parameter.” OpenAI Cookbook (archived recipe). https://cookbook.openai.com/examples/reproducible_outputs_with_the_seed_parameter
- Google Cloud. “Content generation parameters.” Vertex AI Generative AI documentation, last updated 28 September 2026. https://cloud.google.com/vertex-ai/generative-ai/docs/multimodal/content-generation-parameters
- He, Horace, and Thinking Machines Lab. “Defeating Nondeterminism in LLM Inference.” Thinking Machines Lab: Connectionism, September 10, 2025. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
- Atil, Berk, et al. “Non-Determinism of ‘Deterministic’ LLM Settings.” arXiv:2408.04667, submitted August 6, 2024, revised April 2, 2025. https://arxiv.org/abs/2408.04667
- PyTorch. “Reproducibility.” PyTorch documentation, developer notes. https://docs.pytorch.org/docs/stable/notes/randomness.html
- Anthropic. “Model deprecations.” Claude Platform documentation. https://platform.claude.com/docs/en/about-claude/model-deprecations
- Heitmann, Martin, Stefan Muench, Brandi Stockton, and Frederick Blumenthal. “ChatGPT, BARD, and Other Large Language Models Meet Regulated Pharma.” Pharmaceutical Engineering (ISPE), July/August 2023. https://ispe.org/pharmaceutical-engineering/july-august-2023/chatgpt-bard-and-other-large-language-models-meet








Your perspective matters—join the conversation.