In This Article
- Executive Summary
- What the Regulatory Documents Actually Say, and What Status They Hold
- Two Readers, Two Explanations: The Data Scientist and the QA Reviewer
- The Evidence Types and Where Each One Breaks
- Tying the Evidence to Intended Use and Risk Classification
- When the Model Belongs to a Vendor and the Internals Are Closed
- Keeping the Evidence Current After Retraining
- The Evidence Package: A Proposed Table of Contents and the Inspector’s Questions
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Pharma and biotech quality teams are being asked a new question in inspections: not only “does the model perform,” but “can you show me why it decided what it decided.” The regulatory texts that frame that question are, as of September 2026, mostly unfinished. Draft EU GMP Annex 22 is still a draft with no final text. FDA’s January 2025 guidance on AI credibility is still a draft, marked not for implementation. The January 2026 FDA and EMA joint principles are principles, not requirements. The ISPE GAMP Guide on Artificial Intelligence (July 2025) is industry good practice. None of these is a rule an inspector can cite as binding, and all of them describe, in overlapping ways, what a reasonable explainability record looks like.
The core insight of this article is that explainability evidence is a document set, not a chart. A SHAP plot satisfies a data scientist. A QA reviewer needs the plot, the rationale for why the highlighted features are process-relevant, the record that a subject matter expert reviewed and approved that rationale, the confidence threshold and what happens below it, and proof that all of it still describes the model in production today. The peer-reviewed literature on explanation methods is clear that these techniques are unstable under small input changes, can be manipulated, and are routinely over-trusted by the people who run them. An evidence package that ignores those limits will not survive a careful reviewer.
The article verifies the current status and actual wording of each regulatory document, then works through seven evidence types and their known failure modes, explains how the depth of evidence should follow intended use and risk, covers vendor models whose internals are closed, and describes how to keep the record current after retraining. It closes with a proposed table of contents for an explainability evidence package and a checklist of questions an inspector is likely to ask.
What the Regulatory Documents Actually Say, and What Status They Hold
Before anyone builds an evidence package, it is worth being precise about which documents exist, what each one actually says about explainability, and what legal weight it carries. Attribution errors are common in this area. Vendors and conference speakers routinely describe “Annex 22 requirements” or “FDA’s explainability rule” in terms the documents do not use. An inspector who has read the text will notice.
Draft EU GMP Annex 22 (still a draft)
The European Commission published draft Annex 22, titled “Artificial Intelligence,” for public consultation on July 7, 2025, alongside a revised Annex 11 and Chapter 4. The consultation closed on October 7, 2025. As of this writing, no final text has been adopted and no implementation date exists. Anything a company builds against Annex 22 today is built against a consultation draft that may change.1
What the draft says is nonetheless useful, because it is the only GMP text anywhere that names explanation methods directly. Its scope is narrow: computerized systems in manufacturing where AI models are used in “critical applications with direct impact on patient safety, product quality or data integrity, e.g. to predict or classify data.” It covers static models with deterministic output. It states that dynamic models that learn during use, probabilistic models, and generative AI and large language models “should not be used in critical GMP applications.”1 So the explainability sections of Annex 22 are written for classifiers and predictors, such as a visual inspection model that rejects vials or a model that predicts a critical quality attribute, not for a document-drafting assistant.
Section 8 of the draft is titled “Explainability” and contains two clauses. Clause 8.1, “Feature attribution,” states that during testing of models used in critical GMP applications, “systems should capture and record the features in the test data that have contributed to a particular classification or decision (e.g. rejection),” and that “where applicable, techniques like feature attribution (e.g. SHAP values or LIME) or visual tools like heat maps should be used to highlight key factors contributing to the outcome.” Clause 8.2, “Feature justification,” states that “in order to ensure that a model is making decisions based on relevant and appropriate features and based on risk, a review of these features should be part of the process for approval of test results.”1
Section 9, “Confidence,” adds two more. Clause 9.1 says the system should, “where applicable, log the confidence score of the model for each prediction or classification outcome.” Clause 9.2 says models “should have an appropriate threshold setting to ensure predictions or classifications are made only when suitable,” and that if the confidence score is very low, “it should be considered whether the model should flag the outcome as ‘undecided’, rather than making potentially unreliable predictions or classifications.”1
Read carefully, those four clauses describe a specific and fairly modest record: attribution captured during testing, a documented review of whether the attributed features make process sense as part of approving test results, per-output confidence logging, and a defended threshold with an “undecided” path. The draft does not say explanations must be produced for every production decision. It does not name a particular method as mandatory. It does not define what an acceptable explanation looks like. Those gaps are where a company’s own rationale has to do the work.
FDA’s January 2025 draft guidance on AI credibility (still a draft)
FDA’s “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products” was published in January 2025 under docket FDA-2024-D-4689. FDA’s guidance page still lists it as a Draft Level 1 Guidance, “Not for implementation. Contains non-binding recommendations.”2 It applies to AI used to produce information or data supporting regulatory decisions on safety, effectiveness, or quality, which includes manufacturing models.
The word “explainability” does not appear in this guidance at all. In its background section, FDA notes that understanding how AI models are developed and how they arrive at their conclusions “may be difficult and necessitate methodological transparency (e.g., detailing in the regulatory submission the methods and processes used to develop a particular AI model),” and that “uncertainty of the accuracy in the deployed models’ output may be difficult to interpret, explain, or quantify.”3
The operative content is the seven-step risk-based credibility assessment framework: define the question of interest, define the context of use, assess model risk, develop a credibility assessment plan, execute it, document results and deviations, and determine adequacy for the context of use. Model risk is defined as a combination of “model influence” (the contribution of the model’s evidence relative to other evidence) and “decision consequence” (the significance of an adverse outcome from an incorrect decision).3
Within Step 4, the guidance lists what a sponsor should describe about the model: inputs and outputs, architecture, features, the feature selection process and loss functions, parameters, and a rationale for the modeling approach. On evaluation, it asks sponsors to “specify the process by which the uncertainty and confidence level of model predictions were estimated,” notes that “information regarding the uncertainty of model output is important because it helps interpret model outputs,” and states that “all performance estimates should be provided with confidence intervals.” It also says that where the context of use involves a human in the loop, evaluation methods should “consider the performance of the human-AI team, rather than just the performance of the model in isolation.”3 There is no mention of SHAP, LIME, model cards, or any named explanation technique.
The FDA and EMA joint principles (January 2026)
On January 14, 2026, FDA and EMA published “Guiding Principles of Good AI Practice in Drug Development,” ten short principles covering AI used to generate or analyze evidence across nonclinical, clinical, post-marketing, and manufacturing phases. The document describes itself as intended “to lay the foundation for developing good practice” and says the areas of collaboration it describes “may help inform regulatory policies and regulatory guidelines in different jurisdictions.” It is not a guidance and creates no obligations.4
Two principles touch explainability directly. Principle 7, “Model design and development practices,” says development should follow best practices in model and system design and use fit-for-use data, “considering interpretability, explainability, and predictive performance,” and that good development “promotes transparency, reliability, generalizability, and robustness.” Principle 10, “Clear, essential information,” says plain language should be used to present “clear, accessible, and contextually relevant information to the intended audience, including users and patients, regarding the AI technology’s context of use, performance, limitations, underlying data, updates, and interpretability or explainability.”4 Principle 9 adds that AI technologies “undergo scheduled monitoring and periodic re-evaluation to ensure adequate performance (e.g., to address data drift).”4
The ISPE GAMP Guide: Artificial Intelligence (July 2025)
ISPE published the GAMP Guide: Artificial Intelligence in July 2025. It is a 290-page industry good practice guide that sits alongside GAMP 5 Second Edition and covers AI-enabled computerized systems across the life cycle, with a risk-based approach and attention to ongoing monitoring and control.5 It is not a regulation. In the authors’ own description of the guide in Pharmaceutical Engineering, “Explainable AI” is presented as a means “to support human-AI-team collaboration in GxP regulated processes,” and “Trustworthy AI” is framed through “transparency, human oversight, and bias mitigation.” The guide’s model governance strategies aim at “decision traceability throughout development and use,” and it emphasizes supplier collaboration and enhanced supplier assessment for AI-enabled systems.6 In a later summary, the same authors note that “general principles like the need for transparency, independence of test data sets, and the relevance of ongoing monitoring are becoming an agreed standard,” and say that the factors to consider when retiring a model include “traceability of model input, the model, and model output, as well as the integration into the AI-enabled computerized system to allow for ex-post assessment.”7
Context: the EU AI Act and the device-side transparency principles
Two further documents are worth knowing about, with caveats. Article 13 of the EU AI Act requires high-risk AI systems to be “sufficiently transparent to enable deployers to interpret a system’s output and use it appropriately,” and requires instructions for use to include the level of accuracy and its metrics, known limitations, and “where applicable, information to enable deployers to interpret the output.”8 Most pharma manufacturing AI is unlikely to be classified as high-risk under the Act, and the high-risk obligations have been deferred (Annex III systems to December 2027 and Annex I systems to August 2028), so this is context rather than a current obligation.
On the device side, FDA, Health Canada, and MHRA published “Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles” in June 2024. It defines transparency as “the degree to which appropriate information about a MLMD (including its intended use, development, performance and, when available, logic) is clearly communicated to relevant audiences,” and lists the “logic of the model, when available” and “device limitations, including biases, confidence intervals and data characterization gaps” among the information that should be shared.9 It is a device document and does not apply to pharma manufacturing, but the “when available” qualifier on model logic is the most candid regulatory statement anywhere that logic may simply not be available, and that other evidence must then carry the weight.
| Document | Status (September 2026) | What it says about explanation | What it does not say |
|---|---|---|---|
| Draft EU GMP Annex 221 | Consultation draft; consultation closed Oct 7, 2025; no final text, no implementation date | Capture attributed features during testing (SHAP, LIME, heat maps named as examples); review features as part of approving test results; log confidence scores; set a threshold and consider an “undecided” outcome | Does not require explanations for every production output; does not mandate a method; does not define an acceptable explanation |
| FDA draft AI credibility guidance2,3 | Draft, not for implementation, docket FDA-2024-D-4689 | Describe architecture, features, feature selection, parameters; state how uncertainty and confidence were estimated; performance estimates with confidence intervals; evaluate the human-AI team | Does not name any explanation technique; does not require model cards |
| FDA and EMA joint principles4 | Principles; not guidance; no obligations | Consider interpretability and explainability in design (Principle 7); communicate context of use, performance, limitations, and interpretability in plain language (Principle 10) | Does not specify evidence, format, or depth |
| ISPE GAMP Guide: AI5,6 | Industry good practice, July 2025 | Explainable AI to support human-AI teaming; transparency and human oversight; decision traceability; supplier assessment | Not a regulation; an inspector cannot cite it as a requirement |
| EU AI Act Article 138 | In force, but high-risk obligations deferred to 2027 and 2028 | Sufficient transparency for deployers to interpret output; instructions must cover accuracy, limitations, interpretation | Most GMP manufacturing models are unlikely to be high-risk under the Act |
Two Readers, Two Explanations: The Data Scientist and the QA Reviewer
Most explainability evidence fails not because the method was wrong but because it was written for the wrong reader. A feature attribution chart is a data scientist’s artifact. It answers a data scientist’s question: which inputs moved this output, and by how much. A QA reviewer or inspector is asking a different question: is this model making its decision for reasons that make sense in this process, and how do I know that is still true today.
What the literature says about the people who run the tools
Kaur and colleagues studied how data scientists actually use two interpretability tools, the InterpretML implementation of generalized additive models and the SHAP package, through a contextual inquiry with 11 participants and a survey of 197. They found that data scientists “over-trust and misuse interpretability tools,” and that few participants could accurately describe the visualizations the tools produced.10 That is a study of the experts, not the novices. If the people who generate SHAP plots misread them, a QA reviewer handed the same plot without an accompanying narrative has little chance.
Jesus and colleagues ran an application-grounded evaluation with professional fraud analysts, comparing decisions made with raw data only, data plus the model score, and data plus score plus explanations from LIME, SHAP, and TreeInterpreter. The analysts made the most accurate decisions with data only, and every explainer configuration underperformed that baseline, though explanations did help relative to seeing only the score. The authors warn that explainers chosen on proxy metrics like fidelity and stability “might be chosen that, in fact, hurt the overall performance of the combined system of ML model + end-users.”11 For a GxP process with a human in the loop, this is the direct warning: an explanation display can change operator behavior for the worse, and the human-AI team, not the model, is what gets evaluated under the FDA draft framework.3
Underneath both studies is a definitional problem that Lipton described in 2018: “interpretability” is not one property but a bundle of loosely related goals, and a claim that a model is interpretable is nearly meaningless unless someone says interpretable to whom and for what purpose.12 Doshi-Velez and Kim proposed distinguishing application-grounded evaluation (real people doing the real task), human-grounded evaluation (lay people on simplified tasks), and functionally-grounded evaluation (proxy metrics with no humans).13 Almost all explainability evidence produced in pharma today is functionally grounded. A reviewer can reasonably ask for at least one piece of evidence that a process expert used the explanation to reach a correct judgment.
What separates an actionable explanation from a technical one
Mechanism
A SHAP summary plot, a partial dependence curve, a saliency map. The artifact shows which inputs moved the output. It is internally consistent and reproducible from the notebook.
Meaning
The same artifact, plus a written statement from a process subject matter expert that the highlighted features are physically or procedurally relevant, why, and which features would be a red flag if they appeared.
Aggregate
Global feature importance across the test set. Average behavior, well summarized.
Instance
For the specific rejected batch or unit, what the model saw, what it attributed, what confidence it reported, and what the operator did next. Aggregate evidence cannot answer a question about a specific deviation.
Snapshot
Explanations generated at validation and filed with the validation report.
Lineage
Explanations tied to a model version, a test dataset version, an explanation-tool version, and a date, with a procedure that regenerates them on retraining and compares them to the last approved set.
The difference in every row is the same: the actionable version adds a human judgment, a record of that judgment, and a link to the model as it exists now. Rudin’s 2019 argument goes further: for high-stakes decisions, post hoc explanations of a black box model are themselves a source of risk, because they are approximations of the model rather than the model, and the better path is to build an inherently interpretable model where the problem permits.14 For many GMP classification problems on tabular data, that is a real option, and a reviewer may reasonably ask why a black box was chosen when an interpretable model would have met the acceptance criteria. The model selection rationale belongs in the evidence package for that reason.
The Evidence Types and Where Each One Breaks
Draft Annex 22 names SHAP, LIME, and heat maps as examples.1 The FDA draft asks for uncertainty and confidence estimation.3 Neither tells a quality team what the known weaknesses of these methods are, and an inspector who has read the literature may. What follows is each evidence type, what it actually shows, and the published limit that a package should acknowledge rather than hide.
Global feature importance
Global importance ranks input features by their average contribution across a dataset, whether from permutation importance, aggregated SHAP values, or model-native measures such as tree split gain. It answers “what does this model rely on in general.” It is the right first page of an evidence package because a process expert can read it without training: if a vial rejection model’s top feature is a lighting artifact rather than a defect signature, that is visible here. Its limit is that it is an average. A model can rely on sensible features on average and on a nonsensical one for a specific subgroup, which is exactly the case draft Annex 22 clause 3.2 asks companies to examine by dividing the input space into subgroups.1 Global importance should therefore be reported per subgroup, not only overall.
Local attribution: SHAP and LIME
SHAP, introduced by Lundberg and Lee in 2017, assigns each feature a contribution to a single prediction using Shapley values from cooperative game theory, and unifies several earlier attribution methods under one framework.15 LIME, introduced by Ribeiro, Singh, and Guestrin in 2016, fits a simple interpretable model in the neighborhood of a single prediction by perturbing the input and observing the output.16 Both are model-agnostic, both are widely implemented, and both are named in the Annex 22 draft. They are also the methods with the best-documented weaknesses.
Alvarez-Melis and Jaakkola argued in 2018 that a basic requirement of any explanation is that “similar inputs should give rise to similar explanations,” proposed metrics for that property, and showed that current methods “do not perform well according to these metrics.”17 In practical terms, two nearly identical images of the same vial can yield visibly different attribution maps. Slack and colleagues showed that a classifier can be deliberately constructed so that “post hoc explanations of the scaffolded classifier look innocuous” while the underlying predictions remain biased, and that such classifiers “can easily fool popular explanation techniques such as LIME and SHAP.”18 The relevance in GxP is not that a vendor is adversarial; it is that a method which can be fooled by construction is also a method that can be misled by ordinary correlation in training data. Kumar and colleagues added that “mathematical problems arise when Shapley values are used for feature importance,” that fixing them requires causal assumptions, and that the resulting explanations may not match what people actually need from an explanation.19
None of this means SHAP and LIME should be excluded. It means the evidence package should record the method version and settings (background dataset, number of perturbation samples, kernel width), report attribution stability on repeated runs and on slightly perturbed inputs, and state in plain language that these are approximations of the model’s behavior rather than a readout of its internals.
Saliency and heat maps for image models
For visual inspection models, heat maps are the natural explanation and the Annex 22 draft names them.1 Two papers should temper confidence in them. Adebayo and colleagues proposed “sanity checks” for saliency methods and found that several popular ones produce nearly the same map whether the model is trained or has randomized weights, meaning the map reflects the image’s edges rather than what the model learned.20 Ghorbani, Abid, and Zou showed that “interpretation of neural networks is fragile”: imperceptible changes to an input can produce a very different saliency map while leaving the prediction unchanged.21 A package that includes heat maps should state which saliency method was used, whether it passed the randomization sanity check, and whether maps for near-duplicate images agree. If nobody has run those checks, that is a finding waiting to happen.
Decision boundary and threshold documentation
For a classifier, the single most consequential design decision is where the threshold sits, and it is the piece of explainability evidence quality teams most often leave out because it does not look like an “explanation.” Annex 22 draft clause 9.2 asks for “an appropriate threshold setting” and for consideration of an “undecided” outcome at low confidence.1 The record should show the threshold value, the operating characteristic it was chosen on (sensitivity versus specificity trade-off, with the confusion matrix at that point), who approved it, the rationale linking it to the process (for example, that false accepts are the failure that matters and false rejects go to a manual re-inspection), the “undecided” band and what procedure handles it, and per-subgroup performance at the chosen threshold. A reviewer can act on that. A reviewer cannot act on “the model uses a 0.5 threshold.”
Worked example libraries
A worked example library is a curated set of representative inputs and the model’s outputs, each with attribution, confidence, and a process expert’s annotation of whether the decision and its reasons are correct. It should span the subgroups defined for intended use, include the rare variations Annex 22 draft clause 3.1 asks companies to characterize, and include cases at and near the threshold.1 This is the evidence type closest to how an inspector already thinks: show me a case, walk me through it. It is also the evidence type that best supports the Annex 22 draft’s “feature justification” review, because the justification is done case by case by a named expert and signed. Its limit is coverage. A library of forty cases proves the model behaved sensibly on forty cases; it should be presented as a qualitative complement to the quantitative test, never a substitute.
Counterfactual examples
Wachter, Mittelstadt, and Russell proposed counterfactual explanations in 2017 as a way to explain an automated decision without opening the model: state the smallest change to the input that would have changed the output.22 For a GMP model this translates to statements such as “this unit would have been accepted if the measured fill volume had been 0.3 mL higher” or “this batch would not have been flagged if the temperature excursion had been 4 minutes shorter.” Counterfactuals are unusually useful for QA review because they are testable: a process expert can say whether the stated change is physically meaningful. Their limit is that many counterfactuals exist for any decision, some of them physically impossible, so the package should state how the counterfactuals were generated and constrained to realistic input ranges.
Uncertainty and confidence reporting
The Annex 22 draft asks for confidence scores to be logged per outcome.1 The FDA draft asks sponsors to specify how uncertainty and confidence were estimated and to report performance with confidence intervals.3 The trap is that a softmax output is not a calibrated probability. Guo and colleagues showed in 2017 that “modern neural networks, unlike those from a decade ago, are poorly calibrated,” and that a simple post-processing step, temperature scaling, “is surprisingly effective at calibrating predictions.”23 A confidence threshold set on an uncalibrated score is a threshold on a number that does not mean what the procedure says it means. The evidence should include a calibration plot (reliability diagram) and the calibration method, or an explicit statement that the score is a ranking rather than a probability and that the threshold was set empirically on test data.
For teams that want a confidence statement with a guarantee, conformal prediction offers one. Angelopoulos and Bates describe it as a distribution-free method that wraps any model and produces prediction sets with a user-chosen coverage level, under an exchangeability assumption.24 In a classification setting, a conformal prediction set that contains both “accept” and “reject” is a principled version of the Annex 22 draft’s “undecided” flag, with a stated error rate rather than an arbitrary cutoff. That assumption of exchangeability is also why it must be re-checked after any change in input distribution.
Model cards and equivalent summaries
Mitchell and colleagues proposed model cards in 2019 as short documents accompanying trained models that report intended use, out-of-scope uses, evaluation data, performance disaggregated by relevant conditions and groups, and ethical considerations.25 In a GxP setting, the model card is the cover sheet of the evidence package: one or two pages a reviewer can read in five minutes that state what the model is for, what it is not for, how it performed on which data, and where the detailed evidence lives. It should not carry any claim not supported by a document behind it. None of the regulatory documents reviewed above requires a model card by that name, and the package should not claim otherwise; it should present the card as the company’s chosen way of meeting the joint principles’ call for “clear, accessible, and contextually relevant information.”4
| Evidence type | What it shows | Published limit | What the package must record |
|---|---|---|---|
| Global feature importance | What the model relies on across the test set | Averages hide subgroup behavior | Importance per subgroup; SME statement on process relevance |
| Local attribution (SHAP, LIME) | What drove one specific output | Unstable under small perturbations17; can be fooled18; mathematical issues19 | Method and version, settings, stability results, plain-language caveat |
| Saliency / heat maps | Image regions associated with the output | Some methods ignore the model20; fragile to tiny input changes21 | Method, sanity-check result, agreement across near-duplicate images |
| Threshold documentation | Where and why the decision line sits | Meaningless if score is uncalibrated23 | Threshold, operating point, approver, undecided band, subgroup performance |
| Worked example library | Case-by-case reasonableness, expert-annotated | Coverage is limited to the cases chosen | Case selection rationale, subgroup coverage, signed SME annotations |
| Counterfactuals | Smallest input change that flips the output22 | Many exist; some are physically impossible | Generation method, realism constraints, SME review |
| Uncertainty and confidence | How sure the model is, per output | Raw scores are not probabilities23 | Calibration plot and method, or conformal coverage and its assumption24 |
| Model card | One-page summary of use, data, performance, limits25 | Only as good as the documents behind it | Cross-references to every supporting record; version and date |
Tying the Evidence to Intended Use and Risk Classification
Every document reviewed above makes the same structural move: define intended use first, then assess risk, then scale the evidence to the risk. Explainability evidence is no exception. The depth of the package should be an output of the risk assessment, not a fixed template applied to every model.
Intended use defines what an explanation is for
Draft Annex 22 clause 3.1 asks for the intended use to be described “in detail,” including “a comprehensive characterisation of the data the model is intended to use as input and all common and rare variations,” with a process subject matter expert responsible for the adequacy of that description. Clause 3.3 adds that where a model gives input to a human decision and the testing effort has been reduced on that basis, the description “should include the responsibility of the operator.”1 The FDA draft’s Step 2 asks for the context of use, and Step 1 for the question of interest the model helps answer.3
That description determines the explanation audience. If the model is fully automated with a downstream release test as the safeguard (the FDA draft’s own fill-volume example, where release testing reduces model influence and lowers model risk to medium despite high decision consequence3), the explanation audience is the QA reviewer doing periodic review and the investigator handling a deviation. If the model presents a recommendation to an operator who makes the call, the explanation audience is the operator in real time, and the evidence must include what the operator sees, how they were trained to read it, and how the human-AI team performed in testing, which is what the FDA draft explicitly asks for.3 The Jesus study is the reason this matters: an explanation display can lower the accuracy of the combined system.11
Risk sets the depth
Using the FDA draft’s two axes, model influence and decision consequence, a practical scaling looks like this:
Low model risk (low influence, low consequence)
Model card, global feature importance with an SME relevance statement, threshold documentation, confidence logging. Worked examples optional. This is the floor, and it already exceeds what many deployed models have today.
Medium model risk (one axis high, mitigated by independent checks)
Everything above, plus per-subgroup importance, a worked example library covering all subgroups and near-threshold cases, calibration evidence, and documented attribution stability for the chosen local method.
High model risk (high influence, high consequence)
Everything above, plus counterfactual analysis, an application-grounded evaluation showing process experts reach correct judgments using the explanations, a documented rationale for why an interpretable model was not used or was rejected, and per-output attribution retained in production for the retention period, not only at test.
Any risk level, human in the loop
Add the operator-facing display specification, the training record for reading it, the human-AI team test results, and the Annex 22 draft clause 10.5 records of the human review process, which may mean a review of every output depending on criticality.
The risk classification itself has to be in the package, with the reasoning. An inspector who sees a thin explainability record for a model that rejects product will want to know whether the thinness was a decision or an oversight. A one-page risk rationale that cites model influence, decision consequence, and the mitigating controls turns a potential observation into a documented judgment call.
When the Model Belongs to a Vendor and the Internals Are Closed
Much of the AI in GMP manufacturing arrives inside a purchased system: a vision inspection platform, a predictive maintenance module, a chromatography data system feature. The regulated company often cannot see the architecture, the training data, or the weights. Draft Annex 22 clause 2.2 does not accept that as a reason for a thinner record. It states that documentation “should be available and reviewed by the regulated user irrespective of whether a model is trained, validated and tested in-house or whether it is provided by a supplier or service provider.”1 The GAMP AI Guide likewise emphasizes enhanced supplier assessment and collaboration for AI-enabled systems.6
What to ask the supplier for
The request list should mirror the FDA draft’s Step 4 description items, because they are the clearest published statement of what a reviewer expects to see about a model: inputs and outputs, architecture class, features, feature selection, parameters (or at least a statement of their count and governance), training data description and how labels were established, the evaluation approach, performance with confidence intervals, and how uncertainty was estimated.3 Suppliers can provide most of that without disclosing weights or proprietary training sets. Where they refuse, the refusal should be documented, and the gap closed with evidence the regulated company generates itself.
Evidence the regulated company can generate without the internals
Every explanation type in the table above except native model-specific importance can be produced against a closed model, because SHAP’s kernel variant, LIME, counterfactual search, and calibration all work on inputs and outputs alone. That is their entire reason for existing. So a closed vendor model does not prevent an explainability record; it changes who produces it and shifts the emphasis toward behavioral evidence:
- Independent test data and independent metrics. The Annex 22 draft’s sections 5 through 7 on test data, independence, and execution apply in full, and the regulated company owns them regardless of where the model came from.1
- Behavioral attribution. Model-agnostic attribution on the company’s own test set, with the stability checks described earlier, so that the “feature justification” review can still be performed by a process expert.
- A worked example library built from the company’s own product and lines, since the supplier’s examples were not drawn from the intended use as this site defines it.
- Counterfactual probing at the decision boundary, which is often the fastest way to discover that a closed model is sensitive to something it should not be.
- Confidence logging and calibration on site data, with the threshold set and approved by the regulated company, not accepted as a vendor default.
- Contractual commitments on change notification, so that a silent model update in a software release does not invalidate the record. The Annex 22 draft’s change control and configuration control clauses (10.1 and 10.2) require the company to detect unauthorized change; it cannot do that if the supplier does not tell it what changed.1
The device-side transparency principles are worth borrowing here for their honesty: model logic is to be communicated “when available.”9 A package for a closed vendor model should say plainly that the logic is not available, explain what was requested and what was received, and show that the behavioral evidence covers the intended use. That is a defensible position. Pretending the vendor’s marketing white paper is an explainability record is not.
Keeping the Evidence Current After Retraining
An explainability record is a statement about a specific model version on a specific test set with a specific explanation tool. Any of the three can change. The regulatory documents agree on the response. Draft Annex 22 clause 10.1 requires a tested model, its system, and the whole process to be under change control before deployment, and says any change to the model, the system, the process, or the physical objects it takes as input “should be documented and evaluated to determine if the model needs to be retested,” with any decision not to retest “fully justified.”1 The FDA draft says that depending on the extent and impact of a change, “some steps in the credibility assessment plan may need to be re-executed, including retraining and retesting the model,” and asks for life cycle plans that specify performance metrics, a risk-based monitoring frequency, and “triggers for model retesting.”3 The joint principles call for “scheduled monitoring and periodic re-evaluation.”4
Retraining as a change control event with an explanation clause
Most organizations already treat retraining as a change. Fewer treat the explanation record as part of what the change affects. A practical procedure adds four things to the retraining change record:
- Regenerate every explanation artifact against the new model version using the same test set version and the same explanation tool version, and file them as a new revision, never overwriting the old.
- Compare the new global importance and subgroup importance with the last approved set, and require a process expert to review and sign off on any material change in the ranking. A model that now depends on a different feature to reach the same accuracy has changed in a way the accuracy metric cannot see.
- Re-run the worked example library and flag every case where the decision, the attribution, or the confidence changed. Cases that flipped are the first things a reviewer will ask about.
- Re-check calibration and, if conformal methods are used, re-check coverage, since both depend on the relationship between model and data that retraining alters.23,24
Explanation drift as a monitoring signal
Annex 22 draft clauses 10.3 and 10.4 require monitoring of model performance metrics and of whether input data remain within the sample space, with defined drift metrics.1 Attribution offers a third monitoring channel that neither of those captures: the distribution of attributed features on production inputs. If a vision model that was approved because it attends to defect regions begins attending to background regions on a new lot of containers, performance metrics may not move for weeks, and the input distribution may look within range on the monitored variables. A periodic sample of production attributions compared against the approved reference set will catch it. This is not required by any of the documents reviewed here; it is a control that follows naturally from having built the record in the first place, and it converts an explainability package from a validation artifact into an operational one.
The Evidence Package: A Proposed Table of Contents and the Inspector’s Questions
The following is a proposed structure for an explainability evidence package for a validated AI or machine learning model used in a GxP decision. It is intended to sit within, or be cross-referenced from, the system’s validation package under Annex 11 or the site’s computer software assurance approach, and to be scaled by risk as described above. Section numbers are suggestions.
Proposed table of contents
| Section | Contents | Owner |
|---|---|---|
| 0. Model card | One to two pages: intended use and out-of-scope uses; model version; data summary; headline performance with confidence intervals per subgroup; known limitations; confidence threshold; where each supporting record lives25 | Data science, approved by QA |
| 1. Regulatory basis and status | Which documents were used as benchmarks and their status as of the package date; the binding requirements the package satisfies (Annex 11, site procedures); explicit statement that draft documents are drafts | QA / regulatory |
| 2. Intended use and audience for explanation | Intended use description per Annex 22 draft 3.1; subgroups per 3.2; human-in-the-loop responsibilities per 3.3; who needs explanations, when, and for what decision1 | Process SME |
| 3. Risk classification and evidence depth rationale | Model influence and decision consequence assessment; mitigating controls; the resulting evidence tier and why3 | QA with process SME |
| 4. Model description and selection rationale | Inputs, outputs, architecture class, features, feature selection, parameters, training summary; why this model type; why an interpretable alternative was or was not chosen3,14 | Data science |
| 5. Global feature importance and SME relevance review | Importance overall and per subgroup; method and version; signed SME statement on process relevance and on features that would be red flags | Data science, process SME |
| 6. Local attribution and stability | Method, settings, version; attribution stability on repeated runs and perturbed inputs; plain-language caveat on approximation17,18,19 | Data science |
| 7. Saliency evidence (image models) | Method; randomization sanity check result; near-duplicate agreement20,21 | Data science |
| 8. Decision threshold and undecided handling | Threshold, operating point, confusion matrix at that point, per-subgroup metrics, approver, undecided band and procedure1 | Process SME, QA |
| 9. Confidence and uncertainty | How confidence is computed; calibration method and reliability diagram, or conformal coverage and assumptions; confidence logging specification23,24 | Data science |
| 10. Worked example library | Case selection rationale; coverage matrix by subgroup and near-threshold; each case with input, output, attribution, confidence, and signed SME annotation | Process SME |
| 11. Counterfactual analysis (higher risk) | Generation method; realism constraints; SME review of physical plausibility22 | Data science, process SME |
| 12. Human-AI team evidence (human in the loop) | Operator display specification; training record; team performance in testing versus model alone; review records per Annex 22 draft 10.51,3 | Operations, QA |
| 13. Supplier evidence and gaps (vendor models) | What was requested, what was received, what was refused, and the behavioral evidence that closes each gap; change notification terms | QA, procurement |
| 14. Currency and change history | Version pins for model, data, and tools; regeneration procedure on retraining; comparison against last approved set; production attribution monitoring plan | Data science, QA |
| 15. Approvals | Process SME, data science lead, QA; date; link to validation summary report | All |
Questions an inspector is likely to ask
These are framed the way they tend to be asked: plainly, in sequence, each one following from the last. The package section that should answer each is noted.
Inspector question checklist
- What does this model decide, and what happens downstream if it is wrong? (Sections 2 and 3)
- Who decided how much explanation evidence was enough, and on what basis? (Section 3)
- Why this type of model? Could a simpler one you can read directly have met the acceptance criteria? (Section 4)
- What does the model rely on, and has someone who knows the process confirmed those features make sense? (Section 5)
- Does it rely on the same things for every line, camera, material, or site? (Sections 5 and 10)
- Show me a rejected unit or batch. What did the model see, what did it attribute, how confident was it, and what did the operator do? (Sections 10 and 12)
- If I ran your attribution tool twice, would I get the same answer? What about on a nearly identical input? (Sections 6 and 7)
- Where is the threshold, who set it, and what happens when the model is not sure? (Section 8)
- Is the confidence score a probability? How do you know? (Section 9)
- What is the smallest change to this input that would have changed the decision, and is that change physically meaningful? (Section 11)
- What does the operator see, how were they trained to read it, and did you test the operator and model together or only the model? (Section 12)
- This model came from a supplier. What did you ask for, what did they give you, and how did you fill the gaps? (Section 13)
- The model was retrained in March. Which of these documents describe the model that is running today? (Section 14)
- How would you know if the model started relying on something different without its accuracy changing? (Section 14)
- You cite Annex 22. Is that in force? (Section 1)
The last question is not a trick. It is a check on whether the people presenting the package understand what they are presenting. The correct answer, given plainly, builds more credibility than any chart in the binder.
Conclusion
The regulatory texts on explainability are converging on a shared picture even while most of them remain unfinished: know the intended use, capture what the model relies on, have a process expert judge whether that is sensible, log and threshold confidence honestly, evaluate the human and model together where there is a human, and keep all of it current through change control and monitoring. What the texts do not supply is a warning about the tools. The peer-reviewed literature supplies that in abundance: attribution methods are unstable, can be misled, are misread by the experts who run them, and can make human decisions worse when displayed carelessly. An evidence package that takes both sources seriously looks different from one built to satisfy a checklist. It is organized around questions a reviewer would ask, it records judgments and who made them, it states the limits of its own methods, and it can be regenerated and compared after the model changes.
Sakara Digital works with pharma and biotech organizations building validation and governance records for AI and machine learning models in GxP processes, including the explainability evidence that has to survive a careful reviewer. If you are preparing a model for inspection, assessing a vendor system whose internals you cannot see, or deciding how much explanation evidence a given model actually needs, we are happy to have that conversation.
For Further Reading
For Further Reading
- Annex 22 Mock Inspection: What a Pharma Quality Team Should Practice Now
- Human-in-the-Loop Requirements for Pharma AI: What FDA and EMA Actually Expect
- The AI Model Risk Assessment for Pharma: A Structured Checklist
- Building an AI Model Registry: What to Track and Why
- Computer System Validation vs. AI Validation: Key Differences That Matter
References & Sources
- European Commission. “Annex 22: Artificial Intelligence.” EudraLex Volume 4, consultation draft, July 2025 (consultation closed October 7, 2025; no final text adopted). https://health.ec.europa.eu/document/download/5f38a92d-bb8e-4264-8898-ea076e926db6_en?filename=mp_vol4_chap4_annex22_consultation_guideline_en.pdf
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products.” Guidance document page, Draft Level 1 Guidance, docket FDA-2024-D-4689, January 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products: Draft Guidance for Industry and Other Interested Parties.” January 2025 (PDF). https://www.fda.gov/media/184830/download
- U.S. Food and Drug Administration and European Medicines Agency. “Guiding Principles of Good AI Practice in Drug Development.” January 2026 (PDF). https://www.fda.gov/media/189581/download
- International Society for Pharmaceutical Engineering. “ISPE GAMP Guide: Artificial Intelligence.” July 2025. https://ispe.org/publications/guidance-documents/gamp-guide-artificial-intelligence
- Stockton, B., Staib, E., and Heitmann, M. “New GAMP Guide Addresses Challenges Posed by AI-Enabled Computerized Systems.” Pharmaceutical Engineering, September/October 2025. https://ispe.org/pharmaceutical-engineering/september-october-2025/new-gampr-guide-addresses-challenges-posed-ai
- Stockton, B., Heitmann, M., and O’Connell, J. “New ISPE Framework Targets Uncertainty In Pharma’s AI Deployment.” BioProcess Online, January 29, 2026. https://www.bioprocessonline.com/doc/new-ispe-framework-targets-uncertainty-in-pharma-s-ai-deployment-0001
- European Union. “Article 13: Transparency and Provision of Information to Deployers.” Regulation (EU) 2024/1689 (EU AI Act), as presented by the EU Artificial Intelligence Act explorer. https://artificialintelligenceact.eu/article/13/
- U.S. Food and Drug Administration, Health Canada, and MHRA. “Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles.” June 2024. https://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles
- Kaur, H., Nori, H., Jenkins, S., Caruana, R., Wallach, H., and Wortman Vaughan, J. “Interpreting Interpretability: Understanding Data Scientists’ Use of Interpretability Tools for Machine Learning.” Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 2020. https://www.microsoft.com/en-us/research/publication/interpreting-interpretability-understanding-data-scientists-use-of-interpretability-tools-for-machine-learning/
- Jesus, S., Belém, C., Balayan, V., Bento, J., Saleiro, P., Bizarro, P., and Gama, J. “How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations.” Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21), 2021. https://arxiv.org/abs/2101.08758
- Lipton, Z. C. “The Mythos of Model Interpretability.” ACM Queue 16(3), 2018 (arXiv preprint 1606.03490). https://arxiv.org/abs/1606.03490
- Doshi-Velez, F., and Kim, B. “Towards A Rigorous Science of Interpretable Machine Learning.” arXiv preprint 1702.08608, 2017. https://arxiv.org/abs/1702.08608
- Rudin, C. “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.” Nature Machine Intelligence 1, 206–215, 2019. https://www.nature.com/articles/s42256-019-0048-x
- Lundberg, S. M., and Lee, S.-I. “A Unified Approach to Interpreting Model Predictions.” Advances in Neural Information Processing Systems 30 (NeurIPS 2017). https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html
- Ribeiro, M. T., Singh, S., and Guestrin, C. “‘Why Should I Trust You?’: Explaining the Predictions of Any Classifier.” Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016 (arXiv 1602.04938). https://arxiv.org/abs/1602.04938
- Alvarez-Melis, D., and Jaakkola, T. S. “On the Robustness of Interpretability Methods.” arXiv preprint 1806.08049, 2018 (ICML 2018 Workshop on Human Interpretability in Machine Learning). https://arxiv.org/abs/1806.08049
- Slack, D., Hilgard, S., Jia, E., Singh, S., and Lakkaraju, H. “Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods.” Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2020) (arXiv 1911.02508). https://arxiv.org/abs/1911.02508
- Kumar, I. E., Venkatasubramanian, S., Scheidegger, C., and Friedler, S. “Problems with Shapley-value-based explanations as feature importance measures.” Proceedings of the 37th International Conference on Machine Learning, PMLR 119, 2020. https://proceedings.mlr.press/v119/kumar20e.html
- Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B. “Sanity Checks for Saliency Maps.” Advances in Neural Information Processing Systems 31 (NeurIPS 2018). https://proceedings.neurips.cc/paper/2018/hash/294a8ed24b1ad22ec2e7efea049b8737-Abstract.html
- Ghorbani, A., Abid, A., and Zou, J. “Interpretation of Neural Networks Is Fragile.” Proceedings of the AAAI Conference on Artificial Intelligence 33(1), 2019. https://ojs.aaai.org/index.php/AAAI/article/view/4252
- Wachter, S., Mittelstadt, B., and Russell, C. “Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR.” Harvard Journal of Law & Technology 31(2), 2018 (arXiv 1711.00399). https://arxiv.org/abs/1711.00399
- Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. “On Calibration of Modern Neural Networks.” Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017. https://proceedings.mlr.press/v70/guo17a.html
- Angelopoulos, A. N., and Bates, S. “A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification.” arXiv preprint 2107.07511, 2021. https://arxiv.org/abs/2107.07511
- Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T. “Model Cards for Model Reporting.” Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019) (arXiv 1810.03993). https://arxiv.org/abs/1810.03993








Your perspective matters—join the conversation.