What the Regulatory Documents Actually Say, and What Status They Hold

Before anyone builds an evidence package, it is worth being precise about which documents exist, what each one actually says about explainability, and what legal weight it carries. Attribution errors are common in this area. Vendors and conference speakers routinely describe “Annex 22 requirements” or “FDA’s explainability rule” in terms the documents do not use. An inspector who has read the text will notice.

Draft EU GMP Annex 22 (still a draft)

The European Commission published draft Annex 22, titled “Artificial Intelligence,” for public consultation on July 7, 2025, alongside a revised Annex 11 and Chapter 4. The consultation closed on October 7, 2025. As of this writing, no final text has been adopted and no implementation date exists. Anything a company builds against Annex 22 today is built against a consultation draft that may change.1

What the draft says is nonetheless useful, because it is the only GMP text anywhere that names explanation methods directly. Its scope is narrow: computerized systems in manufacturing where AI models are used in “critical applications with direct impact on patient safety, product quality or data integrity, e.g. to predict or classify data.” It covers static models with deterministic output. It states that dynamic models that learn during use, probabilistic models, and generative AI and large language models “should not be used in critical GMP applications.”1 So the explainability sections of Annex 22 are written for classifiers and predictors, such as a visual inspection model that rejects vials or a model that predicts a critical quality attribute, not for a document-drafting assistant.

Section 8 of the draft is titled “Explainability” and contains two clauses. Clause 8.1, “Feature attribution,” states that during testing of models used in critical GMP applications, “systems should capture and record the features in the test data that have contributed to a particular classification or decision (e.g. rejection),” and that “where applicable, techniques like feature attribution (e.g. SHAP values or LIME) or visual tools like heat maps should be used to highlight key factors contributing to the outcome.” Clause 8.2, “Feature justification,” states that “in order to ensure that a model is making decisions based on relevant and appropriate features and based on risk, a review of these features should be part of the process for approval of test results.”1

Section 9, “Confidence,” adds two more. Clause 9.1 says the system should, “where applicable, log the confidence score of the model for each prediction or classification outcome.” Clause 9.2 says models “should have an appropriate threshold setting to ensure predictions or classifications are made only when suitable,” and that if the confidence score is very low, “it should be considered whether the model should flag the outcome as ‘undecided’, rather than making potentially unreliable predictions or classifications.”1

Read carefully, those four clauses describe a specific and fairly modest record: attribution captured during testing, a documented review of whether the attributed features make process sense as part of approving test results, per-output confidence logging, and a defended threshold with an “undecided” path. The draft does not say explanations must be produced for every production decision. It does not name a particular method as mandatory. It does not define what an acceptable explanation looks like. Those gaps are where a company’s own rationale has to do the work.

FDA’s January 2025 draft guidance on AI credibility (still a draft)

FDA’s “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products” was published in January 2025 under docket FDA-2024-D-4689. FDA’s guidance page still lists it as a Draft Level 1 Guidance, “Not for implementation. Contains non-binding recommendations.”2 It applies to AI used to produce information or data supporting regulatory decisions on safety, effectiveness, or quality, which includes manufacturing models.

The word “explainability” does not appear in this guidance at all. In its background section, FDA notes that understanding how AI models are developed and how they arrive at their conclusions “may be difficult and necessitate methodological transparency (e.g., detailing in the regulatory submission the methods and processes used to develop a particular AI model),” and that “uncertainty of the accuracy in the deployed models’ output may be difficult to interpret, explain, or quantify.”3

The operative content is the seven-step risk-based credibility assessment framework: define the question of interest, define the context of use, assess model risk, develop a credibility assessment plan, execute it, document results and deviations, and determine adequacy for the context of use. Model risk is defined as a combination of “model influence” (the contribution of the model’s evidence relative to other evidence) and “decision consequence” (the significance of an adverse outcome from an incorrect decision).3

Within Step 4, the guidance lists what a sponsor should describe about the model: inputs and outputs, architecture, features, the feature selection process and loss functions, parameters, and a rationale for the modeling approach. On evaluation, it asks sponsors to “specify the process by which the uncertainty and confidence level of model predictions were estimated,” notes that “information regarding the uncertainty of model output is important because it helps interpret model outputs,” and states that “all performance estimates should be provided with confidence intervals.” It also says that where the context of use involves a human in the loop, evaluation methods should “consider the performance of the human-AI team, rather than just the performance of the model in isolation.”3 There is no mention of SHAP, LIME, model cards, or any named explanation technique.

The FDA and EMA joint principles (January 2026)

On January 14, 2026, FDA and EMA published “Guiding Principles of Good AI Practice in Drug Development,” ten short principles covering AI used to generate or analyze evidence across nonclinical, clinical, post-marketing, and manufacturing phases. The document describes itself as intended “to lay the foundation for developing good practice” and says the areas of collaboration it describes “may help inform regulatory policies and regulatory guidelines in different jurisdictions.” It is not a guidance and creates no obligations.4

Two principles touch explainability directly. Principle 7, “Model design and development practices,” says development should follow best practices in model and system design and use fit-for-use data, “considering interpretability, explainability, and predictive performance,” and that good development “promotes transparency, reliability, generalizability, and robustness.” Principle 10, “Clear, essential information,” says plain language should be used to present “clear, accessible, and contextually relevant information to the intended audience, including users and patients, regarding the AI technology’s context of use, performance, limitations, underlying data, updates, and interpretability or explainability.”4 Principle 9 adds that AI technologies “undergo scheduled monitoring and periodic re-evaluation to ensure adequate performance (e.g., to address data drift).”4

The ISPE GAMP Guide: Artificial Intelligence (July 2025)

ISPE published the GAMP Guide: Artificial Intelligence in July 2025. It is a 290-page industry good practice guide that sits alongside GAMP 5 Second Edition and covers AI-enabled computerized systems across the life cycle, with a risk-based approach and attention to ongoing monitoring and control.5 It is not a regulation. In the authors’ own description of the guide in Pharmaceutical Engineering, “Explainable AI” is presented as a means “to support human-AI-team collaboration in GxP regulated processes,” and “Trustworthy AI” is framed through “transparency, human oversight, and bias mitigation.” The guide’s model governance strategies aim at “decision traceability throughout development and use,” and it emphasizes supplier collaboration and enhanced supplier assessment for AI-enabled systems.6 In a later summary, the same authors note that “general principles like the need for transparency, independence of test data sets, and the relevance of ongoing monitoring are becoming an agreed standard,” and say that the factors to consider when retiring a model include “traceability of model input, the model, and model output, as well as the integration into the AI-enabled computerized system to allow for ex-post assessment.”7

Context: the EU AI Act and the device-side transparency principles

Two further documents are worth knowing about, with caveats. Article 13 of the EU AI Act requires high-risk AI systems to be “sufficiently transparent to enable deployers to interpret a system’s output and use it appropriately,” and requires instructions for use to include the level of accuracy and its metrics, known limitations, and “where applicable, information to enable deployers to interpret the output.”8 Most pharma manufacturing AI is unlikely to be classified as high-risk under the Act, and the high-risk obligations have been deferred (Annex III systems to December 2027 and Annex I systems to August 2028), so this is context rather than a current obligation.

On the device side, FDA, Health Canada, and MHRA published “Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles” in June 2024. It defines transparency as “the degree to which appropriate information about a MLMD (including its intended use, development, performance and, when available, logic) is clearly communicated to relevant audiences,” and lists the “logic of the model, when available” and “device limitations, including biases, confidence intervals and data characterization gaps” among the information that should be shared.9 It is a device document and does not apply to pharma manufacturing, but the “when available” qualifier on model logic is the most candid regulatory statement anywhere that logic may simply not be available, and that other evidence must then carry the weight.

DocumentStatus (September 2026)What it says about explanationWhat it does not say
Draft EU GMP Annex 221Consultation draft; consultation closed Oct 7, 2025; no final text, no implementation dateCapture attributed features during testing (SHAP, LIME, heat maps named as examples); review features as part of approving test results; log confidence scores; set a threshold and consider an “undecided” outcomeDoes not require explanations for every production output; does not mandate a method; does not define an acceptable explanation
FDA draft AI credibility guidance2,3Draft, not for implementation, docket FDA-2024-D-4689Describe architecture, features, feature selection, parameters; state how uncertainty and confidence were estimated; performance estimates with confidence intervals; evaluate the human-AI teamDoes not name any explanation technique; does not require model cards
FDA and EMA joint principles4Principles; not guidance; no obligationsConsider interpretability and explainability in design (Principle 7); communicate context of use, performance, limitations, and interpretability in plain language (Principle 10)Does not specify evidence, format, or depth
ISPE GAMP Guide: AI5,6Industry good practice, July 2025Explainable AI to support human-AI teaming; transparency and human oversight; decision traceability; supplier assessmentNot a regulation; an inspector cannot cite it as a requirement
EU AI Act Article 138In force, but high-risk obligations deferred to 2027 and 2028Sufficient transparency for deployers to interpret output; instructions must cover accuracy, limitations, interpretationMost GMP manufacturing models are unlikely to be high-risk under the Act
The status point matters in the room. If an inspector asks “how do you meet Annex 22 section 8,” the accurate answer is that Annex 22 is a consultation draft, that the company has used it as a benchmark for its own procedure, and that the binding requirements come from Annex 11 (validation, and the requirement that critical systems be fit for intended use) and from the site’s own quality system. The evidence package should say this in its introduction. A package that cites a draft as a regulation invites the question of what else in it is imprecise.

Two Readers, Two Explanations: The Data Scientist and the QA Reviewer

Most explainability evidence fails not because the method was wrong but because it was written for the wrong reader. A feature attribution chart is a data scientist’s artifact. It answers a data scientist’s question: which inputs moved this output, and by how much. A QA reviewer or inspector is asking a different question: is this model making its decision for reasons that make sense in this process, and how do I know that is still true today.

What the literature says about the people who run the tools

Kaur and colleagues studied how data scientists actually use two interpretability tools, the InterpretML implementation of generalized additive models and the SHAP package, through a contextual inquiry with 11 participants and a survey of 197. They found that data scientists “over-trust and misuse interpretability tools,” and that few participants could accurately describe the visualizations the tools produced.10 That is a study of the experts, not the novices. If the people who generate SHAP plots misread them, a QA reviewer handed the same plot without an accompanying narrative has little chance.

Jesus and colleagues ran an application-grounded evaluation with professional fraud analysts, comparing decisions made with raw data only, data plus the model score, and data plus score plus explanations from LIME, SHAP, and TreeInterpreter. The analysts made the most accurate decisions with data only, and every explainer configuration underperformed that baseline, though explanations did help relative to seeing only the score. The authors warn that explainers chosen on proxy metrics like fidelity and stability “might be chosen that, in fact, hurt the overall performance of the combined system of ML model + end-users.”11 For a GxP process with a human in the loop, this is the direct warning: an explanation display can change operator behavior for the worse, and the human-AI team, not the model, is what gets evaluated under the FDA draft framework.3

Underneath both studies is a definitional problem that Lipton described in 2018: “interpretability” is not one property but a bundle of loosely related goals, and a claim that a model is interpretable is nearly meaningless unless someone says interpretable to whom and for what purpose.12 Doshi-Velez and Kim proposed distinguishing application-grounded evaluation (real people doing the real task), human-grounded evaluation (lay people on simplified tasks), and functionally-grounded evaluation (proxy metrics with no humans).13 Almost all explainability evidence produced in pharma today is functionally grounded. A reviewer can reasonably ask for at least one piece of evidence that a process expert used the explanation to reach a correct judgment.

What separates an actionable explanation from a technical one

SATISFIES A DATA SCIENTIST

Mechanism

A SHAP summary plot, a partial dependence curve, a saliency map. The artifact shows which inputs moved the output. It is internally consistent and reproducible from the notebook.

A QA REVIEWER CAN ACT ON

Meaning

The same artifact, plus a written statement from a process subject matter expert that the highlighted features are physically or procedurally relevant, why, and which features would be a red flag if they appeared.

SATISFIES A DATA SCIENTIST

Aggregate

Global feature importance across the test set. Average behavior, well summarized.

A QA REVIEWER CAN ACT ON

Instance

For the specific rejected batch or unit, what the model saw, what it attributed, what confidence it reported, and what the operator did next. Aggregate evidence cannot answer a question about a specific deviation.

SATISFIES A DATA SCIENTIST

Snapshot

Explanations generated at validation and filed with the validation report.

A QA REVIEWER CAN ACT ON

Lineage

Explanations tied to a model version, a test dataset version, an explanation-tool version, and a date, with a procedure that regenerates them on retraining and compares them to the last approved set.

The difference in every row is the same: the actionable version adds a human judgment, a record of that judgment, and a link to the model as it exists now. Rudin’s 2019 argument goes further: for high-stakes decisions, post hoc explanations of a black box model are themselves a source of risk, because they are approximations of the model rather than the model, and the better path is to build an inherently interpretable model where the problem permits.14 For many GMP classification problems on tabular data, that is a real option, and a reviewer may reasonably ask why a black box was chosen when an interpretable model would have met the acceptance criteria. The model selection rationale belongs in the evidence package for that reason.

The Evidence Types and Where Each One Breaks

Draft Annex 22 names SHAP, LIME, and heat maps as examples.1 The FDA draft asks for uncertainty and confidence estimation.3 Neither tells a quality team what the known weaknesses of these methods are, and an inspector who has read the literature may. What follows is each evidence type, what it actually shows, and the published limit that a package should acknowledge rather than hide.

Global feature importance

Global importance ranks input features by their average contribution across a dataset, whether from permutation importance, aggregated SHAP values, or model-native measures such as tree split gain. It answers “what does this model rely on in general.” It is the right first page of an evidence package because a process expert can read it without training: if a vial rejection model’s top feature is a lighting artifact rather than a defect signature, that is visible here. Its limit is that it is an average. A model can rely on sensible features on average and on a nonsensical one for a specific subgroup, which is exactly the case draft Annex 22 clause 3.2 asks companies to examine by dividing the input space into subgroups.1 Global importance should therefore be reported per subgroup, not only overall.

Local attribution: SHAP and LIME

SHAP, introduced by Lundberg and Lee in 2017, assigns each feature a contribution to a single prediction using Shapley values from cooperative game theory, and unifies several earlier attribution methods under one framework.15 LIME, introduced by Ribeiro, Singh, and Guestrin in 2016, fits a simple interpretable model in the neighborhood of a single prediction by perturbing the input and observing the output.16 Both are model-agnostic, both are widely implemented, and both are named in the Annex 22 draft. They are also the methods with the best-documented weaknesses.

Alvarez-Melis and Jaakkola argued in 2018 that a basic requirement of any explanation is that “similar inputs should give rise to similar explanations,” proposed metrics for that property, and showed that current methods “do not perform well according to these metrics.”17 In practical terms, two nearly identical images of the same vial can yield visibly different attribution maps. Slack and colleagues showed that a classifier can be deliberately constructed so that “post hoc explanations of the scaffolded classifier look innocuous” while the underlying predictions remain biased, and that such classifiers “can easily fool popular explanation techniques such as LIME and SHAP.”18 The relevance in GxP is not that a vendor is adversarial; it is that a method which can be fooled by construction is also a method that can be misled by ordinary correlation in training data. Kumar and colleagues added that “mathematical problems arise when Shapley values are used for feature importance,” that fixing them requires causal assumptions, and that the resulting explanations may not match what people actually need from an explanation.19

None of this means SHAP and LIME should be excluded. It means the evidence package should record the method version and settings (background dataset, number of perturbation samples, kernel width), report attribution stability on repeated runs and on slightly perturbed inputs, and state in plain language that these are approximations of the model’s behavior rather than a readout of its internals.

Saliency and heat maps for image models

For visual inspection models, heat maps are the natural explanation and the Annex 22 draft names them.1 Two papers should temper confidence in them. Adebayo and colleagues proposed “sanity checks” for saliency methods and found that several popular ones produce nearly the same map whether the model is trained or has randomized weights, meaning the map reflects the image’s edges rather than what the model learned.20 Ghorbani, Abid, and Zou showed that “interpretation of neural networks is fragile”: imperceptible changes to an input can produce a very different saliency map while leaving the prediction unchanged.21 A package that includes heat maps should state which saliency method was used, whether it passed the randomization sanity check, and whether maps for near-duplicate images agree. If nobody has run those checks, that is a finding waiting to happen.

Decision boundary and threshold documentation

For a classifier, the single most consequential design decision is where the threshold sits, and it is the piece of explainability evidence quality teams most often leave out because it does not look like an “explanation.” Annex 22 draft clause 9.2 asks for “an appropriate threshold setting” and for consideration of an “undecided” outcome at low confidence.1 The record should show the threshold value, the operating characteristic it was chosen on (sensitivity versus specificity trade-off, with the confusion matrix at that point), who approved it, the rationale linking it to the process (for example, that false accepts are the failure that matters and false rejects go to a manual re-inspection), the “undecided” band and what procedure handles it, and per-subgroup performance at the chosen threshold. A reviewer can act on that. A reviewer cannot act on “the model uses a 0.5 threshold.”

Worked example libraries

A worked example library is a curated set of representative inputs and the model’s outputs, each with attribution, confidence, and a process expert’s annotation of whether the decision and its reasons are correct. It should span the subgroups defined for intended use, include the rare variations Annex 22 draft clause 3.1 asks companies to characterize, and include cases at and near the threshold.1 This is the evidence type closest to how an inspector already thinks: show me a case, walk me through it. It is also the evidence type that best supports the Annex 22 draft’s “feature justification” review, because the justification is done case by case by a named expert and signed. Its limit is coverage. A library of forty cases proves the model behaved sensibly on forty cases; it should be presented as a qualitative complement to the quantitative test, never a substitute.

Counterfactual examples

Wachter, Mittelstadt, and Russell proposed counterfactual explanations in 2017 as a way to explain an automated decision without opening the model: state the smallest change to the input that would have changed the output.22 For a GMP model this translates to statements such as “this unit would have been accepted if the measured fill volume had been 0.3 mL higher” or “this batch would not have been flagged if the temperature excursion had been 4 minutes shorter.” Counterfactuals are unusually useful for QA review because they are testable: a process expert can say whether the stated change is physically meaningful. Their limit is that many counterfactuals exist for any decision, some of them physically impossible, so the package should state how the counterfactuals were generated and constrained to realistic input ranges.

Uncertainty and confidence reporting

The Annex 22 draft asks for confidence scores to be logged per outcome.1 The FDA draft asks sponsors to specify how uncertainty and confidence were estimated and to report performance with confidence intervals.3 The trap is that a softmax output is not a calibrated probability. Guo and colleagues showed in 2017 that “modern neural networks, unlike those from a decade ago, are poorly calibrated,” and that a simple post-processing step, temperature scaling, “is surprisingly effective at calibrating predictions.”23 A confidence threshold set on an uncalibrated score is a threshold on a number that does not mean what the procedure says it means. The evidence should include a calibration plot (reliability diagram) and the calibration method, or an explicit statement that the score is a ranking rather than a probability and that the threshold was set empirically on test data.

For teams that want a confidence statement with a guarantee, conformal prediction offers one. Angelopoulos and Bates describe it as a distribution-free method that wraps any model and produces prediction sets with a user-chosen coverage level, under an exchangeability assumption.24 In a classification setting, a conformal prediction set that contains both “accept” and “reject” is a principled version of the Annex 22 draft’s “undecided” flag, with a stated error rate rather than an arbitrary cutoff. That assumption of exchangeability is also why it must be re-checked after any change in input distribution.

Model cards and equivalent summaries

Mitchell and colleagues proposed model cards in 2019 as short documents accompanying trained models that report intended use, out-of-scope uses, evaluation data, performance disaggregated by relevant conditions and groups, and ethical considerations.25 In a GxP setting, the model card is the cover sheet of the evidence package: one or two pages a reviewer can read in five minutes that state what the model is for, what it is not for, how it performed on which data, and where the detailed evidence lives. It should not carry any claim not supported by a document behind it. None of the regulatory documents reviewed above requires a model card by that name, and the package should not claim otherwise; it should present the card as the company’s chosen way of meeting the joint principles’ call for “clear, accessible, and contextually relevant information.”4

Evidence typeWhat it showsPublished limitWhat the package must record
Global feature importanceWhat the model relies on across the test setAverages hide subgroup behaviorImportance per subgroup; SME statement on process relevance
Local attribution (SHAP, LIME)What drove one specific outputUnstable under small perturbations17; can be fooled18; mathematical issues19Method and version, settings, stability results, plain-language caveat
Saliency / heat mapsImage regions associated with the outputSome methods ignore the model20; fragile to tiny input changes21Method, sanity-check result, agreement across near-duplicate images
Threshold documentationWhere and why the decision line sitsMeaningless if score is uncalibrated23Threshold, operating point, approver, undecided band, subgroup performance
Worked example libraryCase-by-case reasonableness, expert-annotatedCoverage is limited to the cases chosenCase selection rationale, subgroup coverage, signed SME annotations
CounterfactualsSmallest input change that flips the output22Many exist; some are physically impossibleGeneration method, realism constraints, SME review
Uncertainty and confidenceHow sure the model is, per outputRaw scores are not probabilities23Calibration plot and method, or conformal coverage and its assumption24
Model cardOne-page summary of use, data, performance, limits25Only as good as the documents behind itCross-references to every supporting record; version and date
Do not present an explanation method as a readout of the model. SHAP, LIME, and saliency maps are approximations produced by a second algorithm run against the first. The published literature shows they can disagree with each other, with themselves on repeated runs, and with the model’s actual dependence on features. The honest formulation, and the one that holds up in an inspection, is: “This is what the attribution method reports; here is how stable it was; here is a process expert’s judgment on whether it makes sense; and here is the independent test evidence that the model performs to its acceptance criteria regardless.”

Tying the Evidence to Intended Use and Risk Classification

Every document reviewed above makes the same structural move: define intended use first, then assess risk, then scale the evidence to the risk. Explainability evidence is no exception. The depth of the package should be an output of the risk assessment, not a fixed template applied to every model.

Intended use defines what an explanation is for

Draft Annex 22 clause 3.1 asks for the intended use to be described “in detail,” including “a comprehensive characterisation of the data the model is intended to use as input and all common and rare variations,” with a process subject matter expert responsible for the adequacy of that description. Clause 3.3 adds that where a model gives input to a human decision and the testing effort has been reduced on that basis, the description “should include the responsibility of the operator.”1 The FDA draft’s Step 2 asks for the context of use, and Step 1 for the question of interest the model helps answer.3

That description determines the explanation audience. If the model is fully automated with a downstream release test as the safeguard (the FDA draft’s own fill-volume example, where release testing reduces model influence and lowers model risk to medium despite high decision consequence3), the explanation audience is the QA reviewer doing periodic review and the investigator handling a deviation. If the model presents a recommendation to an operator who makes the call, the explanation audience is the operator in real time, and the evidence must include what the operator sees, how they were trained to read it, and how the human-AI team performed in testing, which is what the FDA draft explicitly asks for.3 The Jesus study is the reason this matters: an explanation display can lower the accuracy of the combined system.11

Risk sets the depth

Using the FDA draft’s two axes, model influence and decision consequence, a practical scaling looks like this:

1

Low model risk (low influence, low consequence)

Model card, global feature importance with an SME relevance statement, threshold documentation, confidence logging. Worked examples optional. This is the floor, and it already exceeds what many deployed models have today.

2

Medium model risk (one axis high, mitigated by independent checks)

Everything above, plus per-subgroup importance, a worked example library covering all subgroups and near-threshold cases, calibration evidence, and documented attribution stability for the chosen local method.

3

High model risk (high influence, high consequence)

Everything above, plus counterfactual analysis, an application-grounded evaluation showing process experts reach correct judgments using the explanations, a documented rationale for why an interpretable model was not used or was rejected, and per-output attribution retained in production for the retention period, not only at test.

4

Any risk level, human in the loop

Add the operator-facing display specification, the training record for reading it, the human-AI team test results, and the Annex 22 draft clause 10.5 records of the human review process, which may mean a review of every output depending on criticality.

The risk classification itself has to be in the package, with the reasoning. An inspector who sees a thin explainability record for a model that rejects product will want to know whether the thinness was a decision or an oversight. A one-page risk rationale that cites model influence, decision consequence, and the mitigating controls turns a potential observation into a documented judgment call.

A subgroup finding is the most common way explainability evidence proves its worth. A model that performs to specification overall can rely on an irrelevant feature within one subgroup, such as a specific line, camera, or material lot. Global importance will not show it. Per-subgroup importance and a worked example library that deliberately includes each subgroup will. That is the Annex 22 draft’s feature justification review doing exactly what it is meant to do, and it is a much better outcome discovered during validation than during an investigation.

When the Model Belongs to a Vendor and the Internals Are Closed

Much of the AI in GMP manufacturing arrives inside a purchased system: a vision inspection platform, a predictive maintenance module, a chromatography data system feature. The regulated company often cannot see the architecture, the training data, or the weights. Draft Annex 22 clause 2.2 does not accept that as a reason for a thinner record. It states that documentation “should be available and reviewed by the regulated user irrespective of whether a model is trained, validated and tested in-house or whether it is provided by a supplier or service provider.”1 The GAMP AI Guide likewise emphasizes enhanced supplier assessment and collaboration for AI-enabled systems.6

What to ask the supplier for

The request list should mirror the FDA draft’s Step 4 description items, because they are the clearest published statement of what a reviewer expects to see about a model: inputs and outputs, architecture class, features, feature selection, parameters (or at least a statement of their count and governance), training data description and how labels were established, the evaluation approach, performance with confidence intervals, and how uncertainty was estimated.3 Suppliers can provide most of that without disclosing weights or proprietary training sets. Where they refuse, the refusal should be documented, and the gap closed with evidence the regulated company generates itself.

Evidence the regulated company can generate without the internals

Every explanation type in the table above except native model-specific importance can be produced against a closed model, because SHAP’s kernel variant, LIME, counterfactual search, and calibration all work on inputs and outputs alone. That is their entire reason for existing. So a closed vendor model does not prevent an explainability record; it changes who produces it and shifts the emphasis toward behavioral evidence:

  • Independent test data and independent metrics. The Annex 22 draft’s sections 5 through 7 on test data, independence, and execution apply in full, and the regulated company owns them regardless of where the model came from.1
  • Behavioral attribution. Model-agnostic attribution on the company’s own test set, with the stability checks described earlier, so that the “feature justification” review can still be performed by a process expert.
  • A worked example library built from the company’s own product and lines, since the supplier’s examples were not drawn from the intended use as this site defines it.
  • Counterfactual probing at the decision boundary, which is often the fastest way to discover that a closed model is sensitive to something it should not be.
  • Confidence logging and calibration on site data, with the threshold set and approved by the regulated company, not accepted as a vendor default.
  • Contractual commitments on change notification, so that a silent model update in a software release does not invalidate the record. The Annex 22 draft’s change control and configuration control clauses (10.1 and 10.2) require the company to detect unauthorized change; it cannot do that if the supplier does not tell it what changed.1

The device-side transparency principles are worth borrowing here for their honesty: model logic is to be communicated “when available.”9 A package for a closed vendor model should say plainly that the logic is not available, explain what was requested and what was received, and show that the behavioral evidence covers the intended use. That is a defensible position. Pretending the vendor’s marketing white paper is an explainability record is not.

Keeping the Evidence Current After Retraining

An explainability record is a statement about a specific model version on a specific test set with a specific explanation tool. Any of the three can change. The regulatory documents agree on the response. Draft Annex 22 clause 10.1 requires a tested model, its system, and the whole process to be under change control before deployment, and says any change to the model, the system, the process, or the physical objects it takes as input “should be documented and evaluated to determine if the model needs to be retested,” with any decision not to retest “fully justified.”1 The FDA draft says that depending on the extent and impact of a change, “some steps in the credibility assessment plan may need to be re-executed, including retraining and retesting the model,” and asks for life cycle plans that specify performance metrics, a risk-based monitoring frequency, and “triggers for model retesting.”3 The joint principles call for “scheduled monitoring and periodic re-evaluation.”4

Retraining as a change control event with an explanation clause

Most organizations already treat retraining as a change. Fewer treat the explanation record as part of what the change affects. A practical procedure adds four things to the retraining change record:

  1. Regenerate every explanation artifact against the new model version using the same test set version and the same explanation tool version, and file them as a new revision, never overwriting the old.
  2. Compare the new global importance and subgroup importance with the last approved set, and require a process expert to review and sign off on any material change in the ranking. A model that now depends on a different feature to reach the same accuracy has changed in a way the accuracy metric cannot see.
  3. Re-run the worked example library and flag every case where the decision, the attribution, or the confidence changed. Cases that flipped are the first things a reviewer will ask about.
  4. Re-check calibration and, if conformal methods are used, re-check coverage, since both depend on the relationship between model and data that retraining alters.23,24

Explanation drift as a monitoring signal

Annex 22 draft clauses 10.3 and 10.4 require monitoring of model performance metrics and of whether input data remain within the sample space, with defined drift metrics.1 Attribution offers a third monitoring channel that neither of those captures: the distribution of attributed features on production inputs. If a vision model that was approved because it attends to defect regions begins attending to background regions on a new lot of containers, performance metrics may not move for weeks, and the input distribution may look within range on the monitored variables. A periodic sample of production attributions compared against the approved reference set will catch it. This is not required by any of the documents reviewed here; it is a control that follows naturally from having built the record in the first place, and it converts an explainability package from a validation artifact into an operational one.

Version everything that produced the explanation. Explanation libraries are software. A SHAP or LIME package upgrade can change attribution values on an unchanged model. The evidence package should pin the model version, test dataset version, explanation tool and version, background or reference dataset, random seeds where applicable, and the date, so that a reviewer can ask “regenerate this” and get the same answer. If the answer differs, the difference is itself a finding to investigate.

The Evidence Package: A Proposed Table of Contents and the Inspector’s Questions

The following is a proposed structure for an explainability evidence package for a validated AI or machine learning model used in a GxP decision. It is intended to sit within, or be cross-referenced from, the system’s validation package under Annex 11 or the site’s computer software assurance approach, and to be scaled by risk as described above. Section numbers are suggestions.

Proposed table of contents

SectionContentsOwner
0. Model cardOne to two pages: intended use and out-of-scope uses; model version; data summary; headline performance with confidence intervals per subgroup; known limitations; confidence threshold; where each supporting record lives25Data science, approved by QA
1. Regulatory basis and statusWhich documents were used as benchmarks and their status as of the package date; the binding requirements the package satisfies (Annex 11, site procedures); explicit statement that draft documents are draftsQA / regulatory
2. Intended use and audience for explanationIntended use description per Annex 22 draft 3.1; subgroups per 3.2; human-in-the-loop responsibilities per 3.3; who needs explanations, when, and for what decision1Process SME
3. Risk classification and evidence depth rationaleModel influence and decision consequence assessment; mitigating controls; the resulting evidence tier and why3QA with process SME
4. Model description and selection rationaleInputs, outputs, architecture class, features, feature selection, parameters, training summary; why this model type; why an interpretable alternative was or was not chosen3,14Data science
5. Global feature importance and SME relevance reviewImportance overall and per subgroup; method and version; signed SME statement on process relevance and on features that would be red flagsData science, process SME
6. Local attribution and stabilityMethod, settings, version; attribution stability on repeated runs and perturbed inputs; plain-language caveat on approximation17,18,19Data science
7. Saliency evidence (image models)Method; randomization sanity check result; near-duplicate agreement20,21Data science
8. Decision threshold and undecided handlingThreshold, operating point, confusion matrix at that point, per-subgroup metrics, approver, undecided band and procedure1Process SME, QA
9. Confidence and uncertaintyHow confidence is computed; calibration method and reliability diagram, or conformal coverage and assumptions; confidence logging specification23,24Data science
10. Worked example libraryCase selection rationale; coverage matrix by subgroup and near-threshold; each case with input, output, attribution, confidence, and signed SME annotationProcess SME
11. Counterfactual analysis (higher risk)Generation method; realism constraints; SME review of physical plausibility22Data science, process SME
12. Human-AI team evidence (human in the loop)Operator display specification; training record; team performance in testing versus model alone; review records per Annex 22 draft 10.51,3Operations, QA
13. Supplier evidence and gaps (vendor models)What was requested, what was received, what was refused, and the behavioral evidence that closes each gap; change notification termsQA, procurement
14. Currency and change historyVersion pins for model, data, and tools; regeneration procedure on retraining; comparison against last approved set; production attribution monitoring planData science, QA
15. ApprovalsProcess SME, data science lead, QA; date; link to validation summary reportAll
4 Clauses in draft Annex 22 sections 8 and 9 that address explainability and confidence (8.1, 8.2, 9.1, 9.2)1
7 Steps in FDA’s draft risk-based credibility assessment framework, from question of interest to adequacy determination3
2 of 10 FDA and EMA joint principles that name interpretability or explainability directly (Principles 7 and 10)4

Questions an inspector is likely to ask

These are framed the way they tend to be asked: plainly, in sequence, each one following from the last. The package section that should answer each is noted.

Inspector question checklist

  • What does this model decide, and what happens downstream if it is wrong? (Sections 2 and 3)
  • Who decided how much explanation evidence was enough, and on what basis? (Section 3)
  • Why this type of model? Could a simpler one you can read directly have met the acceptance criteria? (Section 4)
  • What does the model rely on, and has someone who knows the process confirmed those features make sense? (Section 5)
  • Does it rely on the same things for every line, camera, material, or site? (Sections 5 and 10)
  • Show me a rejected unit or batch. What did the model see, what did it attribute, how confident was it, and what did the operator do? (Sections 10 and 12)
  • If I ran your attribution tool twice, would I get the same answer? What about on a nearly identical input? (Sections 6 and 7)
  • Where is the threshold, who set it, and what happens when the model is not sure? (Section 8)
  • Is the confidence score a probability? How do you know? (Section 9)
  • What is the smallest change to this input that would have changed the decision, and is that change physically meaningful? (Section 11)
  • What does the operator see, how were they trained to read it, and did you test the operator and model together or only the model? (Section 12)
  • This model came from a supplier. What did you ask for, what did they give you, and how did you fill the gaps? (Section 13)
  • The model was retrained in March. Which of these documents describe the model that is running today? (Section 14)
  • How would you know if the model started relying on something different without its accuracy changing? (Section 14)
  • You cite Annex 22. Is that in force? (Section 1)

The last question is not a trick. It is a check on whether the people presenting the package understand what they are presenting. The correct answer, given plainly, builds more credibility than any chart in the binder.

Conclusion

The regulatory texts on explainability are converging on a shared picture even while most of them remain unfinished: know the intended use, capture what the model relies on, have a process expert judge whether that is sensible, log and threshold confidence honestly, evaluate the human and model together where there is a human, and keep all of it current through change control and monitoring. What the texts do not supply is a warning about the tools. The peer-reviewed literature supplies that in abundance: attribution methods are unstable, can be misled, are misread by the experts who run them, and can make human decisions worse when displayed carelessly. An evidence package that takes both sources seriously looks different from one built to satisfy a checklist. It is organized around questions a reviewer would ask, it records judgments and who made them, it states the limits of its own methods, and it can be regenerated and compared after the model changes.

Sakara Digital works with pharma and biotech organizations building validation and governance records for AI and machine learning models in GxP processes, including the explainability evidence that has to survive a careful reviewer. If you are preparing a model for inspection, assessing a vendor system whose internals you cannot see, or deciding how much explanation evidence a given model actually needs, we are happy to have that conversation.

For Further Reading