In This Article
- Executive Summary
- Where the Rules Stand at the End of September 2026
- What Each Source Asks a Context of Use to Contain
- Why Most Context of Use Statements Fail an Inspector
- The Five Parts of a Statement an Inspector Can Follow
- Three Worked Examples
- Testing the Statement: The Inspector Walk-Through
- Keeping the Statement True After Go-Live
- Where the Statement Belongs in Your Document Set
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Every AI tool used in GxP work needs a short written answer to one question: what exactly is this model allowed to do, and for which decision? FDA’s January 2025 draft guidance calls that answer the context of use. Draft EU GMP Annex 22 calls it the intended use. Both documents are still drafts at the time of writing, but ICH M15, now adopted, uses the same idea for model-informed drug development, and the joint FDA and EMA principles from January 2026 list a clear context of use as one of ten principles.
The problem is rarely that a team has no AI context of use statement. The problem is that the statement describes the tool’s capabilities instead of its job, leaves out the human who acts on the output, and never connects to the evidence in the validation file. An inspector reads it and cannot tell what the model decides, who is accountable, or what would happen if it were wrong.
This article gives a five-part structure built from the regulators’ own words: what the model does, the decision it supports, who acts on it, the risk, and the evidence. It works the structure through three examples, sets out a walk-through test to run before an inspection, and explains how to keep the statement accurate after go-live and where it belongs in your document set.
Where the Rules Stand at the End of September 2026
Before writing a single line of a context of use statement, it helps to be exact about which documents exist and what force they have. Much of the commentary on AI in GxP work treats draft texts as though they were already binding. They are not, and a statement that cites a draft as a requirement will draw a question from any careful auditor.
FDA’s AI Credibility Guidance Is Still a Draft
In January 2025, FDA published the draft guidance Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, prepared by CDER with CBER, CDRH, CVM, the Oncology Center of Excellence, the Office of Combination Products and the Office of Inspections and Investigations.1 FDA’s guidance page for the document, checked while writing this article, still labels it a draft that is not for implementation, under docket FDA-2024-D-4689.2 FDA’s announcement said the draft was informed by the agency’s experience with more than 500 drug and biological product submissions with AI components since 2016, and by more than 800 comments on its 2023 discussion papers on AI in drug development and manufacturing.3
The draft sets out a seven-step, risk-based credibility assessment framework. Steps 1 and 2 are to define the question of interest and the context of use. Step 3 assesses model risk. Steps 4 through 7 plan, execute, document and judge the credibility work.1 The context of use matters most, because every step after step 2 is sized to it.
Two details matter for GxP teams who assume this guidance is only about submissions. First, the draft defines regulatory decision-making to include actions taken by sponsors and other interested parties in conformance with FDA’s authority, and it names current good manufacturing practice as an example.1 Second, it says the credibility assessment report may be held and made available to FDA on request, with an inspection given as the example.1 In other words, the reader of your context of use statement may well be an investigator on site, not a reviewer at a desk.
Draft Annex 22 Is Still a Consultation Text
Annex 22, “Artificial Intelligence,” is a proposed new annex to the EU GMP Guide. It was published as a consultation draft in 2025. Its section 3 is titled “Intended Use,” and it asks for a detailed description of what the model is designed to assist or automate.4 The EMA GMP/GDP Inspectors Working Group’s three-year work plan gives Q4 2026 as the target to provide the European Commission with a final text, in parallel with Annex 11 and Chapter 4.5 That is a target for a handover, not a publication date, and no date for coming into force has been set.
The scope of the draft is also under review. The draft limits itself to static models with deterministic output in critical GMP applications and says generative AI and large language models should not be used in critical GMP applications.4 EMA’s page for its June 30 to July 1, 2026 expert workshop says the 2025 consultation suggested support for potentially enabling generative AI or LLMs in medicines manufacturing, that EMA is still considering the implications, and that the drafting group sought expert input on control and mitigation measures such as guardrails as part of a proposed risk-based approach.6 Whatever the final scope, the intended use section is the part of the draft least likely to disappear, because every other section (acceptance criteria, test data, monitoring) refers back to it.
What Is Final: ICH M15 and the Joint Principles
Two documents that use the same vocabulary are not drafts. ICH M15, General Principles for Model-Informed Drug Development, was adopted by the ICH regulatory members under Step 4 on January 29, 2026. EMA’s Step 5 version lists July 23, 2026 as its date for coming into effect.9 M15 names artificial intelligence and machine learning among the modeling approaches it covers, and its assessment framework begins with a question of interest and a context of use, followed by model influence, consequence of wrong decision and model risk.9
In January 2026, FDA and EMA also published ten guiding principles of good AI practice in drug development. Principle 4, “Clear context of use,” reads in full: “AI technologies have a well-defined context of use (role and scope for why it is being used).”7 EMA’s announcement of the principles, dated January 14, 2026, says they will underpin future AI guidance in the different jurisdictions.8 The principles are not binding requirements, but they tell you which words regulators on both sides of the Atlantic now expect to see.
What Applies Today Regardless
None of this means there is no requirement today. FDA’s own draft says, in a footnote, that the use of AI in manufacturing must follow current good manufacturing practice, and that the quality control unit responsibilities in 21 CFR 211.22 and 211.68 apply.1 In the EU, draft Annex 22 describes itself as additional guidance to Annex 11, which means the existing Annex 11 expectations for computerized systems already cover a system with an embedded model.4 A context of use statement is the simplest way to show an inspector that you know what your validated system is for. That has always been expected. The drafts only give the expectation a name.
Premise check. This article treats FDA’s AI credibility guidance and Annex 22 as drafts, because that is what their issuing bodies’ own pages show at the time of writing. The structure recommended here does not depend on either becoming final. It rests on requirements that already apply (CGMP and Annex 11) and on vocabulary that is already final in ICH M15.
What Each Source Asks a Context of Use to Contain
The regulators use slightly different words, but when you lay their texts side by side they ask for much the same content. That is good news: one well-written statement can serve an FDA investigator, an EU inspector and your own quality unit.
FDA: Role, Scope and Other Evidence
FDA’s draft is direct about the core of it: “The COU defines the specific role and scope of the AI model used to address a question of interest.” It continues: “The description of the COU should describe in detail what will be modeled and how model outputs will be used.” The COU should also say whether other information will be used alongside the model output to answer the question.1 FDA’s press release put it in one line: “Context of use is defined as how an AI model is used to address a certain question of interest.”3
The draft also separates two ideas that teams often blur. The question of interest comes first and describes “the specific question, decision, or concern being addressed by the AI model.”1 The context of use comes second and describes the model’s part in answering it. Then model risk combines model influence (how much the model’s output counts relative to other evidence) with decision consequence (how bad the outcome of a wrong decision would be).1
Annex 22: Intended Use, Input Space and the Operator
Draft Annex 22 section 3.1 says: “The intended use of a model and the specific tasks it is designed to assist or automate should be described in detail based on an in-depth knowledge of the process the model is integrated in.”4 It then asks for a full description of the input data the model will see, including common and rare variations (what the draft calls the input sample space), and for known limitations and possible erroneous or biased inputs to be identified. Ownership is explicit: “A process subject matter expert (SME) should be responsible for the adequacy of the description, and it should be documented and approved before the start of acceptance testing.”4
Section 3.2 asks for subgroups of the input space where relevant, such as the decision output, the site or equipment, material characteristics, or defect types. Section 3.3 covers the human: where a model gives input to a decision made by a human operator and testing effort has been reduced on that basis, “the description of the intended use should include the responsibility of the operator.”4
ICH M15: Concise, Clear and Explicit
ICH M15 says the context of use “should be outlined as a concise, clear, and explicit description of the role and scope of the model(s) used to answer the question of interest.” It should include a description of the model, the data used to build it, and any additional data or evidence that will inform the answer.9 M15 also gives a plain rule for model influence: when model outcomes are the sole source supporting the decision, influence should be considered high.9
Where the Idea Came From
None of this is new. FDA’s AI draft notes that its question of interest, context of use and model risk concepts were informed by the ASME V&V 40 standard for computational models of medical devices.1 FDA’s final 2023 guidance on computational modeling credibility for device submissions, which covers physics-based and mechanistic models rather than machine learning, opened with a complaint that still applies: “Regulatory submissions often lack a clear rationale for why models can be considered credible for the context of use (COU).”13 On the drug side, CDER’s Biomarker Qualification Program asks for a single, concise context of use per qualification effort, written in a set structure of the biomarker category followed by its intended use in drug development.10 Outside health care, the NIST AI Risk Management Framework asks organizations to document intended purposes and deployment settings, the targeted application scope, the system’s knowledge limits, and how its output will be used and overseen by humans.14
| Element | FDA draft (Jan 2025) | Draft Annex 22 | ICH M15 (final) |
|---|---|---|---|
| The question or task | Question of interest (step 1) | Specific tasks the model assists or automates (3.1) | Question of interest |
| What the model does | What will be modeled; role of the model (step 2) | Intended use, based on in-depth process knowledge (3.1) | Description of the model and its role |
| Limits of the model’s part | Scope of the model (step 2) | Input sample space, limitations, erroneous and biased inputs (3.1, 3.2) | Scope; data used to build the model |
| Other evidence used | Statement on other information used with the output (step 2) | Not named as such; human review records where testing is reduced (10.5) | Additional data or evidence |
| The human | Human-AI team performance in evaluation (step 4) | Responsibility of the operator (3.3) | Not specific |
| Risk | Model influence plus decision consequence (step 3) | Risk to patient safety, product quality and data integrity (2.3) | Model influence plus consequence of wrong decision |
| Owner and approval | Sponsor, with early FDA engagement encouraged | Process SME, approved before acceptance testing (3.1) | Documented in a model analysis plan |
Read across the rows and a common shape appears. Each source wants to know the task, the model’s part in it, the boundaries, the other evidence, the human, the risk, and who stands behind the description. That shape is the basis for the five parts described below.
Why Most Context of Use Statements Fail an Inspector
An inspector reading a context of use statement is trying to answer four questions quickly. What does this thing decide? Who is accountable for the decision? What happens if it is wrong? Where is the evidence that it works for this purpose? A statement fails when it leaves any of those unanswered, or answers them in words that cannot be checked against records.
The failure patterns below are common, and none of them requires bad intent. They come from writing the statement too early, from the wrong seat, or by copying a vendor’s language.
It Describes Capability Instead of Use
“The system uses a convolutional neural network to detect visual defects with high accuracy” is a description of a product. It says nothing about which defects, on which line, for which product, feeding which decision. FDA’s draft asks what will be modeled and how the outputs will be used.1 A capability sentence answers neither. Vendor documents are written this way on purpose, because the vendor sells to many uses. Your statement must describe one.
It Hides the Decision
Many statements say the tool “supports” or “assists” a process. Support what, exactly? If the model output feeds a release decision, a batch disposition, a deviation classification or a data review, name the decision and name the record where it is made. An inspector will follow the output downstream. If the statement does not say where it goes, the inspector will find out from the operators, and the answer may not match your validation file.
It Leaves Out the Human, or Overstates Them
Two opposite problems show up here. Some statements omit the human entirely, which makes the model look like the sole decision-maker even when it is not. Others claim a “human in the loop” as a blanket control without saying who that person is, what training they have, what they check, and how their checks are recorded. Draft Annex 22 section 3.3 asks for the operator’s responsibility to be in the intended use description when reduced testing relies on human review.4 FDA’s draft asks that evaluation consider “the performance of the human-AI team, rather than just the performance of the model in isolation.”1
The research on automation bias is a good reason to take this seriously. A systematic review of 74 studies on automation bias, drawn from several research fields with a focus on health care, found that trust, workload, task complexity and time pressure all affect how much people over-rely on automated advice, and that training and emphasizing user accountability were among the factors that reduced it.15 A statement that says only “reviewed by a human” gives an inspector no reason to believe the review is meaningful.
It Has No Boundaries
A statement that does not say what the tool is not for invites scope creep. The deviation-drafting assistant that starts summarizing investigations ends up proposing root causes. The trend model built for one product gets pointed at a second product with a different process. Annex 22 asks for limitations and possible erroneous inputs to be identified up front.4 Boundaries written in plain language (“not used for products other than X,” “does not assign root cause”) give operators and inspectors a line they can see.
It Does Not Connect to Evidence
A statement that could be lifted into any validation package has not been tied to this one. The acceptance criteria, the test data, and the monitoring metrics should all trace back to specific words in the context of use. If the statement says the model classifies fill levels for 2 mL and 10 mL vials, the test data should include both, and the acceptance criteria should be set for both. Annex 22 makes the link explicit: test metrics measure performance “according to the intended use,” acceptance criteria are set for that intended use, and test data should cover its full sample space.4
Five warning signs in a draft statement.
- It could describe the same product at another company without edits.
- The word “assist” or “support” appears without naming a decision.
- The human reviewer has no role title, no training reference and no record.
- There is no sentence starting “The model is not used to…”
- You cannot point to a test case for each condition the statement names.
The Five Parts of a Statement an Inspector Can Follow
The structure below takes the common content from FDA’s draft, draft Annex 22 and ICH M15, and orders it the way an inspector will read it: from the task, to the decision, to the person, to the risk, to the proof. Each part is short. The whole statement should fit on a page, with references to the documents that hold the detail.
What the Model Does
The task in operational terms: inputs, output, the process step, the products and sites in scope, and the stated boundaries.
The Decision It Supports
The question of interest, the GxP decision the output feeds, the record where that decision is made, and what other evidence contributes.
Who Acts on It
The role that uses the output, what they check, their training, what they do when they disagree, and how their action is recorded.
The Risk
Model influence and decision consequence, each rated with a reason, and the resulting model risk, plus the control that keeps influence where you say it is.
The Evidence
Where the proof lives: acceptance criteria, test data, results, monitoring metrics and the change triggers that would require the statement to be revisited.
Ownership and Version
System and model version, process SME owner, quality approver, approval date, and the date the statement was last confirmed accurate.
Part 1: What the Model Does
Write this in the language of the process, not the language of machine learning. “Classifies each vial image from camera station 3 on filling line 2 as within or outside the fill-height specification” is better than “binary image classifier.” Name the inputs (image, sensor stream, text field), the output (a class, a score, a ranked list, a draft paragraph), and the process step where the model is used.
Then name the scope in terms an operator could check on the floor: which products, which presentations, which sites, which equipment. This is where the Annex 22 idea of an input sample space becomes practical. If the model was built and tested on clear glass vials, the statement should say so, and it should say that amber vials are out of scope until tested.4
Close Part 1 with boundaries. One or two sentences beginning “The model is not used to…” do more to prevent misuse than a page of procedure. Also state the model type in one line (for example, a fixed, trained classifier whose parameters do not change during use), because under draft Annex 22 the difference between a static and a dynamic model, and between a deterministic and a probabilistic output, decides which rules apply.4
Part 2: The Decision It Supports
Start with the question of interest in one sentence, written as a question. FDA’s manufacturing example is a good model: “Do vials of Drug B meet established fill volume specifications?”1 Then name the GxP decision the answer feeds (rejection of a unit, batch disposition, classification of a deviation, acceptance of a data set) and the record in which that decision is documented.
Next, state what other evidence contributes to the same decision. This is the part teams skip most often, and it is the part that most affects model risk. In FDA’s vial example, independent release testing of a representative sample means “the AI-based model will not be the sole determinant for the release of product.”1 That one sentence moved FDA’s model influence rating from what could have been high to low, and the overall model risk to medium, even though the decision consequence was high.1
Part 3: Who Acts on It
Name the role, not the person. “QC reviewer, trained per the named QC review SOP” can be checked against a training matrix. “The analyst” cannot. Say what that role does with the output: accept, reject, confirm, override, escalate. Say what they look at besides the model output, because a reviewer who sees only the model’s answer is not independent of it.
Then say what happens on disagreement. If the reviewer overrides the model, where is that recorded, and does anyone look at override rates? Draft Annex 22 says that where human review is used to reduce testing effort, records should be kept, and that depending on criticality this may mean consistent review of every model output according to a procedure.4 It also says the training and consistent performance of the operator should be monitored like any other manual process.4 Part 3 is where you point to that procedure and that monitoring.
FDA’s first qualified AI drug development tool shows how plainly this can be written. In announcing the December 2025 qualification of the AIM-NASH tool, FDA wrote that “pathologists are fully responsible for final interpretation, reviewing the whole slide image and AIM-NASH outputs before accepting or rejecting the AI-generated scores.”11 EMA’s March 2025 qualification opinion for the same tool described readings verified by one expert pathologist.12 Neither description leaves any doubt about who decides.
Part 4: The Risk
Use the two-factor structure that FDA’s draft and ICH M15 share. Rate model influence (low, medium or high) with a one-line reason. Rate decision consequence the same way. Combine them into model risk.1, 9
Two details from FDA’s draft help keep this honest. First, “Model risk is the possibility that the AI model output may lead to an incorrect decision that could result in an adverse outcome, and not risk intrinsic to the model.” A complex model is not high risk because it is complex. Second, FDA says the “decision consequence should consider the question of interest, but should not consider the COU of the model.”1 In practice, this means you cannot lower the consequence of a wrong release decision by pointing to your model’s controls. Controls lower influence. The consequence stays where the question puts it.
Finish Part 4 by naming the control that holds model influence at the level you claimed. If influence is low because of independent release testing, name the test and the SOP. If it is low because a trained reviewer confirms every output, name the procedure and the record. An inspector will test this claim first, because it is the one most likely to be untrue in daily practice.
Part 5: The Evidence
Part 5 is a map, not a summary. List where each piece of evidence lives: the acceptance criteria and who set them, the test data set and how its independence was protected, the test report, the monitoring metrics and their review frequency. Draft Annex 22 asks that acceptance criteria be set by a process SME before acceptance testing, and that they be at least as good as the performance of the process the model replaces.4 FDA’s draft asks for performance estimates with confidence intervals and for evaluation of the human-AI team where a human is in the loop.1
End with the triggers that would reopen the statement: a new product, a new site, a new camera, a new model version, a sustained change in override rate, or an input drift alarm. These triggers connect the statement to change control, covered later in this article.
A one-page template. Header: system, model version, owner (process SME), quality approver, approval date, last confirmed date. Part 1: task, inputs, output, process step, scope, boundaries, model type. Part 2: question of interest, decision fed, decision record, other evidence. Part 3: acting role, training reference, what they check, disagreement handling, record. Part 4: influence rating and reason, consequence rating and reason, model risk, control holding influence. Part 5: evidence map and change triggers.
Three Worked Examples
The examples below are illustrations. The first restates FDA’s own hypothetical manufacturing example in the five-part structure. The second and third are composites of common GxP uses, written to show how the same structure handles very different tools. None describes a specific company.
Example 1: Automated Fill-Level Inspection
FDA’s draft describes a manufacturer proposing an AI-based visual analysis system for 100 percent automated assessment of fill level in multidose vials of a parenteral product, where volume is a critical quality attribute.1 Written as a five-part statement, it might read as follows. FDA’s example stops at the risk assessment, so the details in the “who acts on it” and “evidence” rows are our own illustrations.
| Part | Statement text |
|---|---|
| What the model does | A fixed, trained image classifier analyzes the image of each filled vial of Drug B from the inline camera and flags vials whose fill level indicates a volume deviation. In scope: Drug B multidose vials on the named line. The model is not used for other products, other vial formats or other lines, and its parameters do not change during use. |
| Decision it supports | Question of interest: Do vials of Drug B meet established fill volume specifications? Output feeds automatic rejection of flagged units. Batch release also relies on independent fill-volume testing of a representative sample per batch, so the model is not the sole determinant of release. |
| Who acts on it | Rejected units go to the reject bin for count reconciliation by production. QC performs release testing independently of the model. QA reviews reject rates per batch against the trending limit. |
| Risk | Decision consequence: high (volume is a critical quality attribute). Model influence: low, because independent release testing measures fill volume on each batch. Model risk: medium. These ratings follow FDA’s own reasoning for this example.1 |
| Evidence | Acceptance criteria set by the process SME before testing, at least equal to the manual inspection performance they replace; independent test set covering all fill-level classes; test report; monthly monitoring of reject rate and input image quality. Triggers: new vial format, camera or lighting change, model retraining. |
Notice what this does for an inspector. The boundaries are visible. The claim that influence is low rests on a named test the inspector can look up in the batch record. The monitoring includes image quality, which reflects the Annex 22 expectation that performance monitoring should detect changes such as a change in lighting condition.4
Example 2: Deviation Triage Scoring
Consider a model that reads the free-text description of each new deviation record and suggests a preliminary classification (minor, major, critical) to help QA prioritize its queue. This is a common early use, and it is a useful test of the structure because the answer to “is this critical?” depends almost entirely on how the statement is written.
If the context of use says the model’s suggestion becomes the classification unless someone objects, the model is effectively deciding the depth of investigation for every deviation, and a wrong “minor” could delay the handling of a quality problem. If instead the statement says the model only orders the queue, that a trained QA reviewer assigns every classification from the full record, and that the suggestion is shown after the reviewer’s own choice, the model’s influence on the classification decision is low.
The five-part statement for the second version would state that the model orders the triage queue and displays a suggested class; that the question of interest is which new deviations QA should review first; that the classification decision is made and recorded by the QA reviewer in the deviation record; that the reviewer is trained per the named SOP and sees the full deviation record; that override rates are trended monthly; and that the model is not used to set investigation timelines, assign root cause or close records. The risk section would rate consequence as medium or high (a delayed review of a serious deviation matters) and influence as low, with the control being reviewer classification before the suggestion is shown.
That last design detail is where the writing of the statement changes the design of the system. If you cannot honestly write “the reviewer classifies before seeing the suggestion,” you cannot honestly claim low influence, and the research on automation bias gives an inspector good reason to ask.15
Example 3: An LLM That Drafts Batch Record Review Comments
Now take a large language model that reads a completed electronic batch record and drafts a list of review comments for a QA reviewer (missing entries, out-of-sequence timestamps, values near limits). Under draft Annex 22 as written, generative AI and LLMs should not be used in critical GMP applications, and where they are used in non-critical applications, qualified and trained personnel should always be responsible for ensuring the outputs are suitable for the intended use.4 EMA is reconsidering that line, but a statement written today should respect it.6
The context of use must therefore make the non-critical position real rather than asserted. Part 1 says the LLM produces a draft list of comments and nothing else, and that it does not change the batch record, disposition the batch, or sign anything. Part 2 says the question of interest is where the reviewer should look first, and that the batch record review itself is performed by QA against the full record per the review SOP, independent of the draft. Part 3 names the QA reviewer role, requires the reviewer to complete the full review, and records which draft comments were accepted, edited or rejected. Part 4 rates influence as low only because the full review is independent of the draft, and says so. Part 5 points to evidence that the draft helps rather than harms: for example, a comparison of review findings with and without the draft on a test set of past records, and a periodic check of whether reviewers are finding issues the draft missed.
If the organization later wants the LLM to replace part of the manual review (for example, reviewers skip sections the draft marks as clean), the context of use changes, influence rises, and under draft Annex 22 as it stands the use would move into territory the draft says should not use an LLM. The statement makes that line visible before someone crosses it.
What the three examples have in common. In each case, the model influence rating depends on a named control that exists outside the model: an independent release test, a reviewer who classifies first, a full review independent of the draft. Write the control into the statement, and then make sure the process delivers it every day.
Testing the Statement: The Inspector Walk-Through
A statement is only as good as its match to daily practice. The walk-through below is a short exercise to run before an inspection, or better, before go-live. It takes a few hours per system and needs three people: the process SME who owns the statement, a quality reviewer who did not write it, and one operator who uses the tool.
Read It Aloud to Someone Outside the Project
Ask a quality colleague from another area to explain back what the model decides and who is accountable. If they cannot, the statement is written for the project team, not for an inspector.
Follow One Output Downstream
Pick a recent output and trace it to the decision record named in Part 2. Confirm the record shows the human action named in Part 3. If the output went somewhere the statement does not mention, the statement is incomplete.
Ask the Operator, Not the Owner
Ask the operator what they do with the output and what they do when they disagree. Compare the answer with Part 3. Differences here are the most common finding, and the most useful one.
Test the Influence Claim
Find the control named in Part 4 in the records for a sample of recent uses. If influence is low because of an independent check, show that the check happened each time.
Match Each Scope Condition to a Test Case
For each product, site, format or input type named in Part 1, point to test data in the validation report. Any scope condition without a test case should be removed from scope or tested.
Ask “What if It Is Wrong Today?”
Walk through a wrong output. Who would notice, how quickly, and through which record? If the honest answer is “no one, until a complaint,” the influence or consequence rating is wrong.
Record the walk-through as a short memo with any gaps and their fixes. That memo is useful evidence in its own right, because it shows that the organization checks its own claims.
Questions to Expect in the Room
Inspectors will not use the words “context of use” in every case. They will ask plain questions that the statement should already answer:
- What does this system do, and what did people do before it existed?
- Which decisions depend on its output, and who signs for them?
- How do you know it works for this product and this line?
- What happens when the operator disagrees with it?
- What would tell you it had stopped working?
- What changed since validation, and how did you assess the change?
If each answer points to a part of the statement, and each part points to a record, the conversation stays short. If the answers come from memory, the conversation gets longer.
Keeping the Statement True After Go-Live
A context of use statement is accurate on the day it is approved. It becomes inaccurate as soon as the process around it changes and nobody updates it. Both FDA’s draft and draft Annex 22 treat the model, the system and the process around them as things that must stay under control after deployment.
Change Control Applies to the Process, Not Only the Model
Draft Annex 22 section 10.1 says the tested model, the system it is implemented in, and “the whole process it is automating or assisting” should be put under change control before deployment. Any change to the model, the system, or the process, including changes to physical objects the model uses as input, should be evaluated to determine whether the model needs retesting.4 FDA’s draft says that in manufacturing, changes to the model or to manufacturing that may affect model performance should be evaluated through the manufacturer’s change management system within its pharmaceutical quality system, and that some changes may need to be reported to the agency.1
The practical step is to add the context of use to the list of documents every change assessment checks. If a change touches anything named in Parts 1 through 3 (product, line, input, reviewer role, decision record), the statement is reviewed as part of the change. If the change alters what the statement says, it is a new context of use, and the credibility evidence in Part 5 must be checked against it.
Monitoring Tells You When the Statement Has Drifted
Draft Annex 22 asks for regular monitoring of model performance against its metrics, and of whether input data are still within the model’s sample space and intended use, with metrics defined for input drift.4 FDA’s draft says model performance metrics should be monitored on an ongoing basis and that the level of oversight should be commensurate with model risk and the context of use.1 The joint FDA and EMA principles call for scheduled monitoring and periodic re-evaluation, with data drift given as an example.7
Monitoring is also how you detect a less visible kind of change: use drifting beyond the statement without any formal change at all. Add two simple checks to periodic review. First, compare actual users and actual uses (from system access logs and a short operator survey) with Parts 1 and 3. Second, trend override or disagreement rates. A falling override rate can mean the model is improving, or it can mean reviewers have stopped checking. Either way, someone should look.
When the Statement Should Be Rewritten
| Event | Effect on the statement | Likely evidence needed |
|---|---|---|
| New product, format, site or line | Part 1 scope changes; may be a new context of use | Test data covering the new condition; acceptance criteria confirmed |
| Model retrained or replaced | Header version changes; Part 5 evidence must be refreshed | Retest against the same acceptance criteria; comparison with prior version |
| Independent check removed or reduced | Part 4 influence rises; model risk may rise | More rigorous credibility evidence, or restore the check |
| Reviewer role or training changes | Part 3 changes | Updated training records; human-AI team performance where relevant |
| Output used for a new decision | Part 2 changes; new question of interest | New risk assessment and credibility plan for that use |
| Input drift or performance alarm | Statement may no longer be true | Investigation; decision to retest, restrict scope or suspend |
FDA’s draft lists options for when credibility is not established for the model risk: reduce model influence by adding other evidence, increase the rigor of credibility work, add controls, change the modeling approach, or revise or reject the context of use.1 The same options apply when monitoring shows the statement is no longer true. Restricting the scope is a legitimate answer, and often the fastest one.
Where the Statement Belongs in Your Document Set
A context of use statement is most useful when it exists once, is controlled, and is referenced by everything else. The worst outcome is several versions: one in the user requirements, a different one in the risk assessment, and a third in the vendor’s configuration notes.
One Controlled Source, Many References
For most organizations, the natural home is the validation plan or an equivalent controlled document for the system, with the statement as its own section and its own approval. The user requirements, risk assessment, test plan, SOP and periodic review then reference it by section rather than restating it. Draft Annex 22 asks that the test plan contain a summary of the intended use, and that test documentation be retained along with the description of the intended use.4 A reference to the controlled statement meets that need without creating a second version.
ICH M15 offers a useful parallel for model-informed work: it expects the question of interest, context of use, model influence, consequence and model risk to be documented in a model analysis plan before the analysis, and reported afterward.9 For GxP systems, the matching documents are the approved statement, written before acceptance testing, and the validation summary that confirms the evidence in Part 5.
Link It to the AI Inventory
If your organization keeps an inventory of AI uses, each entry should carry a short form of the statement (the one-sentence task and the question of interest) and a link to the controlled version. This is how quality leadership sees, across the portfolio, which uses are critical, which rely on human review, and which are due for periodic review. The ISPE authors writing on AI governance in GxP environments recommend defining clear roles and responsibilities for developing, deploying and managing AI systems, and setting a policy for human oversight and for handling incorrect or harmful recommendations.16 The inventory, pointing to controlled statements, is where those roles and policies meet individual systems.
Suppliers Do Not Write Your Statement
Vendor documentation can describe the product’s intended purpose in general terms, and it can supply much of the model description and test evidence. It cannot describe your process, your decision, your reviewer or your independent checks. Draft Annex 22 says documentation should be available and reviewed by the regulated user whether a model is trained and tested in-house or provided by a supplier.4 ISPE’s July 2025 article introducing its GAMP Guide on artificial intelligence says the guide aims to help ensure AI-enabled computerized systems are fit for intended use.17 Fitness is judged against your intended use, which is why the regulated company must write and own the statement even when the vendor wrote the model.
Who Signs
Draft Annex 22 places responsibility for the adequacy of the intended use description on a process SME, with documentation and approval before acceptance testing.4 That is a sensible default even outside the EU. The process SME signs as author because they know the process. The quality unit approves because it is accountable for the decisions the model feeds. Data science and IT review for accuracy of the model description, but they should not be the authors, because the statement is about the process, not the algorithm. Draft Annex 22 also calls for close cooperation among process SMEs, QA, data scientists, IT and consultants across the model’s life.4
A simple ownership rule. The person who can best answer “what decision does this feed, and who acts on it?” writes the statement. The quality unit approves it. Everyone else reviews it. If nobody in the room can answer the question, the system is not ready for acceptance testing.
Conclusion
The regulatory texts on AI in GxP work are still moving. FDA’s credibility guidance remains a draft, draft Annex 22 is still being revised, and EMA is openly reconsidering the scope for generative AI. What is not moving is the first question every one of these documents asks: what is this model for, and what decision does it affect? ICH M15 has already put that question into a final guideline, and FDA and EMA have named a clear context of use as a shared principle. A statement that answers it plainly, with the decision, the person, the risk and the evidence all visible, will serve under whatever final text appears.
The value of writing the statement well goes beyond inspections. It is the moment when design choices become visible: whether the reviewer works independently of the model, whether the independent check still happens, whether the scope matches the test data. Sakara Digital works with pharma and biotech organizations on exactly this kind of AI governance and validation work. If you are preparing context of use statements for AI tools already in GxP use, or planning the first one, and want an independent view on whether an inspector could follow them, we are happy to have that conversation.
For Further Reading
For Further Reading
- FDA’s AI Credibility Guidance Is Still a Draft: What to Do While You Wait
- The AI Model Risk Assessment for Pharma: A Structured Checklist
- Human-in-the-Loop Requirements for Pharma AI: What FDA and EMA Actually Expect
- EMA Reopens the GenAI Question in Annex 22: Reading the June 2026 Expert Workshop
- When an AI Model Is Retrained: A Change Control Decision Tree
References & Sources
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products.” Draft Guidance for Industry and Other Interested Parties, January 2025. https://www.fda.gov/media/184830/download
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products.” Guidance document page (draft status, docket FDA-2024-D-4689). https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological
- U.S. Food and Drug Administration. “FDA Proposes Framework to Advance Credibility of AI Models Used for Drug and Biological Product Submissions.” Press announcement, January 6, 2025. https://www.fda.gov/news-events/press-announcements/fda-proposes-framework-advance-credibility-ai-models-used-drug-and-biological-product-submissions
- European Commission. “Annex 22: Artificial Intelligence.” EudraLex Volume 4, consultation draft, 2025. https://health.ec.europa.eu/document/download/5f38a92d-bb8e-4264-8898-ea076e926db6_en?filename=mp_vol4_chap4_annex22_consultation_guideline_en.pdf
- European Medicines Agency. “The 3-Year Work Plan for the Inspectors Working Group.” https://www.ema.europa.eu/en/documents/other/3-year-work-plan-inspectors-working-group_en.pdf
- European Medicines Agency. “Good Manufacturing Practice: Multistakeholder Workshop on Expert Contributions to Artificial Intelligence Guidance Development (Annex 22).” Event page, June 30 to July 1, 2026. https://www.ema.europa.eu/en/events/good-manufacturing-practice-multistakeholder-workshop-expert-contributions-artificial-intelligence-guidance-development-annex-22
- U.S. Food and Drug Administration. “Guiding Principles of Good AI Practice in Drug Development.” January 2026. https://www.fda.gov/about-fda/artificial-intelligence-drug-development/guiding-principles-good-ai-practice-drug-development
- European Medicines Agency. “EMA and FDA Set Common Principles for AI in Medicine Development.” News, January 14, 2026. https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0
- European Medicines Agency. “ICH M15 Guideline on General Principles for Model-Informed Drug Development, Step 5.” EMA/CHMP/ICH/496426/2024, February 9, 2026. https://www.ema.europa.eu/en/documents/scientific-guideline/ich-m15-guideline-general-principles-model-informed-drug-development-step-5_en.pdf
- U.S. Food and Drug Administration. “Context of Use.” Biomarker Qualification Program, CDER. https://www.fda.gov/drugs/biomarker-qualification-program/context-use
- U.S. Food and Drug Administration. “FDA Qualifies First AI Drug Development Tool, Will Be Used in ‘MASH’ Clinical Trials.” December 8, 2025. https://www.fda.gov/drugs/drug-safety-and-availability/fda-qualifies-first-ai-drug-development-tool-will-be-used-mash-clinical-trials
- European Medicines Agency. “EMA Qualifies First Artificial Intelligence Tool to Diagnose Inflammatory Liver Disease (MASH) in Biopsy Samples.” News, March 20, 2025. https://www.ema.europa.eu/en/news/ema-qualifies-first-artificial-intelligence-tool-diagnose-inflammatory-liver-disease-mash-biopsy-samples
- U.S. Food and Drug Administration. “Assessing the Credibility of Computational Modeling and Simulation in Medical Device Submissions.” Final Guidance, November 2023. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/assessing-credibility-computational-modeling-and-simulation-medical-device-submissions
- National Institute of Standards and Technology. “Artificial Intelligence Risk Management Framework (AI RMF 1.0).” NIST AI 100-1, January 2023. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- Goddard K, Roudsari A, Wyatt JC. “Automation Bias: A Systematic Review of Frequency, Effect Mediators, and Mitigators.” Journal of the American Medical Informatics Association 19(1), 2012. https://pubmed.ncbi.nlm.nih.gov/21685142/
- Mintanciyan A, Budihandojo R, English J, Lopez O, Matos J, McDowall R. “Artificial Intelligence Governance in GxP Environments.” Pharmaceutical Engineering, ISPE, July/August 2024. https://ispe.org/pharmaceutical-engineering/july-august-2024/artificial-intelligence-governance-gxp-environments
- Stockton B, Staib E, Heitmann M. “New GAMP Guide Addresses Challenges Posed by AI-Enabled Computerized Systems.” Pharmaceutical Engineering iSpeak, ISPE, July 24, 2025. https://ispe.org/pharmaceutical-engineering/ispeak/new-gampr-guide-addresses-challenges-posed-ai-enabled








Your perspective matters—join the conversation.