In This Article
- Executive Summary
- Why the Pre-Inspection Scramble Persists
- Two Regulator Signals That Draw the Boundary
- Where AI Helps: Five Readiness Tasks Worth Automating
- What Stays Human, and Why
- The Automate, Assist, or Keep Human Map
- A Readiness Operating Model: Daily, Weekly, Monthly
- Controlling the AI Tools So Readiness Does Not Become a Finding
- Making It Stick: Measures and Management Review
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Most pharma and biotech quality organizations still treat inspection readiness as an event. A notification arrives, or a pre-approval inspection is expected, and the site spends weeks pulling records, closing overdue reviews, and rehearsing answers. That model was tolerable when records were on paper and inspections were rare. It is no longer tolerable when audit trails, deviation systems, training records, and document management systems all generate data every day that an inspector can ask for at any moment. Readiness has to become a standing state, and AI is the first technology that makes continuous review of that data practical at a site scale.
The core question is not whether to use AI in readiness work but where. Two regulator positions from 2026 draw the boundary clearly. The FDA warning letter to Purolea Cosmetics Lab in April 2026 cited a firm that let AI create specifications, procedures, and master records without anyone verifying them2. The MHRA Inspectorate post in June 2026 described inspection responses that cited guidance that does not exist and set five expectations for any submission, whatever tool drafted it1. Read together, the message is consistent: AI can review, surface, sort, and draft. A named, accountable person decides, explains, and signs.
This article gives a working map of which readiness tasks to automate, which to assist, and which to keep fully human. It covers continuous audit trail review, CAPA and deviation trend surveillance, document currency checks, mock-inspection question banks built from public 483 and warning letter text, and first-pass response drafting. It then lays out a daily, weekly, and monthly operating model, and closes with how to control the AI tools themselves so the readiness program does not turn into an observation of its own.
Why the Pre-Inspection Scramble Persists
The scramble is not a failure of effort. Most quality teams work hard before an inspection. It is a failure of design. The regulations already require a great deal of ongoing review, but they express it in periodic terms, so organizations build periodic processes around it. The annual product quality review, the periodic evaluation of computerized systems, the biennial SOP review: each one is a real requirement, and each one creates a calendar date that can slip. When the calendar slips and an inspection arrives, everything that slipped becomes visible at once.
The requirements were always continuous in spirit
Look at what the rules actually say. 21 CFR 211.180(e) requires that written records be maintained so that the data in them “can be used for evaluating, at least annually, the quality standards of each drug product to determine the need for changes in drug product specifications or manufacturing or control procedures”8. “At least annually” is a floor, not a target. EU GMP Chapter 1 paragraph 1.10 calls for “regular periodic or rolling quality reviews” with the objective of verifying process consistency and highlighting trends, normally documented annually11. The word “rolling” is doing real work in that sentence. ICH Q10 section 3.2.1 asks companies to “plan and execute a system for the monitoring of process performance and product quality to ensure a state of control is maintained”9. A state of control is a condition that persists, not a report that gets written.
The same is true of computerized systems. EU GMP Annex 11 paragraph 9 says audit trails “need to be available and convertible to a generally intelligible form and regularly reviewed,” and paragraph 11 says systems “should be periodically evaluated to confirm that they remain in a valid state and are compliant with GMP”13. The WHO guideline on data integrity says the nature of the data should be “considered when determining the frequency of the audit trail review” and that systems used to record and store data “should be periodically reviewed for effectiveness”14. None of these documents says “review everything the month before the inspector arrives.” They describe an ongoing obligation that most sites have implemented as a series of deadlines.
What the public inspection data tells you
FDA publishes inspection observation data by fiscal year, with the area of regulation and the number of times it was cited on a Form FDA 483. Spreadsheets are available from fiscal year 2006 through fiscal year 2025, though FDA is careful to note that they “are not a comprehensive listing of all inspectional observations,” since some 483s are prepared manually and are not captured6. The FDA Data Dashboard adds inspection classifications (NAI, VAI, and OAI), a citations dataset, and a dataset of published 483s, updated weekly7. FDA also cautions that inspection frequency should not be read as a measure of facility quality.
For a readiness program, the value of this data is not the league table of most-cited sections. It is that the same categories recur year after year: investigations, quality unit responsibilities, procedures not followed, records incomplete. These are not exotic failures. They are the routine outputs of a quality system that reviews things on a schedule rather than as they happen. An inspector who asks for the deviation log and finds three investigations open past their due date has found a readiness gap that a daily query would have surfaced weeks earlier.
Periodic requirements were written when reviewing data continuously was impossible. Sites built periodic processes to meet them. The processes now generate the gaps that inspections find. Continuous readiness does not change the requirements. It changes the cadence of the review from “before the deadline” to “as the data arrives,” and that is exactly the kind of work machines do well and people do badly.
Two Regulator Signals That Draw the Boundary
Any leader planning to bring AI into readiness work should read two primary documents before writing a single requirement. Both are short. Both were verified against the regulator’s own site for this article, and both say less, and something more useful, than the commentary around them suggests.
The Purolea warning letter: AI-generated records nobody checked
On April 2, 2026, FDA’s Center for Drug Evaluation and Research issued a warning letter to Purolea Cosmetics Lab following an inspection conducted October 28 to 30, 20252. The letter cites the kinds of failures that appear in warning letters every month: finished products released without microbiological testing, components not tested for identity and purity, and a quality unit that did not review batch records before release. What made the letter unusual was a passage on artificial intelligence.
The firm told investigators it had used AI agents to help it comply with FDA regulations. In the letter’s words, “you used AI to create drug product specifications, procedures, and master production or control records.” FDA’s response was direct: “If you use AI as an aid in document creation, you must review the AI generated documents to ensure they were accurate and actually compliant with CGMP.” The letter cited 21 CFR 211.22(c), the quality unit’s responsibility for procedures. It also noted that the firm had not validated its manufacturing process under 21 CFR 211.100 and that management said it was unaware of the requirement because its AI agent had never told them it was required2. The response items FDA requested address products the firm had discontinued and ask for a commitment to notify FDA before resuming manufacturing.
Two things matter here for readiness. First, FDA did not object to AI. It objected to the absence of review. The sentence “you must review the AI generated documents” is the entire regulatory position, and it is the same position the agency has always taken on any document, whatever produced it. Second, the failure mode was not a subtle model error. It was a company that treated a drafting tool as a source of regulatory truth and never asked a qualified person whether the output was right. That is the exact failure a continuous readiness program has to be designed to prevent, because a readiness program that runs on AI will produce a great deal of AI-generated text about the state of the quality system.
The MHRA Inspectorate post: five expectations for any submission
On June 29, 2026, Peter Brown of the MHRA Inspectorate published a post on the use of AI for GxP inspection responses1. The Inspectorate had received responses that referenced MHRA guidance that does not exist, cited inappropriate regulatory frameworks, and in one case ran to more than 90 pages without addressing the deficiencies raised. In another case, a response to a deficiency with patient safety impact contained material inaccuracies, including references to sources that do not exist, and the Inspectorate’s review time rose from about four hours to more than twenty.
The post does not ban AI. It restates what every submission must be, and the list is worth quoting exactly because it doubles as an acceptance checklist for any AI-assisted readiness output. Submissions must be “factually accurate and verifiable,” “technically reviewed by appropriately experienced people,” “signed off by someone with authority and accountability,” “supported by evidence for factual claims,” and “appropriate to the specific regulatory context”1. The Inspectorate makes disclosure of AI use voluntary. If a company chooses to disclose, it should say so briefly at the start of the submission, identify the sections where AI assisted, and confirm that a person verified and approved them. The post states that “Inspectors will consider this positively when assessing organisational compliance.” On consequences, the Inspectorate says it may reject or return inaccurate or incomplete responses, treat the organization as higher risk for future inspection, and refer cases to its Inspection Action Group.
FDA’s March 2026 draft guidance on 483 responses
A third document rounds out the picture. In March 2026, FDA issued a draft guidance, “Responding to FDA Form 483 Observations at the Conclusion of a Drug CGMP Inspection,” under docket FDA-2025-D-1504, issued jointly by CDER, CBER, CVM, and the Office of Inspections and Investigations3. As of this writing it remains a draft, so its recommendations are not binding. The draft recommends that establishments choosing to respond do so within 15 business days, and states that a response received in that window will generally get a detailed review before FDA decides on further action4. It says responses “should be as accurate, clear, concise, and well-organized as necessary to convey an establishment’s position,” and it describes a detailed investigation report that includes scope, the associated drugs and lots, identified root causes and related systemic issues, the CAPA plan, completed and interim actions, and a planned effectiveness evaluation. It defines executive management as the senior person or persons with the authority to mobilize resources, and says the investigation team should have full support from executive management4.
The draft guidance does not mention AI. It does not need to. The MHRA post and the FDA draft converge on the same target: a short, accurate, evidence-backed response, owned by someone with authority, delivered fast. AI is useful for the “fast” part and the “well-organized” part. It is a hazard for the “accurate” and “verifiable” parts unless a person closes the loop.
FDA’s own FAQ states that a Form 483 “does not constitute a final Agency determination of whether any condition is in violation of the FD&C Act or any of its relevant regulations,” and that FDA considers the 483, the Establishment Inspection Report, the evidence collected on site, and “any responses made by the company” when deciding what further action is appropriate5. The response is part of the record. That is why ownership of it cannot be delegated to a tool.
Where AI Helps: Five Readiness Tasks Worth Automating
The tasks below share three properties. Each involves a large volume of structured or semi-structured data. Each has a clear rule for what “normal” looks like, or at least a reliable way to describe an outlier. And in each case the AI output is an input to a human decision, not the decision itself. That combination is what makes them good candidates.
1. Continuous audit trail review and anomaly flagging
Audit trail review is the readiness task where the gap between requirement and practice is widest. Annex 11 says audit trails must be “regularly reviewed”13, and the WHO guideline asks that the nature of the data drive the frequency of that review14. In practice, review is often a signature on a batch record confirming that someone looked at the audit trail for that batch. Nobody looks at the system-level audit trail across batches, users, and months, because no person can.
Automated review changes what is possible. A rules layer catches the patterns regulators have been citing for a decade: data modified after review, results reprocessed until they pass, entries by users whose training record expired, changes made under a shared login, deletions with no documented reason, and activity at times when the instrument should have been idle. A statistical or machine learning layer then looks for what the rules do not describe: a user whose edit rate doubled this month, an instrument whose aborted-run count is rising, a laboratory where reprocessing clusters around one product. The output is not a finding. It is a short list of events a data integrity reviewer should open, with the reason each one was flagged.
The reviewer still reviews. What changes is that they review the twenty events that matter out of two hundred thousand, and they do it this week rather than at the annual periodic evaluation. The audit trail of the review tool itself records what was flagged, what was dismissed, and by whom, which is precisely the evidence an inspector will want when they ask how the site reviews audit trails.
2. CAPA and deviation trend surveillance
ICH Q10 section 3.2.2 asks for a CAPA system that acts on “trends from process performance and product quality monitoring” and says the investigation should use “a structured approach” aimed at root cause9. Trending is where deviation and CAPA systems most often fall short. Sites trend by product, by category, or by month, using whatever fields the QMS exposes, and they miss the connections that live in free text.
Language models are good at reading free text at scale. Pointed at a year of deviation descriptions and investigation narratives, they can cluster events that were categorized differently but describe the same failure, surface recurring equipment, room, shift, or supplier references, and identify investigations whose root cause statement is a restatement of the symptom. Pointed at the CAPA register, they can flag actions whose effectiveness check was due and not done, actions that were extended more than once, and actions whose stated verification method could not possibly demonstrate effectiveness. FDA’s draft 483 guidance calls determining CAPA effectiveness “a fundamental part of evaluating whether the actions establishments take successfully address the associated issues”4. A tool that tells the quality head every Monday which effectiveness checks are overdue is not sophisticated, and that is the point.
3. Document currency checks
Document control failures are cheap to find and expensive to explain. EU GMP Chapter 4 paragraph 4.3 requires that instructional documents be “approved, signed and dated by appropriate and authorised persons” and that “the effective date should be defined”12. Every document management system knows the effective date, the review-by date, and the approval history of every SOP. Every learning management system knows which employees have completed training on the current version. Every change control system knows which changes are open and which reference documents they touch.
The readiness job is to join those three systems and ask the questions an inspector asks: Which SOPs are past their review date? Which employees performed a task in the last thirty days under an SOP version they were never trained on? Which open change controls have been open longer than their own procedure allows? Which periodic reviews of computerized systems are overdue, and which were closed with no findings in less time than a genuine review would take? None of these questions needs a language model. They need queries that run every night and a list that lands with the document control lead every morning. AI helps at the margin, for example by reading an SOP and flagging references to superseded documents, retired equipment, or organizational roles that no longer exist.
4. Mock-inspection question banks built from public 483 and warning letter text
FDA posts warning letters in full, and its inspection observation data lists cited sections by fiscal year6. The Data Dashboard adds a dataset of published 483s7. This is a large, public, regulator-authored corpus of what inspectors actually asked and what they found lacking. Most sites use it only as reading material.
A language model can turn it into a question bank. For each observation in the corpus that maps to a process the site runs, it can generate the question an inspector would have asked to arrive at that observation, the record they would have requested, and the follow-up they would ask if the first answer was weak. Filtered by dosage form, process type, and the site’s own deviation history, that becomes a mock-inspection script that is specific, current, and grounded in real enforcement language rather than a consultant’s memory. The questions are drafts. A quality leader who knows the site should cut the ones that do not apply and sharpen the ones that do. But the raw material, produced in an afternoon rather than a month, is a real gain.
5. First-pass response drafting that a named person then owns
This is the task where AI helps most and where the MHRA post was written. The Inspectorate saw responses that were long, wrong, and unowned. The remedy is not to avoid AI drafting. It is to constrain what the model is allowed to draft from and to make ownership explicit.
A good pattern is this. The model receives the observation text, the site’s actual investigation record, the actual CAPA plan, and nothing else. It is instructed to produce a response that follows the structure in FDA’s draft guidance: what the observation was, what the investigation found, root cause and systemic extent, actions taken and planned, and how effectiveness will be checked4. It is instructed not to cite any regulation or guidance not already present in the source material. The draft goes to a named owner, whose name appears on the response, who checks every factual statement against the underlying record, and who signs. The draft, the sources it was given, the reviewer’s edits, and the signature are retained. That retention is what lets the site answer the question “how did you prepare this?” with a record rather than a shrug.
In every case the AI narrows the field and a person decides. The audit trail tool flags; the reviewer opens. The trend tool clusters; the investigator judges. The document query lists; the document control lead acts. The question bank drafts; the quality head edits. The response drafter structures; the owner verifies and signs. If a proposed use of AI in readiness does not fit that sentence pattern, treat it as a use that has not yet been thought through.
What Stays Human, and Why
The boundary is not drawn by what AI cannot do. Models can write a persuasive escalation memo or a confident explanation of a finding. The boundary is drawn by what a regulator will hold a person to, and by what a person needs to have understood in order to be held to it. Four things stay human.
The accountable response
Both regulators say the same thing in different words. MHRA requires that a submission be “signed off by someone with authority and accountability”1. FDA’s draft 483 guidance ties the investigation to executive management with the authority to mobilize resources4. The Purolea letter cites 21 CFR 211.22(c), the quality unit’s own responsibility2. In every case the accountable party is a person or a defined organizational unit. A model can draft the response. It cannot be the respondent, and any process that lets the draft become the response without a person reading every line has recreated the Purolea failure with better formatting.
The decision to escalate
ICH Q10 section 3.2.4 says management review “should include a timely and effective communication and escalation process to raise appropriate quality issues to senior levels of management”9. A tool can tell you that a trend has crossed a threshold. Deciding that the trend is a quality issue senior management needs to hear about this week, rather than at the next scheduled review, requires knowing what senior management already knows, what else is happening at the site, what the commercial and patient consequences are, and what the organization’s own escalation procedure says. That is a judgment made by a person who will be asked later why they escalated, or why they did not.
Anything an inspector will ask a person to explain
Inspections are conversations. An inspector asks the laboratory manager why a result was reprocessed, asks the production supervisor how a deviation was contained, asks the quality head why a CAPA was extended. The person answering has to understand the answer, not recite it. A readiness program that produces excellent AI-drafted explanations that the people on the floor have not internalized has prepared documents, not people. The mock-inspection question bank is valuable precisely because a person has to answer the questions out loud, in their own words, before an inspector asks them.
The judgment on what a finding means
ICH Q9(R1), adopted in January 2023, is explicit that subjectivity in risk assessment “can directly impact the effectiveness of risk management activities and the decisions made” and that “it is important that subjectivity is managed and minimized,” and that the level of formality applied “may reflect the degree of importance of the decision, as well as the level of uncertainty and complexity”10. Deciding whether an anomaly is a data integrity event or a training gap, whether a cluster of deviations is a process problem or a documentation problem, and whether a finding at one site implies a systemic issue across the network is exactly that kind of judgment. AI can lay out the evidence on both sides. It cannot carry the accountability for choosing.
Ask of any readiness task: if this goes wrong, who will an inspector ask to explain it, and will that person be able to? If the honest answer is “nobody, the system did it,” the task is either in the wrong column or the control around it is missing. The Purolea letter is a record of a company that gave the second answer2.
The Automate, Assist, or Keep Human Map
The table below maps common readiness tasks to three categories. Automate means the task runs without a person in the loop and produces an output a person consumes. Assist means AI produces a draft, a ranking, or a summary that a person edits and approves before it is used. Keep human means AI may inform the task but does not produce its output. The right-hand column names the control that makes each row defensible.
| Readiness task | Category | What the AI or automation does | Control that makes it defensible |
|---|---|---|---|
| Rule-based audit trail screening (edits after review, shared logins, deletions without reason) | Automate | Runs nightly across all GxP systems; produces a flagged-event list | Rules documented and approved; tool validated; flagged list reviewed by a named data integrity reviewer with disposition recorded |
| Statistical or ML anomaly detection on audit trail activity | Assist | Ranks unusual user, instrument, or product activity for review | Model performance monitored; false-positive rate tracked; every flag dispositioned by a person; no automatic conclusions |
| Overdue CAPA, effectiveness check, and deviation due-date reporting | Automate | Queries the QMS daily; distributes an overdue list to owners and the quality head | Query logic validated against the QMS; list retained as a record; escalation path defined in procedure |
| Deviation and investigation free-text clustering and recurrence detection | Assist | Groups similar events, surfaces recurring references, flags weak root cause statements | Output treated as a lead; investigator confirms or rejects each cluster; clusters that trigger a CAPA go through the normal CAPA process |
| SOP review-date, training-currency, and open change control checks | Automate | Joins DMS, LMS, and change control data; produces a daily exception list | Join logic validated; exception list owned by document control; closure of each exception recorded |
| Reading SOPs for stale references (retired equipment, superseded documents, defunct roles) | Assist | Language model reads each procedure and lists suspect references | Every suggestion reviewed by the SOP owner before any revision; no automatic edits to controlled documents |
| Mock-inspection question bank generation from public 483 and warning letter text | Assist | Generates candidate questions, requested records, and follow-ups mapped to site processes | Quality lead curates the bank; sources retained; questions labeled as drafts until approved |
| First-pass drafting of an inspection response | Assist | Structures a draft from the observation, investigation record, and CAPA plan only | Closed source set; no citations beyond the sources; named owner verifies every statement and signs; draft and edits retained |
| Final inspection response content and sign-off | Keep human | None beyond formatting | Signed by a person with authority and accountability, per MHRA and FDA expectations |
| Deciding whether an anomaly is a data integrity event | Keep human | Presents the evidence | Decision and rationale recorded by the data integrity reviewer; quality risk management applied per ICH Q9(R1) |
| Deciding to escalate to senior management | Keep human | May summarize the trend and the procedure’s criteria | Decision recorded against the escalation procedure; ICH Q10 management review evidence |
| Explaining a process, record, or decision to an inspector | Keep human | Supports preparation only | Mock inspections rehearsed by the people who will answer; no reliance on live AI during the inspection |
| Determining systemic extent of a finding across sites or products | Keep human | May search records for similar events | Assessment owned by quality leadership; documented in the investigation per the FDA draft guidance structure |
Notice where the line falls. Everything in the “Automate” column is a query or a rule. Everything in “Assist” produces text or a ranking that a person has to accept. Everything in “Keep human” is a decision or an explanation for which a person will be asked to account. That is not a coincidence. It is the shape both regulators described.
A Readiness Operating Model: Daily, Weekly, Monthly
A map of tasks is not a program. The program is the cadence: what runs, who reads it, and what happens when it shows something. The model below is written for a single manufacturing or laboratory site with a functioning QMS, DMS, LMS, and at least one system with an electronic audit trail. It scales to a network by adding a monthly cross-site layer.
Daily: automated queries, human triage
Overnight, the audit trail screening rules run across all GxP systems, the QMS overdue query runs, and the document currency join runs. By the start of the shift, three short lists exist: flagged audit trail events, overdue CAPA and deviation items, and document and training exceptions. Each list has a named owner who triages it within the day. Triage means opening each item, deciding whether it needs action, and recording that decision. Most days the lists are short and most items are noise. The record that someone looked is the point.
Weekly: assisted review and trend reading
Once a week, the anomaly detection output and the deviation clustering output are reviewed by the data integrity reviewer and the investigations lead respectively. Each looks at what the models surfaced, confirms or rejects each cluster or ranked event, and records the outcome. The quality head receives a one-page summary: how many items were flagged, how many were confirmed, what was opened as a result, and which overdue items from the daily lists are still open after five business days. Overdue items that survive a week go to the site head, in line with the escalation procedure.
Monthly: readiness review and rehearsal
Once a month, the site holds a readiness review that does three things. It reviews the month’s daily and weekly outputs as a trend: are flagged events rising, are overdue items being closed, are the same procedures appearing on the exception list. It runs one mock-inspection session drawn from the question bank, with the people who would answer in a real inspection answering out loud. And it reviews the AI tools themselves: false-positive rates, any model or prompt changes, and any tool failures. The output feeds ICH Q10 management review, which is required to include “the results of regulatory inspections and findings, audits and other assessments, and commitments made to regulatory authorities”9.
Quarterly and annually: the periodic requirements, now pre-populated
The product quality review under 211.180(e) and EU GMP Chapter 1 paragraph 1.10, and the periodic evaluation of computerized systems under Annex 11 paragraph 11, still happen on their calendars. The difference is that the data has been reviewed all year. The annual review becomes a synthesis of twelve monthly reviews rather than a discovery exercise. When Chapter 1 paragraph 1.11 asks the manufacturer to “evaluate the results of the review” and decide whether CAPA or revalidation is needed11, the answer is already partly known.
What the model produces when an inspection is announced
The test of a continuous readiness program is what happens when the notification arrives. Under this model, the site does not start pulling records. It prints the last ninety days of daily and weekly outputs, the last three monthly readiness reviews, and the current exception lists. Those documents show an inspector, before they ask, that audit trails are reviewed, that overdue items are tracked and escalated, that documents are current, and that the site rehearses. The scramble is replaced by a briefing.
| Cadence | What runs | Who reads it | What is recorded |
|---|---|---|---|
| Daily | Audit trail rules; QMS overdue query; DMS, LMS, and change control join | Data integrity reviewer; CAPA coordinator; document control lead | Each flagged item and its disposition |
| Weekly | Anomaly ranking; deviation clustering; overdue survivors report | Data integrity reviewer; investigations lead; quality head | Confirmed and rejected clusters; escalations to site head |
| Monthly | Readiness review; mock-inspection session; tool performance review | Site quality leadership; function heads; the people who will answer | Trend summary; rehearsal outcomes; tool metrics; input to management review |
| Quarterly or annual | Product quality review; periodic system evaluation; management review | Site and network leadership | The regulatory records, built from the monthly reviews |
Controlling the AI Tools So Readiness Does Not Become a Finding
A readiness program built on AI creates a new set of computerized systems, and those systems are used to make GMP-relevant decisions about GMP records. They are in scope for exactly the controls the program is meant to verify elsewhere. The risk is obvious once stated: an inspector asks how the site reviews audit trails, the site describes its AI screening tool, and the inspector asks for the validation package. If the answer is thin, the readiness program has become the observation.
Classify each tool by what it does, not by what it is called
Annex 11 paragraph 1 says decisions on the extent of validation “should be based on a justified and documented risk assessment of the computerised system,” taking into account patient safety, data integrity, and product quality13. The tools in the readiness program are not all equal. A SQL query that lists overdue CAPAs is a simple reporting function against a validated QMS. A machine learning model that ranks audit trail events is a non-deterministic component whose output changes as it is retrained. A language model that drafts response text is a generative system whose output is different every time it runs. Each deserves a different level of control.
The ISPE GAMP Guide: Artificial Intelligence, published in July 2025, is the most complete industry framework for this classification. It is designed to be used alongside GAMP 5 Second Edition, addresses the full lifecycle of “design, development, operation, and use of AI,” and gives particular attention to “effective monitoring and maintenance, demonstrating ongoing control of complex AI-enabled computerized systems”15. Its central move is to apply GAMP’s existing risk-based thinking to components whose behavior is learned rather than coded, and to shift much of the assurance effort from one-time testing to ongoing performance monitoring. That fits a readiness program well, because the monthly tool review already exists.
Borrow the risk-based assurance logic where it applies
FDA’s guidance on Computer Software Assurance for Production and Quality Management System Software, finalized in September 2025 and reissued in February 2026, is written for medical device production and quality systems under 21 CFR Part 820, not for drug CGMP16. It is still worth reading for the way it describes “a risk-based approach to establish confidence in the automation used for production or quality management systems” and identifies where additional rigor is appropriate. The logic transfers even where the rule does not: concentrate scripted, documented testing on the functions whose failure would let a real problem go unseen, and use lighter methods for the rest.
Applied to the readiness program, that means the audit trail rules get formal testing with known-bad seeded data, because a rule that silently stops firing is a rule that hides exactly what it was built to find. The overdue-item queries get reconciliation against the QMS at each release. The anomaly model gets a documented baseline performance, a defined retraining trigger, and monthly monitoring of what it flags and what reviewers confirm. The language model used for drafting gets the least formal validation and the most formal procedural control, because its output is never used without a person reading it.
Design for the hallucination problem directly
The MHRA post described responses citing guidance that does not exist1. That failure is well documented outside pharma. A 2025 experimental study in JMIR Mental Health found that of 176 citations generated by a large language model across simulated literature reviews, 35 (19.9 percent) were fabricated, and of the 141 real citations, 64 (45.4 percent) contained errors. Fabrication was higher for more specialized topics18. A July 2026 preprint that checked bibliographies at major computer science conferences found reference-level hallucination rates usually below one percent, but reported that roughly one in twenty papers at two leading venues contained at least two likely hallucinated references19. Regulatory guidance is a specialized topic. A model asked to cite it will get some of it wrong.
The control is structural. The drafting tool is given only the site’s own documents and is instructed not to cite anything else. Any regulatory reference in a draft is checked by the reviewer against the regulator’s own site before it survives. The reviewer’s check is recorded. This is the same discipline any competent regulatory writer already applies. What changes is that the procedure now says so, and the record shows it was done.
Govern the tools as a set
The NIST AI Risk Management Framework, released in January 2023 and voluntary by design, organizes AI risk work into four functions: Govern, Map, Measure, and Manage. NIST added a Generative AI Profile in July 2024 that identifies risks specific to generative systems and proposes actions for them17. For a pharma quality organization, the framework’s value is as a checklist for the governance layer that sits above individual tool validation: who owns each tool, what data it may see, how changes to models and prompts are controlled, how performance is measured, and how incidents involving the tools are handled. The WHO guideline puts the accountability where it belongs: “Senior management is responsible for the establishment, implementation and control of an effective data governance system”14. AI tools that touch GxP data are part of that system.
Overdue items, document currency, rule-based audit trail screening
Treat as reporting functions on validated systems. Test with seeded data. Reconcile at each release. Keep the rule set under change control. Retain outputs as records.
Anomaly ranking, deviation clustering
Document baseline performance. Define retraining triggers and approval. Monitor confirmed-versus-flagged monthly. Never allow an automatic conclusion; every output is dispositioned by a person.
Response drafts, question banks, SOP reference review
Restrict sources to the site’s own records. Prohibit uncited regulatory claims. Require named human review and sign-off. Retain prompt, sources, draft, edits, and approval.
Ownership, access, change, monitoring, incidents
One inventory of readiness tools. One owner per tool. Model and prompt changes under change control. Tool performance on the monthly readiness review agenda. Failures handled as deviations.
If the site cannot produce, on request, the risk assessment for its audit trail screening tool, the test evidence for its rules, the disposition record for its flagged events, and the change history for its models and prompts, then the program is not ready for the inspection it was built to prepare for. Build those records from the first day, not after the first question.
Making It Stick: Measures and Management Review
Continuous programs fail for a simpler reason than periodic ones. Nobody cancels them. They fade, because the daily list gets long and the person who triaged it moves on, and after a few months the outputs exist but nobody reads them. An inspector who finds six months of unreviewed flagged events has found something worse than no tool at all: evidence that the site knew and did not act.
Measure the reviewing, not just the flagging
The program’s health measures are about human follow-through. What share of daily flagged items were dispositioned within one business day? What share of overdue items survived past five business days, and were they escalated as the procedure requires? How many mock-inspection sessions were held, and how many of the questions could the intended person answer without notes? What is the confirmed rate on anomaly flags, and is it stable? These are the numbers that show a person is in the loop. They are also the numbers that let leadership tune the tools, because a flag confirmed rate that falls toward zero means the model is producing noise and the reviewer will soon stop reading.
Put the program inside management review
ICH Q10 section 3.2.4 says management review “should provide assurance that process performance and product quality are managed over the lifecycle” and may be “a series of reviews at various levels of management”9. The monthly site readiness review is one of those levels. Its trend summary, rehearsal outcomes, and tool metrics belong on the agenda of the quarterly or annual management review, alongside inspection results and regulatory commitments. When they are there, the program has a sponsor who will notice when it fades.
Where this connects to quality management maturity
FDA’s Center for Drug Evaluation and Research runs a voluntary Quality Management Maturity program that “aims to encourage drug manufacturers to implement quality management practices that go beyond current good manufacturing practice (CGMP) requirements.” In February 2026 the agency announced a third cohort of up to nine establishments for its prototype assessment protocol20. The program is voluntary and separate from compliance determinations. It matters here because a standing readiness state, with continuous review and rehearsed people, is close to what “beyond baseline CGMP” looks like in practice. A site that can show twelve months of daily triage records and monthly rehearsals is demonstrating behavior, which is what maturity assessment looks at.
- Today, without preparation, could the site produce a list of every GxP audit trail event flagged for review in the last thirty days and what was decided about each one?
- Does anyone know, this morning, which CAPA effectiveness checks are overdue?
- When was the last time the people who would answer an inspector’s questions answered them out loud, in a rehearsal, using questions drawn from real observations?
- For every AI tool used in readiness work, is there a named owner, a risk assessment, test evidence, and a change history that would survive being asked for?
Conclusion
The move from a pre-inspection scramble to a standing state is not a technology project, although it needs technology. It is a decision about cadence: to review data as it arrives rather than before a deadline, and to rehearse people continuously rather than once. AI makes the reviewing practical at the volume modern systems produce. The two regulator positions from 2026 make the limits clear. FDA’s Purolea letter shows what happens when AI-generated records reach the quality system without a person checking them. The MHRA Inspectorate post lists what every submission must be regardless of what drafted it, and every item on that list is something a person does. Between those two markers, there is a great deal of useful automation: audit trail screening, trend surveillance, document currency checks, question banks, and first drafts. Outside them, there is the accountable response, the escalation decision, the explanation to an inspector, and the judgment about what a finding means. Those were always human, and nothing in 2026 changed that.
Sakara Digital works with pharma and biotech organizations building this kind of readiness program: mapping the tasks, designing the daily-to-annual cadence, and putting the right level of control around the AI tools so the program itself holds up under inspection. If you are exploring continuous inspection readiness and want an independent perspective on where to start, we are happy to have that conversation.
For Further Reading
For Further Reading
- MHRA on Using AI to Draft GxP Inspection Responses: The New Standard
- Annex 22 Mock Inspection: What a Pharma Quality Team Should Practice Now
- The 2026 AI Warning Letter Trend Analysis: What Followed Purolea
- CAPA Effectiveness Reviews: The 90-Day Look-Back I Recommend
- Periodic Review for AI Systems: What GxP Requires
- Inspection Readiness Is a Mindset
References & Sources
- Brown, Peter. “Use of AI for GXP inspection responses: setting standards without stifling innovation.” MHRA Inspectorate blog, June 29, 2026. https://mhrainspectorate.blog.gov.uk/2026/06/29/use-of-ai-for-gxp-inspection-responses-setting-standards-without-stifling-innovation/
- U.S. Food and Drug Administration, Center for Drug Evaluation and Research. “Purolea Cosmetics Lab, Warning Letter 722591.” April 2, 2026. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/purolea-cosmetics-lab-722591-04022026
- U.S. Food and Drug Administration. “Responding to FDA Form 483 Observations at the Conclusion of a Drug CGMP Inspection.” Draft Guidance for Industry, docket FDA-2025-D-1504, March 2026. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/responding-fda-form-483-observations-conclusion-drug-cgmp-inspection
- U.S. Food and Drug Administration. “Responding to FDA Form 483 Observations at the Conclusion of a Drug CGMP Inspection: Guidance for Industry (Draft).” Full text PDF, March 2026. https://www.fda.gov/media/191427/download
- U.S. Food and Drug Administration. “FDA Form 483 Frequently Asked Questions.” https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/inspection-references/fda-form-483-frequently-asked-questions
- U.S. Food and Drug Administration. “Inspection Observations.” Fiscal year 2006 through 2025 datasets. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/inspection-references/inspection-observations
- U.S. Food and Drug Administration. “FDA Data Dashboard: Inspections.” https://datadashboard.fda.gov/oii/cd/inspections.htm
- U.S. Government Publishing Office. “21 CFR 211.180: General requirements (Subpart J, Records and Reports).” Code of Federal Regulations, 2025 edition. https://www.govinfo.gov/content/pkg/CFR-2025-title21-vol4/pdf/CFR-2025-title21-vol4-sec211-180.pdf
- International Council for Harmonisation. “ICH Harmonised Tripartite Guideline Q10: Pharmaceutical Quality System.” Step 4, June 4, 2008. https://database.ich.org/sites/default/files/Q10%20Guideline.pdf
- International Council for Harmonisation. “ICH Harmonised Guideline Q9(R1): Quality Risk Management.” Final version adopted January 18, 2023. https://database.ich.org/sites/default/files/ICH_Q9%28R1%29_Guideline_Step4_2023_0126.pdf
- European Commission. “EudraLex Volume 4, Chapter 1: Pharmaceutical Quality System.” January 2013. https://health.ec.europa.eu/system/files/2016-11/vol4-chap1_2013-01_en_0.pdf
- European Commission. “EudraLex Volume 4, Chapter 4: Documentation.” January 2011. https://health.ec.europa.eu/system/files/2016-11/chapter4_01-2011_en_0.pdf
- European Commission. “EudraLex Volume 4, Annex 11: Computerised Systems.” Revision 1, 2011. https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf
- World Health Organization. “TRS 1033, Annex 4: WHO Guideline on data integrity.” WHO Expert Committee on Specifications for Pharmaceutical Preparations, 2021. https://www.who.int/publications/m/item/annex-4-trs-1033
- International Society for Pharmaceutical Engineering. “ISPE GAMP Guide: Artificial Intelligence.” July 2025. https://ispe.org/publications/guidance-documents/gamp-guide-artificial-intelligence
- U.S. Food and Drug Administration. “Computer Software Assurance for Production and Quality Management System Software: Guidance for Industry and FDA Staff.” Final, February 2026 (original September 2025), docket FDA-2022-D-0795. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/computer-software-assurance-production-and-quality-management-system-software
- National Institute of Standards and Technology. “AI Risk Management Framework.” AI RMF 1.0, January 26, 2023; Generative AI Profile (NIST AI 600-1), July 26, 2024. https://www.nist.gov/itl/ai-risk-management-framework
- Linardon, Jake, et al. “Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study.” JMIR Mental Health, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12658395/
- Russinovich, Mark, Ram Shankar Siva Kumar, and Ahmed Salem. “Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences.” arXiv preprint 2607.00738, July 2026. https://arxiv.org/abs/2607.00738
- U.S. Food and Drug Administration, Center for Drug Evaluation and Research. “CDER Quality Management Maturity.” https://www.fda.gov/drugs/pharmaceutical-quality-resources/cder-quality-management-maturity








Your perspective matters—join the conversation.