What the Purolea Letter Actually Says

On April 2, 2026 the FDA issued a warning letter to Purolea Cosmetics Lab. It is the first FDA warning letter with a deficiency section written specifically about artificial intelligence, under the heading “Inappropriate Use of Artificial Intelligence in Pharmaceutical Manufacturing.”1 Because it is being cited widely and loosely, it is worth being precise about what it says and what it does not.

The firm told the investigator that it had used AI agents to help it comply with FDA regulations. Specifically, AI was used to create drug product specifications, procedures, and master production or control records that were intended to meet FDA requirements. Those AI-generated documents were then used without further review. The agency’s position is stated plainly: if you use AI as an aid in creating documents, you must review the resulting documents to confirm they are accurate and actually compliant with CGMP, and the failure to do so is a violation of 21 CFR 211.22(c), the provision that makes the quality control unit responsible for approving or rejecting procedures and specifications affecting identity, strength, quality, and purity.12

The second finding is the one people skim past, and it is the more instructive of the two. The investigator found the firm had not conducted process validation before distributing its drug products, as required under 21 CFR 211.100. When told of the requirement, the firm’s response was that it had not known process validation was mandatory because its AI tool never mentioned it.1 That is a different failure from the first. The first is a firm accepting a wrong document. The second is a firm treating an AI system as the authority on what its regulatory obligations are, and inheriting an omission it never noticed because nothing in the output flagged an absence.

Be accurate about the scope of this letter. Purolea is a cosmetics and over-the-counter drug operation, not a clinical-stage biotech or a large-molecule manufacturer. The letter also cites insanitary conditions and unapproved new drugs, and the firm has since ceased drug production.1 The reason it matters to a well-run pharma quality organization is not that the circumstances are comparable. It is that the regulatory logic the agency applied is completely general: the responsible unit owns the document regardless of what produced the draft, and an AI system’s silence on a requirement is not a defense.

The transferable finding

Strip away the specifics and the letter establishes two things that apply to every regulated organization drafting validation documents with AI. First, the accountability for a document does not move to the tool that drafted it. Second, an AI-drafted document can fail by omission as easily as by error, and omission is far harder for a reviewer to catch, because the reviewer is looking at what is on the page rather than at what should have been on the page and is not.

Most review processes are built to catch the first failure mode and are structurally blind to the second. A reviewer reading a well-formed user requirements specification with sixty-two requirements has no natural prompt to ask which requirements are missing. That is the problem this article is really about.

The Validation Document Set, Broken Into Its Parts

Computerized system validation produces a recognizable set of documents. In GAMP terms, and in the sequence most quality management systems follow, the set runs roughly as follows: a validation plan, a user requirements specification, functional and configuration specifications, a risk assessment, test protocols and scripts covering installation, operational, and performance qualification, a requirements traceability matrix, deviation records raised during execution, and a validation summary report.3

These documents are not the same kind of thing as each other, and treating them as one category is the error that produces bad AI policy. Some of them are records of a decision. Some are transformations of an earlier document into a different shape. Some are compilations of evidence produced by execution. The AI question has a different answer for each.

Annex 11 of the EU GMP guide, which remains the applicable European expectation for computerized systems, frames the set in a way that makes the distinction visible. Clause 4.1 says validation documentation and reports should cover the relevant steps of the life cycle, and that manufacturers should be able to justify their standards, protocols, acceptance criteria, procedures and records based on their risk assessment. Clause 4.4 says user requirements specifications should describe the required functions of the system and be based on documented risk assessment and GMP impact, and that user requirements should be traceable throughout the life cycle. Clause 4.7 says evidence of appropriate test methods and test scenarios should be demonstrated, that system parameter limits, data limits and error handling should be considered, and that automated testing tools and test environments should have documented assessments for their adequacy.4

Read those three clauses together and the shape of the answer appears. The words carrying the weight are “justify,” “based on documented risk assessment,” “considered,” and “assessments for their adequacy.” Every one of those is a human act of judgment. The documents themselves are the visible residue of those acts. AI can produce residue that looks identical whether or not the act occurred.

Three categories, not one

CATEGORY A

Derived documents

The content is fully determined by an existing approved artifact. A test script derived from an approved requirement, a traceability matrix derived from a requirements list and a test inventory, a configuration specification derived from a system’s actual configuration export. The source exists, is approved, and is checkable.

CATEGORY B

Compiled documents

The content is an assembly of evidence that already exists in executed form. A validation summary report compiled from executed protocols, deviation records, and their dispositions. Nothing new is asserted. The work is retrieval, arrangement, and accurate restatement.

CATEGORY C

Decision documents

The content is a judgment that no prior artifact contains. The risk classification of a function, the decision about what to test and how deeply, the acceptability of a deviation observed during execution, the conclusion that a system is fit for its intended use.

THE LINE

Where the policy splits

Categories A and B are transformation problems, and AI drafting is a real efficiency gain with a verifiable output. Category C is a judgment problem, and an AI draft there produces a document that reads as though a decision was made when it was not.

Where AI Drafting Genuinely Works

Start with the honest case for AI in this work, because it is stronger than the skeptics allow. Validation documentation is structurally repetitive, heavily templated, and full of restatement. A large share of the effort in producing a validation package is not thought. It is the mechanical conversion of an approved statement into a differently formatted statement, repeated a few hundred times.

The DIA Global Forum published an industry review of AI applied to computer system validation in November 2025 that lists the same candidate deliverables most teams arrive at independently: validation plans and summary reports, user requirements specifications, traceability matrices, test plans and test scripts, and checklists of documents required for a given system type. The authors are clear that all of it requires review by subject matter experts and quality professionals before finalization, and that errors go undetected when AI is relied on exclusively.5 That is the right framing, and the reason to take the efficiency case seriously is that the tasks it names are genuinely mechanical.

Traceability matrices

This is the clearest win in the entire document set, and it is worth explaining why. A requirements traceability matrix asserts a relationship between items that already exist: requirement to specification, specification to test, test to result. Every cell in the matrix has a determinate correct value that can be checked against the source documents. There is no judgment in the matrix itself. The judgment happened when someone decided what the requirements were and what tests would challenge them.

An AI system building a traceability matrix from an approved requirements list and an approved test inventory is performing a lookup and a formatting operation. Errors are possible, but they are the kind of error a reviewer can catch by sampling, because each assertion is independently checkable in a few seconds against a document sitting on the same desk. Annex 11 clause 4.4 requires user requirements to be traceable throughout the life cycle, and the traceability artifact is exactly the kind of document where automation reduces transcription error rather than introducing new risk.4 GAMP 5 second edition encourages the use of automated tools for testing and traceability for the same reason.6

Configuration specifications

A configuration specification records what the settings of a system actually are. Where the system can export its own configuration, the specification is a transformation of that export into a readable, structured, reviewable document. This is one of the better uses of AI drafting, because the ground truth is machine-readable and the verification step is a comparison rather than an assessment.

The caution is narrow and specific: the specification must be generated from the configuration export, not from the vendor’s documentation of what the configuration options are. Those two sources look similar and diverge in exactly the places that matter. A specification generated from a product manual describes a system that could exist. A specification generated from an export describes the system you have.

Validation summary reports

A summary report compiles executed evidence. The protocols have been run, the results recorded, the deviations raised and dispositioned. The report restates that record in a form a reader can follow. AI drafting handles the restatement well, and the restatement is a large share of the page count.

There is one clause in a summary report that does not belong to this category, and it is the last one: the conclusion that the system is fit for its intended use and may be released. That sentence is a decision, and it belongs in Category C no matter what produced the surrounding forty pages. The practical rule is that AI may draft the body of a summary report and must not draft the conclusion statement or the disposition of any deviation the report summarizes.

First-pass requirements from an existing source

Requirements are a mixed case, and the mixing is where teams go wrong. Where a requirement already exists in another form, a URS from a comparable system, a process description, a regulatory clause, an approved user story, restating it as a testable requirement is a transformation. Where the requirement does not yet exist anywhere, drafting it with AI is not transformation. It is invention presented as recall, and it is where an omission of the Purolea kind enters.

The test to apply. Before letting AI draft any validation document, ask one question: could a competent reviewer verify every assertion in this document against a source that already exists and is already approved? If yes, AI drafting is appropriate and the review is a verification. If no, the document contains a judgment, and drafting it with AI means the judgment either was never made or was made by the model.

Where It Fails: Documents That Carry an Unmade Judgment

Now the harder half. The documents AI drafts worst are precisely the documents teams most want help with, and the reason is not a limitation that better models will fix. It is a category error about what the document is for.

Risk classification

A GxP risk assessment records a decision about how much a particular function matters to patient safety, product quality, and data integrity, and therefore how much assurance effort is warranted. GAMP 5 second edition is built around this: the appropriate approach is defined by knowledgeable and experienced subject matter experts applying critical thinking to the specific system, its intended use, and its context.6 The output is a classification. The input is knowledge of your process, your product, your patients, and your operational reality.

An AI system drafting a risk assessment has the classification vocabulary and none of the input. What it produces is a plausible distribution of high, medium, and low across a list of functions, weighted by how similar functions are usually classified in the material it was trained on. That distribution will be right often enough to be dangerous. It will be wrong in the places where your process differs from the general case, which is exactly where a risk assessment earns its existence.

The failure is not that the model gets a classification wrong. Humans get classifications wrong too, and the quality system has mechanisms for that. The failure is that the document arrives already filled in, which removes the occasion on which the subject matter expert would have had to think. A blank risk assessment forces a conversation. A pre-filled risk assessment invites agreement.

Test scope

Deciding what to test is the central act of risk-based validation, and it is the decision the entire computer software assurance approach exists to sharpen. FDA’s guidance on computer software assurance, reissued February 3, 2026 as “Computer Software Assurance for Production and Quality Management System Software” and superseding the September 2025 version, describes an approach built on risk-based testing, unscripted testing, continuous performance monitoring, and reliance on activities performed by other parties such as developers and suppliers.7 The guidance is addressed to device and biologics manufacturers rather than to drug CGMP directly, but the reasoning has shaped how pharma validation groups think, and GAMP 5 second edition moved in the same direction.6

That whole approach depends on someone deciding, for this function in this system with this intended use, whether a failure poses a high process risk and what assurance activity is proportionate. Writing in Pharmaceutical Engineering, Walia and Neri describe critical thinking as planning first and creating documentation from that plan, rather than producing paperwork that does not itself produce quality.8 An AI system asked to propose test scope inverts that order. It produces the documentation and leaves the plan implied.

Deviation acceptability during execution

This is the sharpest case, and the one worth writing into policy in the strongest terms. During protocol execution, a test fails or produces an unexpected result. Someone must decide whether the deviation affects the validity of the qualification, whether it needs correction and retest, whether it is a script defect rather than a system defect, and whether the system remains fit for its intended use.

Annex 11 clause 4.2 requires validation documentation to include reports on any deviations observed during the validation process.4 The report is the record of a judgment made under conditions of genuine uncertainty, usually with commercial pressure to conclude the deviation is minor. This is the single worst place in the document set to accept an AI draft, because the model has a strong tendency to produce a well-argued justification for whatever conclusion the prompt implies, and the prompt is being written by someone who wants the qualification to close.

A rule worth writing verbatim into the SOP. No AI system may draft the assessment or disposition of a deviation observed during validation execution, and no AI-generated text may appear in a deviation record other than a factual restatement of what was observed. The reason is not that the model would be wrong. It is that the record must show a person weighed the evidence, and a fluent AI-drafted justification is indistinguishable on the page from a considered one.

The fitness-for-intended-use conclusion

Every validation package ends with a person asserting that the system does what it was specified to do and may be used. FDA’s computer software assurance guidance describes the record as including a conclusion statement declaring acceptability of the software for its intended use, plus a record of who performed the testing or assessment and the date, and established review and approval where appropriate.7 The conclusion statement is the point of the whole exercise. It is an assertion by a named individual with signatory authority. Delegating its drafting to a system that cannot hold accountability makes the signature decorative.

The Specific Hazard of AI-Generated Test Scripts

Test scripts deserve their own treatment because they occupy an unusual position. They are formally Category A, derived from approved requirements, and therefore look like a safe automation target. They are also the document where AI drafting fails in the most dangerous way available, because the failure is invisible to ordinary review.

The hazard is this. An AI-generated test script derived from a requirement will reliably produce steps that exercise the described function and record that it worked. What it will not reliably produce is a test that could have failed if the requirement were not met. A script can be complete, well-formed, traceable to its requirement, executable, and entirely incapable of detecting the defect it exists to detect.

The evidence from software engineering research

This is not speculation. It is one of the most consistently reproduced findings in the research on machine-generated tests, and the numbers are sobering.

40.2% Average mutation score of LLM-generated unit tests on a benchmark of real-world functions, against 45.1% statement coverage9
41.3% Average accuracy of those same generated tests in the same study9
Context-dependent Replication study finding on whether coverage and mutation scores predict real fault detection for generated suites10

Mutation score measures whether a test suite would actually catch a defect: small faults are deliberately introduced into the code and the suite is scored on how many it detects. It is the closest available proxy for the question a validation professional cares about, which is whether the test would fail if the system were broken. A benchmark study of unit test generation from real-world functions found generated tests averaged 41.32% accuracy, 45.10% statement coverage, 30.22% branch coverage, and 40.21% mutation score, substantially below what earlier and easier benchmarks had suggested.9

A 2026 replication study at ISSTA examined whether coverage and mutation scores of generated test suites correlate with their real effectiveness, and found the usefulness of both is highly context-dependent. In regression settings where the code is assumed correct, the metrics give a meaningful comparative signal. Where the code under test may already contain defects, which is the situation every qualification test is in, they stop being reliable indicators.10 A systematic literature review of large language models in unit testing reaches the underlying reason: generating reliable test oracles is the hard part, because it requires capturing the intended design specification rather than reflecting the behavior the implementation happens to have.11

That last sentence is the whole problem restated in software engineering terms. A test derived from what the system does will always pass. A test derived from what the system was required to do is the only kind worth executing. An AI system given both a requirement and access to the system has a strong pull toward the first.

What this looks like in a qualification package. Every script traces to a requirement. Every script executed. Every script passed. Coverage of requirements is 100%. And the acceptance criteria were written to describe the behavior observed rather than the limit the requirement imposed, so nothing in the package would have failed if the configuration had been wrong. This package will pass an internal review. It will not survive an inspector who picks one script and asks what result would have constituted a failure.

What Annex 11 already requires here

European expectations anticipated this more directly than most teams realize. Annex 11 clause 4.7 states that evidence of appropriate test methods and test scenarios should be demonstrated, that system parameter limits, data limits and error handling should be considered in particular, and that automated testing tools and test environments should have documented assessments for their adequacy.4

Two obligations follow. First, the boundary and error cases are called out by name, and those are the cases a generated script is least likely to construct, because they require reasoning about what the system should refuse to do rather than what it should do. Second, if you use an AI tool to generate test scripts, that tool is an automated testing tool, and Annex 11 already expects a documented assessment of its adequacy. That assessment is not a vendor questionnaire. It is a demonstration, on your own requirements, that scripts produced by the tool detect seeded defects.

The countermeasure that actually works

The countermeasure is not more review of the script text. Reviewers reading generated scripts approve them, because the scripts are well-formed and traceable and there is nothing on the page that signals weakness. The countermeasure is to change what the reviewer is asked to do.

For every AI-drafted script covering a requirement classified as high risk, the reviewer answers one question in writing before approval: what result, if observed during execution, would cause this script to fail? If the answer is “the system does not display the expected screen,” the script is weak. If the answer names a specific limit, a specific rejected input, or a specific error condition, the script is doing work. This question takes a competent reviewer under two minutes per script and catches the failure mode that no amount of proofreading will.

For a sample of high-risk scripts, go further and run them against a deliberately misconfigured instance. If the script passes against a system configured wrongly, you have measured its value directly. This is the documented adequacy assessment Annex 11 asks for, and it is a one-time exercise per tool and script pattern rather than a per-script burden.

The Review Model That Keeps a Named Human Accountable

The Purolea letter’s core finding is about review, so the review model is where policy has to be strongest. The problem is that “reviewed and approved by a qualified person” is already in every validation SOP, and it did not prevent the failure it was written to prevent. Something has to change in what review means.

Why ordinary review fails on AI drafts

Human review of machine-generated material has a measurable and specific failure pattern, and understanding it changes the design of the control. A randomized experiment with 2,784 participants examining how people evaluate AI-generated suggestions found two results that matter here. Requiring corrections for flagged AI errors reduced engagement and increased the tendency to accept incorrect suggestions. And attitude toward AI predicted performance more strongly than demographics: participants skeptical of AI detected errors more reliably and achieved higher accuracy, while those favorable toward automation showed marked overreliance.12

Both findings are uncomfortable for the standard control design. The first says that making review more burdensome makes it worse, not better, which undercuts the instinct to require line-by-line sign-off on everything. The second says the quality of your review depends on who is doing it in a way that is not captured by qualification records, and that the enthusiastic early adopter running your AI pilot is the least reliable reviewer of its output.

Related work on how people judge text by its labeled origin found raters strongly favored content labeled as human-generated over content labeled as AI-generated, and the pattern held even when the labels were swapped.13 Perceived provenance moves judgment independently of content. In a validation context that means the same document gets a different review depending on whether the reviewer knows how it was drafted, which is an argument for disclosure being mandatory rather than optional.

Verification, not approval

The change that fixes this is small to state and substantial in practice. The reviewer of an AI-drafted validation document is not asked whether they agree with it. They are asked to independently verify specific things and to record what they verified.

DocumentWhat the reviewer independently verifiesWhat the reviewer may not simply approve
User requirements specificationThat every requirement traces to a stated business or regulatory need; that the process owner has confirmed the list is complete for the intended useCompleteness. A reviewer cannot verify completeness by reading. It must be confirmed against an independent source.
Functional and configuration specificationThat each stated setting matches the actual system configuration export, by sampling or full comparisonAny setting described from vendor documentation rather than the live configuration
Risk assessmentThat the classification rationale references this process and this product, not generic languageThe classification itself. It must be made by the subject matter expert, not confirmed after the fact.
Test scriptThat the acceptance criterion states a condition the system could fail; that boundary, limit, and error cases appear where the requirement implies themTraceability alone. A traced script is not a challenging script.
Traceability matrixThat a defined sample of cells matches the source documents in both directionsNothing. This document is fully verifiable and is the safest in the set.
Deviation recordThat the factual description matches what was observed during executionThe assessment and disposition, which must be authored by a person
Validation summary reportThat every claim in the body is supported by an executed record, by samplingThe conclusion statement, which must be authored and signed by the accountable individual

The three structural rules

1

The drafter and the reviewer are different people, and neither is the model

A person prompts, edits, and submits the draft as their own work product. A second person verifies. The AI system is a tool used by the first person, not a party to the review. This sounds obvious and is violated constantly when one person generates a document and routes it straight to quality approval.

2

Disclosure is mandatory and specific

The reviewer is told which sections were AI-drafted before they review, not after. Given the evidence that perceived origin changes how carefully people read, telling them is the point. A reviewer who knows a section was generated reads it differently, and that difference is the control.

3

The reviewer records what was checked, not that it was checked

A signature saying “reviewed and approved” produces no evidence that verification happened. A short record naming the source compared against, the sample size, and the discrepancies found produces evidence an inspector can evaluate. This is also the only version of the control that resists the pattern where review effort decays over time.

Avoiding the burden trap

One caution follows directly from the experimental evidence: adding verification steps everywhere makes review worse. If every AI-drafted sentence requires an independent check, reviewers stop checking and start signing, and the recorded control becomes a fiction that is harder to detect than an absent one.12

Scale verification depth to the risk classification of the requirement the document covers. Full independent verification for anything supporting a high-risk function. Defined-percentage sampling for medium. Spot checks for low. This is the same proportionality logic that governs testing depth, applied to review, and it is defensible for exactly the same reason.

The Scope Statement for Your Validation SOP

Here is the substance a company can adapt into its validation SOP or into an AI-use annex referenced from it. It is written to be read by a validation lead, a quality reviewer, and an inspector, and it deliberately makes the permitted and prohibited uses specific enough to audit against. Adapt the wording to your document naming, and keep the structure.

Section 1: What AI may draft

An approved AI tool may be used to produce a first draft of the following, provided the source material named in each case exists and is approved before drafting begins:

  • Requirements restated from an existing approved source, including a prior URS, an approved process description, a cited regulatory clause, or an approved change request.
  • Functional and configuration specifications generated from a system configuration export or an approved design document.
  • Test scripts derived from an approved requirement, subject to the challenge review in Section 3.
  • Requirements traceability matrices generated from an approved requirements list and an approved test inventory.
  • The body and evidence-summary sections of a validation summary report, generated from executed protocols and completed records.
  • Formatting, structural, and consistency edits to any validation document, including terminology alignment and template conformance.

Section 2: What AI may never draft

No AI tool may be used to draft, propose, or suggest content for the following. These must be authored by a named qualified individual:

  • The risk classification of any GxP function, and the rationale supporting that classification.
  • The determination of test scope, test depth, or which requirements warrant formal scripted testing.
  • The assessment, impact evaluation, or disposition of any deviation observed during validation execution.
  • The conclusion statement in a validation summary report, or any statement that a system is fit for its intended use.
  • A requirement that does not exist in any prior approved source. New requirements are elicited from the process owner, not generated.
  • Any statement about what a regulation requires. Regulatory requirements are cited from the regulation, and their applicability is determined by a qualified person.

The last item is written directly against the Purolea failure pattern, in which a firm treated an AI system’s silence as evidence that a requirement did not apply.1

Section 3: What the reviewer attests to

The reviewer of any AI-drafted validation document signs an attestation that states, in this form rather than as a general approval:

  • I was informed which sections of this document were produced with AI assistance before reviewing it.
  • I independently verified the items listed in the verification record attached to this document, against the sources named there.
  • For each test script covering a high-risk requirement, I have recorded the observed result that would constitute a failure of that script.
  • I confirmed that no content in this document falls within the prohibited categories in Section 2.
  • I am accountable for the content of this document as though I had authored it.

Section 4: What is recorded about the tool and the prompt

For each AI-drafted validation document, the validation record includes:

  • The tool name and version or model identifier, and the date of use.
  • The identity of the person who prompted and edited the draft.
  • The prompt or prompt template used, retained in full.
  • The source documents supplied to the tool, identified by document number and version.
  • The sections of the finished document that originated as AI output, identified at section level.
  • The verification record produced under Section 3, including sample sizes and discrepancies found.

These records are retained for the same period and under the same controls as the validation package itself.

Section 5: Qualification of the tool itself

An AI tool used to generate test scripts is an automated testing tool within the meaning of Annex 11 clause 4.7 and requires a documented assessment of its adequacy before use.4 That assessment demonstrates, on a representative sample of the organization’s own requirements, that scripts produced by the tool detect deliberately introduced configuration faults. The assessment is repeated when the tool or model version changes.

Why the SOP is written this way

Three design choices in the above are worth explaining, because they are the ones people push back on.

First, the prohibited list is a list of decisions, not a list of documents. This matters because it survives changes in your document set. If your quality management system merges the risk assessment into the validation plan next year, the rule still applies to the classification decision wherever it now lives.

Second, the record captures the prompt. Teams resist this, on the grounds that prompts are working material rather than records. The counterargument is straightforward: the prompt is the specification given to the tool that produced a GxP document, and it is the only artifact that shows what the tool was asked to do. If a document later proves wrong, the prompt is what tells you whether the tool was misused or the tool is inadequate. Those two conclusions lead to completely different corrective actions.

Third, the attestation is written in first person and names specific acts. Approval language that says “reviewed for accuracy and completeness” produces no evidence and no discomfort. Language that says “I am accountable for the content of this document as though I had authored it” produces both, and the discomfort is the control working as intended.

What Gets Recorded About the Tool and the Prompt

The last piece is the record itself, and it deserves separate treatment because it is where two regulatory expectations meet and where most implementations are thinnest.

FDA’s computer software assurance guidance sets out what an assurance record should generally include: the intended use of the software feature or function, the result of the risk-based analysis, a description of the testing conducted, issues found during testing, a conclusion statement declaring acceptability for the intended use, a record of who performed the testing or assessment and the date it was performed, and established review and approval where appropriate, including a signature and date from an individual with signatory authority. The guidance also says documentation need not include more evidence than necessary to show the function performs as intended for the risk identified, and recommends using digital records such as system logs and audit trail entries rather than duplicating results already retained digitally by the software.7

That structure is directly usable for AI-drafted documents, with one addition. The FDA record answers what was tested, by whom, and with what conclusion. An AI-assisted validation record has to answer one further question: how do I know a person made the decision this document records?

The four fields that answer it

FieldWhat it capturesWhy an inspector asks for it
ProvenanceWhich sections were AI-drafted, at section granularity, with tool and versionEstablishes the scope of the enhanced review, and lets an inspector target sampling
InstructionThe prompt and the identified source documents supplied to the toolDistinguishes tool inadequacy from tool misuse when an error is found
VerificationWhat the reviewer checked, against which source, at what sample size, and what discrepancies were foundConverts an approval signature into evidence that verification occurred
DecisionThe named individual who made each Category C judgment, and the basis they recordedDemonstrates the judgment was made by a person, which is the finding in the Purolea letter

The fourth field is the one that answers the enforcement question directly. If an inspector asks how you know the risk classification in this assessment was a human decision rather than an accepted default, the answer needs to be a record, not a policy statement. A field naming the subject matter expert and holding two or three sentences of process-specific rationale answers it. A signature block does not.

Discrepancies are evidence, not embarrassment

One practical point that changes how these records perform under inspection. Teams instinctively record verification as clean: sample checked, no findings, approved. A verification record that never finds anything is not reassuring. It reads as a control that is not operating.

Record the discrepancies you find and what you did about them. A traceability matrix verification that found three mismatched cells out of a fifty-cell sample, corrected them, and expanded the sample is a demonstrably functioning control. The same verification recorded as “sample checked, acceptable” is indistinguishable from no verification at all. The same logic applies to the challenge review of test scripts: recording that four scripts were rewritten because their acceptance criteria could not have failed is the strongest evidence you can offer that the review is real.

A reasonable first target. If you are starting from nothing, do not begin with a policy covering every document. Begin with traceability matrices and summary report bodies, which are the safest and highest-volume wins. Add test scripts only once the challenge review and the tool adequacy assessment are running. Leave risk assessments, test scope, deviations, and conclusion statements out of scope entirely, and say so explicitly in the SOP so that the exclusion is a documented decision rather than an oversight.

Where the regulatory picture is heading

Two things are worth watching without over-reading either. The ISPE GAMP Guide on artificial intelligence, published in 2025, gives the industry its first structured framework for applying GAMP concepts to AI in GxP contexts, and it is the reference most validation groups will build against.14 Separately, the European Commission consulted in 2025 on a new Annex 22 dedicated to artificial intelligence in GMP, alongside revisions to Annex 11 and Chapter 4.15 Annex 22 is still draft, has no final text, and has no implementation date, so nothing in it is an obligation today.

The point for a validation SOP being written now is that none of the controls described in this article depend on Annex 22 being finalized. They follow from 21 CFR 211.22(c), from Annex 11 clauses 4.1, 4.2, 4.4 and 4.7, and from the ordinary expectation that a validation record shows who decided what.24 A company that puts this scope statement in place is not anticipating a future rule. It is meeting a current one in a situation the current rule did not explicitly imagine.

Conclusion

The useful question about AI in validation documentation is not whether to allow it. That decision has already been made in most organizations, usually without a policy. The useful question is which documents in the set are transformations of something that already exists and which ones carry a judgment nobody has made yet. Test scripts, traceability matrices, configuration specifications, and summary report bodies are transformations, and AI drafting there is a genuine and defensible efficiency gain. Risk classifications, test scope, deviation dispositions, and fitness-for-use conclusions are judgments, and drafting them with AI produces a document that reads exactly like a decision without one having occurred. That is a harder failure to detect than a wrong answer, and it is what the Purolea letter is really about.

The one place that reasoning needs an extra layer is test scripts, because they look like the safe category and behave like the dangerous one. A generated script that traces cleanly to its requirement and passes on execution can be incapable of failing, and the research on machine-generated tests says this is the normal case rather than the edge case. Asking one question per high-risk script, what observed result would make this fail, catches it. Running a sample against a deliberately misconfigured system proves it, and doubles as the tool adequacy assessment Annex 11 already expects.

Sakara Digital works with pharma and biotech organizations writing this kind of scope into their validation procedures, and with quality groups trying to work out what changed about review now that the drafts arrive already written. If you are deciding where the line between AI drafting and human judgment belongs in your own validation SOP, and you want an independent read on it, we are happy to have that conversation.

For Further Reading