What the Regulations Actually Require

Two regulatory texts drive the annual review, and they are structured very differently. The US requirement is short and outcome focused. The EU requirement is long and input focused. Reading them together tells you almost everything you need to know about what to automate.

The US requirement: 21 CFR 211.180(e)

The US regulation is a single paragraph inside the general records requirements of Part 211. It states that written records shall be maintained so that the data in them can be used for evaluating, at least annually, the quality standards of each drug product, to determine the need for changes in drug product specifications or manufacturing or control procedures. It then requires written procedures for that evaluation, including a review of a representative number of batches, whether approved or rejected, and a review of complaints, recalls, returned or salvaged drug products, and investigations conducted under 211.192.1

Two things are worth noticing. First, the regulation names very few inputs. Batches, complaints, recalls, returns, and investigations. That is the whole list. Second, the operative verb is evaluating, and the operative purpose is to determine the need for changes. The US regulation does not ask for a document. It asks for a determination, and it asks that the determination be supported by records.

The EU requirement: Chapter 1, sections 1.10 and 1.11

EU GMP Part I Chapter 1 takes the opposite approach. Section 1.10 states that regular periodic or rolling quality reviews of all authorized medicinal products should be conducted with the objective of verifying the consistency of the existing process, verifying the appropriateness of current specifications for both starting materials and finished product, highlighting any trends, and identifying product and process improvements. Reviews should normally be conducted and documented annually, taking previous reviews into account, and should include at least twelve named elements.2

Rendered in plain terms, the twelve elements are: starting and packaging materials, with attention to new sources and to supply chain traceability of active substances; critical in-process controls and finished product results; all batches that failed to meet specification and their investigations; all significant deviations or non-conformances, their investigations, and the effectiveness of the resulting corrective and preventive actions; all changes to processes or analytical methods; marketing authorization variations submitted, granted, or refused, including export dossiers; the results of the stability monitoring program and any adverse trends; all quality related returns, complaints, and recalls, and the investigations performed at the time; the adequacy of any other previous product, process, or equipment corrective actions; post-marketing commitments for new authorizations and variations; the qualification status of relevant equipment and utilities such as HVAC, water, and compressed gases; and any contractual arrangements as defined in Chapter 7, to confirm they are up to date.

Section 1.11 is the sentence that matters most for this discussion. The manufacturer and, where different, the marketing authorization holder should evaluate the results of the review, and an assessment should be made as to whether corrective and preventive action or any revalidation should be undertaken under the pharmaceutical quality system. There should be management procedures for the ongoing management and review of those actions, and the effectiveness of those procedures should be verified during self-inspection. Section 1.11 also permits grouping reviews by product type where scientifically justified, and requires a technical agreement defining responsibilities where the marketing authorization holder is not the manufacturer.2

The structural point. Section 1.10 is a list of inputs. Section 1.11 is a decision. 21 CFR 211.180(e) collapses both into one sentence, but the same structure is there. Every automation decision in this article follows from the difference between the two: inputs can be assembled by a machine, and the decision cannot.

Active substances and ICH Q7

For active pharmaceutical ingredients, ICH Q7 section 2.5 sets a parallel expectation. Regular quality reviews of APIs should be conducted to verify the consistency of the process, and should normally be conducted and documented annually. The Q7 list is shorter than the EU list and covers critical in-process control and critical API test results, all batches that failed to meet specification, all critical deviations or non-conformances and their investigations, changes to processes or analytical methods, results of the stability monitoring program, and all quality related returns, complaints, and recalls.3 Q7 also states that the results of the review should be evaluated and an assessment made of whether corrective action or revalidation should be undertaken. The same structure again.

The practical questions the guidance already answers

The EMA good manufacturing practice questions and answers page resolves several of the scoping arguments that consume time in review programs.4 Grouping of products, sometimes described as bracketing or matrixing, may be appropriate where properly justified, and is usually only suited to situations where annual batch numbers are low and the grouped products share dosage form, the same or very similar active ingredients, and the same equipment. Where no manufacturing has occurred in the review period, the quality and regulatory review should still be conducted, covering stability results, returns, complaints, recalls, deviations including those arising from qualification and validation activities, and the regulatory background. Where a chain of contracts exists, all contracts in the chain are to be reviewed as part of the review process.

Settle these scoping questions in your procedure before you automate anything. A review tool built on an ambiguous scope definition will produce a fast, consistent, and wrong extract.

The requirement is enforced, and it is enforced against the quality unit

It is worth remembering that this is not a dormant regulation. FDA cites 211.180(e) regularly, and it usually appears inside a broader finding against the quality control unit under 21 CFR 211.22.22 In an October 2025 warning letter, FDA listed the failure to perform periodic, at least annual, product review among a set of quality unit oversight failures that also covered production and control record review and change control.20 In an April 2025 letter, the finding was framed as a failure to establish written records describing the evaluation of quality standards of each drug product at least annually, again within a broader quality unit deficiency.21 Analyses of FDA’s fiscal year 2025 output recorded 112 warning letters citing GMP deficiencies under 21 CFR 211, the highest count in more than two decades, with quality control unit responsibilities the most frequently cited section.19

The framing matters for anyone building a review program. A missing or inadequate review is not treated as a documentation lapse. It is treated as evidence that the quality unit is not exercising oversight, which is a much harder finding to close.

12 Elements EU GMP Chapter 1.10 requires in every Product Quality Review, at minimum
112 Warning letters citing 21 CFR 211 GMP deficiencies in FDA fiscal year 2025, the highest count in over two decades
211.22(c) The citation FDA used against a firm that let AI generate GMP documents without quality unit review

Where the Hours Go: An Illustrative Effort Model

The title of this article names a reduction from 200 hours to 40. That framing needs an honest caveat before anything else.

The 200 to 40 figures are an illustrative model, not a published benchmark. There is no independent, peer reviewed study establishing an industry average effort for a product quality review, and this article does not attribute the numbers to one. Software vendors publish reduction percentages for their own products. Those figures are marketing material, they are not independently verified, and none of them are cited here. What follows is a worked model with its assumptions stated openly, so that you can substitute your own numbers. The point of the model is the shape of the effort distribution, not the absolute totals.

Assumptions of the model

The model below assumes one commercial finished product, manufactured at one site, with roughly sixty batches disposed in the review period. Source data lives in eight systems: an ERP for batch and material records, a manufacturing execution system or electronic batch record system, a LIMS for QC results, a separate stability system or stability module, a quality management system for deviations and CAPAs, a change control module, a complaints and returns system, and a regulatory information management system for variations and commitments. Some of that data exists only in PDFs or spreadsheets. The review is compiled by one quality professional, with input requested from manufacturing, QC, validation, and regulatory colleagues. The hours below count all of that effort, not just the compiler’s time.

If your product is manufactured at three sites with two contract packers, the numbers go up. If you have a single validated data warehouse that already holds everything, they go down sharply. Substitute accordingly.

ActivityIllustrative hoursWhat drives the effortAutomation potential
Extract batch, yield, and disposition data30Reconciling bulk, sub-lot, and pack-stage records into a single batch viewHigh
Extract and reconcile QC results, including OOS and OOT25Multiple test methods, multiple specification versions, results held per sample rather than per batchHigh
Pull deviations and link them to product and batch25Free-text product references, deviations raised against equipment or area rather than productMedium to high
Pull CAPAs, status, and due dates12CAPAs linked to deviations rather than to product; open items spanning review periodsHigh
Pull change controls and assess relevance15Deciding which site or system changes actually touched this productMedium
Pull complaints, returns, and recalls12Complaints coded to trade name and market, not to manufacturing product codeMedium to high
Extract and tabulate stability data15Study codes, time points, multiple ongoing studies per productHigh
Confirm qualification status of equipment and utilities10Records held by engineering, often outside the quality systemMedium
Collect regulatory variations and post-marketing commitments10Regulatory system organized by dossier, not by manufacturing productMedium
Review contractual arrangements for currency8Agreements held in a contracts repository with no product linkageLow to medium
Build charts and tables18Manual charting, manual limit annotation, manual formattingHigh
Write narrative, interpret trends, draft conclusion, review and approve20Analysis and judgmentNone
Total200

What the model shows

In this distribution, roughly 162 hours sit in extraction, reconciliation, and presentation. Roughly 18 hours sit in charting. Roughly 20 hours sit in analysis and decision. That is the whole argument of this article in one table. The overwhelming majority of the effort in a product quality review is spent moving data, not thinking about it.

A well built assembly and trending capability removes most of the extraction, reconciliation, and charting effort. It does not remove all of it, because someone still has to verify that the assembled dataset is complete and correct before relying on it, and that verification is a real activity with real hours attached. In this model, verification of the assembled data plus the analysis and conclusion work lands around 40 hours.

The number that should not fall. If your post-automation figure comes out at ten hours rather than forty, the likely explanation is not that your tool is better. It is that the analysis and conclusion work has quietly disappeared, either because the tool is drafting the narrative or because the reviewer is signing an assembled document without interrogating it. The assembly hours should collapse. The judgment hours should hold, and in a mature program they often go up slightly, because reviewers finally have time to look at the borderline items properly.

Measure your own baseline before you build anything

Before starting a review automation project, run one full cycle with time recorded against the components above. Two hours of setup buys you a real baseline, a defensible business case, and a way to tell afterward whether the change worked. It also frequently changes the project scope. Programs that assume the pain is in report writing usually discover that report writing is a small fraction of the total, and that the real burden sits in two or three specific extractions that nobody had itemized before.

The Decomposition: Assembly, Trending, Judgment

Every component of the review falls into one of three classes. The classification is not about technical difficulty. It is about whether the activity has a single correct answer that can be derived deterministically from records that already exist.

CLASS 1

Automatable assembly

Retrieval, filtering, joining, and formatting of records that already exist and already carry an approved status. There is one correct output for a given period and product. Automate fully, then verify completeness rather than re-derive by hand.

CLASS 2

Semi-automatable trending

Computation against defined rules, presented for interpretation. The math automates. The rules must be written into the procedure in advance. The meaning of an exception stays with a person.

CLASS 3

Irreducible judgment

Adequacy assessments, effectiveness assessments, the overall evaluation, and the decision on corrective action or revalidation. These are the regulated output. No tool should draft them, and no tool should pre-populate them.

TEST

How to classify a component

Ask whether two competent reviewers with the same data would necessarily produce the same answer. If yes, it is assembly. If they would produce the same numbers but might reasonably differ on what the numbers mean, it is trending. If they might reasonably reach different conclusions, it is judgment.

The full component map

The table below maps each regulated component to its class, states what the tool should do, and states what stays with a person. It covers the twelve EU GMP elements plus the US-specific items.

Review componentBasisClassWhat the tool should doWhat stays with a person
Batches manufactured, yields, and disposition 211.180(e)(1); 1.10(ii) Assembly Produce the complete batch list with yields, disposition, and dates; flag any batch present in one system and absent from another Confirm the batch list matches expectation; explain any reconciliation gap
Critical in-process controls and finished product results 1.10(ii); Q7 2.5 Assembly, then trending Assemble results per batch per attribute; compute distribution statistics, capability, and comparison to the prior period Interpret drift, decide whether variation is acceptable
Batches failing specification, and their investigations 211.180(e)(1); 1.10(iii) Assembly List every failing batch, link to the investigation record, show closure status and root cause code Assess whether the investigations were adequate and whether failures share a cause
Out of specification and out of trend laboratory results 211.192; 1.10(ii) Assembly, then trending List all OOS and OOT events with method, stage, and outcome; count by method and by period Judge whether laboratory performance is a contributing factor
Significant deviations and non-conformances 211.192; 1.10(iv) Assembly, then trending List deviations linked to the product; count by category, by severity, by root cause code, by period Decide what recurrence means and whether the categorization itself is credible
Effectiveness of resulting corrective and preventive actions 1.10(iv) Judgment Surface CAPA records, effectiveness check dates, and post-CAPA recurrence counts State whether the actions were effective, with reasoning
Adequacy of previous product, process, or equipment corrective actions 1.10(ix) Judgment List prior-period actions and their current status Assess adequacy. The regulation uses the word adequacy, which is an evaluative term
Changes to processes or analytical methods 1.10(v) Assembly, with human scoping List change controls with an explicit product or equipment linkage Decide which site-level or system-level changes actually affected this product
Starting and packaging materials, new sources, supply chain traceability 1.10(i) Assembly, with human scoping List materials used, suppliers, lots, and any newly approved sources in the period Assess supply chain traceability of active substances and the significance of new sources
Marketing authorization variations submitted, granted, or refused 1.10(vi) Assembly List variations by market with status and date Confirm the manufacturing reality matches the approved dossier
Post-marketing commitments 1.10(x) Assembly List open commitments with due dates and status Confirm commitments are being met and escalate those that are not
Stability monitoring results and adverse trends 1.10(vii); Q7 2.5 Assembly, then trending Assemble results by study, time point, and attribute; fit and plot trends against shelf-life specification Judge whether any trend threatens the approved shelf life
Complaints, returns, and recalls 211.180(e)(2); 1.10(viii) Assembly, then trending List and categorize; normalize counts against units distributed; compare to prior periods Interpret category shifts and decide whether a complaint pattern is a quality signal
Qualification status of equipment and utilities 1.10(xi) Assembly Report current qualification status and next due dates for relevant equipment, HVAC, water, and gases Assess whether any lapse affected the product in the period
Contractual arrangements 1.10(xii) Mixed List agreements, parties, effective dates, and review dates; flag expired or overdue agreements Confirm the scope of each agreement still matches what the parties actually do
Overall evaluation and conclusion 1.11; 211.180(e) Judgment Nothing. Present the assembled evidence and leave the field empty State whether the process remains in a state of control, whether specifications remain appropriate, and whether CAPA or revalidation is required

The components people misclassify

Three rows in that table are routinely put in the wrong class, and each mistake creates a specific problem.

CAPA effectiveness gets treated as assembly. A tool can report that an effectiveness check was completed and closed as effective. It cannot report that the action was effective, because the effectiveness check itself may have been superficial, and because a recurrence in a different form may not have been counted. The review is a second look, not a repeat of the first look. If your review section on CAPA effectiveness simply restates the closure status recorded in the quality system, you have not performed the assessment 1.10(iv) asks for.

Change control relevance gets treated as fully automatable. Change records are usually linked to equipment, systems, facilities, or documents rather than to finished products. A rule that pulls every change touching a piece of equipment used by the product will return a long list containing many changes with no product impact, and will miss changes to a shared utility or a shared analytical method that the rule did not know to look at. Automate the retrieval, then keep a human scoping step, and record the scoping rationale so the next reviewer can follow it.

Contractual arrangements get skipped. Element (xii) is easy to satisfy in form and hard to satisfy in substance. Reporting that an agreement exists and has not expired is assembly. Confirming that the agreement still describes what the parties actually do is judgment, and it is the check that finds the real problems: a testing scope that moved, a site that was added, a responsibility that was assumed informally and never documented.

The Conclusion Is the Product, and It Cannot Be Generated

Everything above builds to this section. The review exists to produce four statements: whether the process remained in a state of control across the period, whether current specifications remain appropriate for starting materials and finished product, whether corrective and preventive action is required, and whether revalidation is required. EU GMP 1.11 asks for exactly this. 21 CFR 211.180(e) asks for the determination of the need for changes in specifications, manufacturing, or control procedures. ICH Q7 2.5 asks for the same evaluation for active substances.

Those statements are regulatory acts performed by named, accountable people. They are not summaries of the data above them. Two reviewers looking at the same assembled dataset can reasonably reach different conclusions, and the reasoning is the evidence that the pharmaceutical quality system is functioning.

What “state of control” actually means

The phrase is defined in FDA’s process validation guidance, which describes a state of control as a condition in which the set of controls consistently provides assurance of continued process performance and product quality. Stage 3 of the lifecycle, continued process verification, exists to give ongoing assurance that the process remains in that condition during routine production.5 The annual review is one of the places where that assurance is recorded and evaluated. It is not a filing exercise sitting alongside continued process verification. It is a periodic, documented judgment about the same question.

Annex 15 makes the linkage explicit on the EU side: process validation is not a one-time event, and periodic evaluation should confirm that processes and procedures remain in a state of control, with the review acting as one input to the decision on whether revalidation is warranted.6

The failure mode FDA has already described

In April 2026, FDA issued a warning letter to a drug manufacturer describing the inappropriate use of artificial intelligence in pharmaceutical manufacturing. The firm told investigators it had used AI agents to help comply with FDA regulations, specifically to create drug product specifications, procedures, and master production or control records. FDA’s response was direct: if AI is used as an aid in document creation, the AI generated documents must be reviewed to confirm they are accurate and actually compliant with CGMP, and the failure to do so is a violation of 21 CFR 211.22(c). The letter also recorded that the firm had not conducted process validation before distribution, and that the firm’s stated reason was that the AI agent it used had never told it the requirement existed.78

The relevance to review automation is precise, and it is worth stating carefully. That warning letter was issued to a small manufacturer with multiple basic CGMP failures, and it should not be read as a general regulatory position on AI in pharma. What it does establish is the principle FDA applied: a generated GMP document that the quality unit did not genuinely review is a quality unit failure, regardless of how the document was produced. A review tool that drafts the conclusion creates exactly that situation. The reviewer stops writing an assessment and starts editing a draft, and editing a plausible draft is not the same activity as forming a judgment.

Design rules that keep the line in place

The separation has to be built into the tool, not left to policy. Four design rules do most of the work.

  • The conclusion field starts empty and stays empty until a person types in it. No suggested text, no template sentence, no “no adverse trends were observed” pre-filled and waiting to be accepted.
  • The tool may state facts and must not state assessments. “Fourteen deviations were recorded, of which three carry the same root cause code as two deviations in the prior period” is a fact and belongs in the assembled output. “The process remains in a state of control” is an assessment and does not.
  • Signals are surfaced, not resolved. Where a defined trend rule fires, the tool should say which rule fired, on which data, and leave the interpretation to the reviewer. A flag that says “review required” is useful. A flag that says “not significant” is the tool making the call.
  • The reviewer records what they examined. Capture which exceptions were opened, which source records were inspected, and what the reviewer concluded about each. That record is the evidence that a human review happened, and it is what an inspector will ask to see when the report looks machine generated.

If generative text is used anywhere near the review, treat it as a separate change with its own risk assessment. The draft EU GMP Annex 22 on artificial intelligence, published for consultation in July 2025, is scoped to static, deterministic models for applications that directly affect product quality or patient safety, and excludes models that adapt during use from GMP-critical applications.9 A large language model drafting narrative text in a GMP document is not obviously inside that scope, and building your program on the assumption that it is would be optimistic. Human-in-the-loop is not a mitigation you can claim. It is a control you have to design, document, and be able to demonstrate.

The Real Difficulty Is Joining the Data

Teams that have not attempted this work expect the hard part to be the report. Teams that have attempted it know the hard part is that batch data, deviations, complaints, and stability results live in systems that do not share a product identifier or a batch identifier, and that no amount of report design fixes that.

The technical literature on manufacturing system integration names the problem directly. Semantic and master data mismatch, where two systems use slightly different names for materials, batches, and units, is the most common integration failure, and it means a batch in the LIMS may not map cleanly to the same batch in the manufacturing execution system. Where QC results live in one system and manufacturing execution data in another with only a thin interface between them, no single system holds the complete record of a batch.10

The four identifiers that break

Product identifier. The ERP holds a material number. The manufacturing execution system holds a recipe or production version. The LIMS holds a product or specification code. The quality management system frequently holds a free-text product name typed by whoever raised the record. The complaints system holds a trade name, which varies by market. The regulatory system is organized by dossier and procedure number. Six identifiers, one product, and no authoritative mapping in most organizations.

Batch identifier. Lot numbering conventions differ by site and sometimes by product. Bulk lots, sub-lots, filled lots, and packed lots may carry different numbers, with a genealogy relationship that exists in one system and not in others. Reworked and repacked lots create additional identifiers pointing at the same material. Yield reconciliation across bulk and pack stages depends entirely on getting this genealogy right.

Reference data and taxonomies. Deviation categories, root cause taxonomies, CAPA classifications, and complaint codes are reference data, and reference data changes. If the deviation categorization scheme was revised in month seven of the review period, a year-over-year count by category is comparing two different taxonomies and the comparison is meaningless unless the change is stated. Record the taxonomy version in force during the period, and state any mid-period change explicitly in the review.

Dates and period boundaries. Manufacturing date, release date, disposition date, and distribution date are all candidates for defining whether a batch falls inside the review period, and different systems default to different ones. The procedure must state which date governs, and every extract must implement the same rule. Two extracts using different date fields will produce two different batch counts, and the discrepancy will surface in an inspection rather than in the build.

The single most useful artifact: an unmatched records report. Every automated review should produce, alongside the review itself, a report of records that could not be mapped to a product or batch. Deviations with no resolvable product link. Complaints whose trade name mapped to nothing. QC results for a batch not present in the batch list. That count is a data quality metric in its own right, it belongs in the review, and it is the honest answer to the question an inspector will eventually ask: how do you know the report is complete? A review that reports zero unmatched records and has never reported a non-zero figure is usually one where the unmatched records are being silently dropped.

Treat the mapping as governed master data, not as a spreadsheet

The cross-reference between product identifiers and between batch identifiers is master data. It needs a named owner, a change process, version history, and a periodic review of its own. In most organizations it starts life as a spreadsheet maintained by one experienced person, which works until that person changes role. Moving it into a governed master data object, with the same controls applied to any other master data domain, is usually the single highest-value step in a review automation program, and it delivers benefits well beyond the annual review.

This is also where the business case gets easier. Product and batch identifier resolution is not an annual review problem. It is the same problem that blocks continued process verification dashboards, deviation trending across sites, supply chain traceability, and any attempt to apply analytics to manufacturing quality data. Funding it as review automation alone understates the return.

Most reviews present data. Comparatively few analyze it. Section 1.10 asks the review to highlight any trends, and section 1.10(vii) asks specifically about adverse trends in stability. Inspectors increasingly ask a harder question than whether a chart is present: what does the trend mean, and what did you do about it.

The MHRA’s published inspection deficiency data has consistently shown Chapter 1, Quality Management, among the most frequently cited areas of the EU GMP guide, alongside the sterile products annex.1112 Published analyses of that data set out the pattern in more detail, and a recurring theme is not the absence of data but the absence of evaluation.13

What genuine trending contains

A trend section that stands up to questioning contains all of the following, and most weak trend sections are missing at least three of them.

  • Comparison to prior periods. Not just this year’s numbers. This year against last year and, where available, the two years before that. Section 1.10 explicitly requires the review to take previous reviews into account.
  • Comparison to specification limits and to control limits. These are different comparisons answering different questions. Specification limits tell you whether product met requirements. Control limits, derived from process performance, tell you whether the process behaved as it normally does. A batch inside specification but outside control limits is a signal, and a chart showing only specification limits will hide it.
  • A distinction between common cause and special cause variation. FDA’s process validation guidance states that trending should be performed in a way that guards against overreaction to individual events, and that data should be statistically trended and reviewed by a statistician or an appropriately trained person.5 Reacting to every point is as much a failure as reacting to none.
  • Process capability where it applies, with its assumptions checked. For numeric attributes with meaningful specification limits, a capability measure states whether a stable process can consistently meet requirements. Capability indices assume an underlying normal distribution, so normality should be tested before the index is quoted, and a capability index describes a future single unit rather than the probability that a future batch will meet specification.14 That distinction matters when a reviewer is asked what a capability figure actually predicts.
  • Denominators on every count. Fourteen deviations across sixty batches is a different picture from fourteen deviations across two hundred batches. Complaint counts should be normalized against units distributed. Counts without denominators are not trends.
  • A stated definition of what would constitute an adverse trend. If the procedure does not define the trend rules in advance, then “no adverse trend was observed” is unfalsifiable. Define the rules, write them into the procedure, implement them in the tool, and let the reviewer interpret the exceptions the rules produce.

Five questions every trend section should answer.

  • What changed compared with the prior period, in numbers?
  • Is the change within the normal variation of this process, or outside it?
  • Is the process capable of meeting its specification, and has capability moved?
  • Is any movement directional, meaning a drift rather than scatter?
  • What action follows, and if none, on what basis?

The trending failures that show up in inspections

Four patterns account for most weak trend sections. Charts drawn without limits, so the reader cannot see whether any point matters. Counts presented without denominators, so year-over-year comparisons are not comparable. A conclusion of “no adverse trend” with no stated definition of an adverse trend. And taxonomies that changed mid-period without disclosure, so the comparison is between two different categorization schemes.

All four are fixable in the procedure rather than in the tool, and all four are worth fixing before automation rather than after. A tool built on undefined trend rules will produce charts faster and answer no more questions than the manual version did.

Connect the review to continued process verification

Where a continued process verification program exists, the annual review should draw on it rather than duplicate it. Continued process verification is an ongoing program of collecting and evaluating process and product data during routine commercial production, and it addresses the same question of whether the process remains in a state of control.15 ISPE’s guidance on the process performance and product quality monitoring system, one of the four elements of a pharmaceutical quality system under ICH Q10, describes the monitoring system that both activities depend on.1617

Running the two independently is a common source of duplicated effort and, occasionally, of contradictory statements in two documents about the same process in the same period. Where both exist, define which one owns the statistical analysis, and have the other reference it.

Validating the Review Tool Itself

A system that extracts regulated data from validated source systems, transforms it, and assembles it into a GMP document that will be presented to inspectors is in scope for computerized system validation. This is not a controversial position, but it is regularly overlooked, particularly where the review is built by a data team using business intelligence tooling rather than by an IT function familiar with GxP.

The proportionate answer is not a heavy validation package. It is a validation effort aimed at the specific risk the tool creates.

Identify the actual risk

The risk is not that the reporting tool crashes. A crash is visible and someone fixes it. The risk is that the tool silently returns an incomplete or incorrect dataset that looks perfectly reasonable, and that a reviewer forms a judgment on it and signs. A missing join that drops eight percent of deviations produces a review showing improved deviation performance. Nothing in the output announces the problem.

GAMP 5 second edition frames exactly this kind of decision. It keeps the risk-based framework and the software categories, and puts more weight on critical thinking by knowledgeable subject matter experts to define an appropriate approach, rather than applying a uniform document set to every system.18 Applied here, critical thinking means concentrating the testing effort on data completeness and correctness, and spending very little on the parts of the tool that are configuration of a supported product.

What to test

  • Extraction completeness. Reconcile record counts against the source system for a defined period. Not a sample. A count.
  • Filter correctness. Period boundaries, product scope, status filters, and site scope. Test the boundaries specifically, including a batch disposed on the first and last day of the period.
  • Join correctness. Confirm that joins neither duplicate nor drop records. A one-to-many join that silently multiplies deviation counts is a common and undetected failure.
  • Calculation accuracy. Verify computed statistics, limits, and capability measures against an independent calculation for a defined sample.
  • Null and unmapped handling. Confirm that records which cannot be mapped are reported rather than dropped, and that null values are not silently treated as zero.
  • Access control and report definition change. Establish who can change the query or report logic, and confirm those changes are controlled. A report whose logic anyone can edit is not a controlled GMP output.
  • Audit trail. Record who generated the report, when, against which data, and with which version of the report logic.
  • Reproducibility. Regenerate a prior period’s review from the same logic and confirm it matches the approved version, or explain every difference. This test is skipped more often than any other and it is the one that most efficiently demonstrates control.

The change control point that programs miss after go-live. An upgrade to a source system, a change to a deviation taxonomy, a revision to a specification, or a new site added to the product all change what the report returns. Add the review tool to the impact assessment for changes to every source system it draws from, and re-run the completeness reconciliation after each one. The most common post-implementation failure in review automation is not a defect in the tool. It is a source system change that nobody assessed against the report.

What not to validate

The validated scope covers assembly, computation, and presentation. It does not cover the reviewer’s conclusion, because the conclusion is not produced by the system. Attempting to validate the narrative confuses the boundary this whole article is about. The reviewer’s judgment is controlled by procedure, by training and qualification, by defined approval authority, and by the record of what was examined. Those are quality system controls, not software controls, and they should be documented as such.

A Staged Path to Getting There

Review automation programs fail in predictable ways: they start with the report, they discover the identifier problem in month four, and they either stall or deliver a tool that produces a fast version of an unclear review. The sequence below front-loads the work that everything else depends on.

1

Fix the procedure before you fix the tooling

Define the review period and the date field that governs inclusion. Define product scope and any grouping justification. Define the trend rules and what constitutes an adverse trend. Define what the conclusion must state. Automating an ambiguous procedure produces a faster ambiguous review, and the ambiguity becomes harder to see once it is embedded in code.

2

Measure your own baseline for one full cycle

Record hours by component through one complete review. The result gives you a defensible business case, a way to prioritize, and a way to know afterward whether anything improved. Most teams find the distribution surprises them.

3

Build the identifier map as governed master data

Product cross-reference across ERP, MES, LIMS, QMS, complaints, and regulatory systems. Batch genealogy across bulk, filled, and packed stages. Named owner, change control, version history. Build the unmatched records report at the same time and treat its output as a metric, not as an error log.

4

Automate assembly for the highest-effort components first

In most organizations that means batch and yield data and QC results, which together account for a large share of the extraction burden. Deliver those, verify completeness against source counts, and let the reviewers work with the output for a cycle before extending.

5

Add computed trending with the rules you already defined

Every chart carries its limits. Every count carries its denominator. Every trend rule that fires says which rule fired and on what data. Nothing in the trending layer states whether a signal matters.

6

Validate proportionately and tie the tool to source system change control

Concentrate testing on completeness, filters, joins, calculations, and reproducibility. Add the tool to the impact assessment for every source system it reads from, and reconcile counts after each source system change.

7

Run in parallel for one cycle, then retire the manual process

Produce the review both ways for one period and reconcile every difference. The differences you find are the defects you would otherwise have discovered during an inspection, and the parallel run itself is useful validation evidence.

8

Leave the conclusion alone, permanently

Empty field, human author, recorded reasoning, recorded evidence of what was examined. This is the one design decision that should survive every future release of the tool.

What changes for the reviewer

The point of the change is not that the review takes less time. The point is that the reviewer’s time moves from collection to analysis. In the illustrative model earlier in this article, a reviewer spends roughly ten percent of the total effort on analysis and judgment. After the change, that share goes from ten percent of two hundred hours to roughly half of forty. The absolute hours spent thinking stay similar or rise slightly. Everything that changes is the ratio.

That reframing also changes how you should describe the project internally. A program presented as a headcount reduction will be judged on whether hours fell. A program presented as moving the quality unit’s effort from assembly to evaluation will be judged on whether the reviews got better, which is both the honest measure and the one that matters in an inspection.

Conclusion

The product quality review has an unusual property among GMP documents: the regulation is very clear about what goes into it and comparatively brief about what comes out of it, and the part that comes out is the part that matters. EU GMP 1.11 asks for an evaluation and an assessment of whether corrective action or revalidation is needed. 21 CFR 211.180(e) asks for a determination of the need for changes. Everything else in the document exists to support those statements. When a quality unit spends four fifths of its review effort assembling inputs, the statements at the end get the leftover attention, and that is the real problem automation should be solving.

So the design question is not how much of the review can be automated. It is which parts, and the answer follows the structure of the regulation itself. Assemble automatically and verify completeness. Compute trends against rules you defined in advance and present them with their limits and denominators. Leave the conclusion, the state of control assessment, the adequacy and effectiveness judgments, and the revalidation decision to accountable people, with the evidence of their examination recorded. A tool that drafts the conclusion has not saved the reviewer time. It has moved the reviewer from authoring to editing, and editing a plausible draft is a different and weaker activity than forming a judgment. That is the failure FDA described in its first AI-related CGMP warning letter, and it is entirely avoidable at design time.

Sakara Digital works with pharma and biotech organizations building this kind of automation into their quality systems, with particular attention to the master data and reference data work that determines whether any of it holds together. If you are looking at product quality review automation and want an independent view on where the line between assembly and judgment should sit, we are happy to have that conversation.

For Further Reading