Why Small, Bounded Pilots Work Better Than a Data Quality Launch

A familiar pattern plays out at pharma and biotech manufacturing sites. Someone runs an assessment and finds data quality issues in the batch records, the equipment lists, the warehouse, and the deviation system. A program is proposed to address all of them. It needs a steering committee, a data governance charter, a tool selection, and a budget cycle. Twelve months later the site has a governance framework on paper and the same batch record corrections it had before.

The problem is not the ambition. It is that “data quality” at a manufacturing site is not one problem. It is dozens of separate problems, each with its own system of record, its own owner, and its own cause. A missing equipment ID in a batch record and a deviation coded to the wrong root cause category have almost nothing in common, except that both make it harder to release product and harder to see trends. Treating them as one program means nobody owns either.

A pilot flips this around. It takes one data set, in one area, with one owner, and asks a narrow question: how good is this data today, what makes it bad, and can we measurably improve it in thirteen weeks? The answer is useful even if it is “no.” A pilot that finds the data is already in good shape has saved the site from spending a year fixing something that was not broken.

What “Pay Back in One Quarter” Means Here

We want to be precise about the promise in this article’s title. We have not found a public benchmark showing that a particular data quality pilot returns its investment within a quarter, and we would be suspicious of anyone who quotes one. What we mean is narrower and more useful: each pilot below is designed so that the return can be counted inside the quarter, using the site’s own records. That return usually takes one of three forms:

  • Hours given back. Review time, reconciliation time, and investigation time that no longer have to be spent because the data is right the first time.
  • Events avoided. A known type of correction, discrepancy, or deviation that happened at a measured rate before the pilot and at a lower rate after.
  • Decisions made possible. A trend report, a review-by-exception approach, or an impact assessment that could not be trusted before and can be now.

The first two can be converted into money if your finance team wants that. The third is often the most valuable and the hardest to price. We cover how to count all three honestly in the section on counting the payback.

What This Article Does Not Repeat

Sakara Digital has already published detailed pieces on several manufacturing data sets. We link to them rather than restate them. The electronic batch record fields article covers which fields cause investigations and how to redesign them. The environmental monitoring article covers sampling location master data and trend programs. The document management pilots article covers SOP metadata, document ownership, master batch record naming, and quality agreement metadata. This article is about the pilot itself: how to scope it, who owns it, what to count, and what a good result looks like, across seven data sets those pieces do not treat as pilots.

Why environmental monitoring is not on the list of seven. It is a strong candidate, and many sites should start there. We left it off only because our environmental monitoring article already lays out the fixes. If EM is your biggest gap, the location hierarchy element from that article scopes well to one grade area for one quarter, and the four tests and thirteen-week plan below apply to it without change.

The Four Tests Every Pilot Must Pass

Before any pilot starts, it should pass four tests. These are not bureaucratic hurdles. Each one exists because pilots that skip it tend to fail in a predictable way.

Test 1

One Named Data Owner

A single person who has the authority to change how the data is created, and who will still be accountable for it after the pilot ends. Not a committee. Not “IT and Quality jointly.”

Test 2

One System of Record per Data Element

For each field being measured, the team agrees which system or record is the reference. If three systems disagree, the pilot measures the disagreement against that reference.

Test 3

A Measure Countable From Existing Records

The baseline must come from records that already exist, so it can be taken in the first two weeks without new data collection. If you have to build a tool to get a baseline, the scope is too big.

Test 4

A Result Visible Within Thirteen Weeks

The data must be created often enough that a change shows up in the counts within the quarter. A record created twice a year cannot show improvement in a quarter.

Naming the Data Owner

The first test is the one most often fudged. The MHRA GXP data integrity guidance (Revision 1, March 2018) says data governance should address data ownership and accountability throughout the lifecycle, and should consider the design, operation and monitoring of processes and systems.1 In practice, the data owner for a manufacturing data set is usually the person who runs the process that creates the data: the production manager for batch record entries, the maintenance or engineering lead for equipment records, the QA lead for release status.

This matters for a pilot because most fixes change how people create data: a redesigned form, a new rule at data entry, a revised category definition. The system administrator cannot make those changes stick. The process owner can. When a pilot is owned by the data team or by IT alone, it tends to produce a clean report and no lasting change.

Each pilot also needs a measure owner, who may be a different person. The measure owner is responsible for pulling the counts the same way every time. On a small site this can be one quality engineer running two pilots. What matters is that the definition of the measure does not change mid-quarter, which is the fastest way to produce a result that looks good and means nothing.

Three Words for What You Are Measuring

Data quality has too many competing vocabularies. For manufacturing pilots we find it helps to borrow the three categories from a harmonized framework developed for health record data by Kahn and colleagues: conformance (do values follow the expected format, allowed values, and relationships?), completeness (is the data present where it should be?), and plausibility (are the values believable given everything else we know?).2 The same paper separates verification, meaning checks you can run against your own organization’s rules, from validation against an external reference.

These three words are enough to describe every measure in this article. An equipment ID that does not match the format standard is a conformance failure. A cleaning record with no end time is a completeness failure. A clean hold time that is negative, because the next use was recorded before cleaning finished, is a plausibility failure. Using the same three words across pilots also makes it easier to report them together to site leadership.

Why Scope Small, Even When the Problem Is Big

Every pilot below is scoped to one product, one suite, one warehouse zone, or one system. That can feel too small when the site leader knows the problem exists everywhere. There are three reasons to hold the line. First, a small scope makes the baseline cheap to take. Second, it keeps the number of people who need to change their behavior small enough that the pilot team can talk to all of them. Third, the fix you design for one suite is usually the fix for every suite; the pilot proves it before you roll it out. Expanding a proven fix is quick. Expanding an unproven one is how a pilot turns into the stalled program it was meant to replace.

The Quarter Plan: Thirteen Weeks in Four Stages

All seven pilots follow the same shape. The timings are a guide, not a rule, but the order matters. In particular, the baseline has to be taken before anyone starts fixing anything.

1

Weeks 1–2: Charter and Baseline

Write a one-page charter: scope, data owner, measure owner, system of record, the exact definition of each measure, and the target. Then pull the baseline from existing records, usually the last one to three months of data. Freeze the measure definitions once the baseline is taken.

2

Weeks 3–4: Find the Causes

Sort the failures by type and count them. In most pilots a handful of causes account for most of the failures. Talk to the people who create the data. Decide which fixes are in scope, and start any change control that the fixes need now, because approvals take time.

3

Weeks 5–10: Fix and Re-measure

Put the fixes in place and re-measure every one or two weeks using the frozen definitions. Prefer fixes that change the point of data creation (a form, a rule, a definition) over fixes that clean data after the fact. Cleanup without a creation fix returns to baseline.

4

Weeks 11–13: Hold and Decide

Keep measuring without adding new fixes, to see whether the improvement holds. Then write a short decision paper: expand to the next area, hold the measure as a routine metric, or stop. Present it to site leadership with the counts, not just the conclusion.

Change control is part of the plan, not an obstacle to it. Several fixes below change GMP records or validated system configuration: a batch record template, a cleaning log form, a monitoring system alarm setpoint. Those changes go through your change control process. Plan for this in weeks 3 and 4. If a fix needs a validated system change that cannot be approved and implemented in the quarter, pick a fix outside the validated boundary for this pilot (a procedure, a training point, a reference list) and log the system change as a follow-on.

Pilots One to Three: Records and Reference Data

The first three pilots deal with data that is written down once and read many times: batch record entries, equipment identifiers, and material status. Errors in this kind of data spread, because every downstream record copies them.

Pilot One: Batch Record Review Corrections for One Product

Scope. Executed batch records for one high-volume product, paper or electronic, for the quarter. The pilot counts every correction request raised during production review and QA review: missing entries, entries in the wrong format, calculation errors, missing second-person checks, and late entries that need a comment.

Data owner. The production manager responsible for the product. The QA batch record review lead is the natural measure owner, because QA reviewers already see every correction.

Why it matters. Under 21 CFR 211.192, the quality control unit must review and approve all production and control records before a batch is released or distributed, and any unexplained discrepancy must be thoroughly investigated.3 Section 211.188 lists what the batch record must document, including dates, the identity of major equipment used, component batches, weights and measures, in-process results, and the identity of the people performing and checking each significant step.4 Every correction is a return trip: the reviewer flags it, the record goes back to production, someone corrects it with a dated comment, and the reviewer checks it again. The product waits the whole time.

Measures.

  • Corrections per executed batch record, split by field or field type (completeness and conformance).
  • The share of all corrections that come from the top five fields. In our experience this share is high, which is what makes the pilot work, but measure it at your site rather than assume it.
  • Review cycle time: elapsed time from batch completion to QA disposition, in working days.

The baseline. Pull the last 20 to 30 executed records for the product and count corrections from the review comments or the correction log. If the site does not keep a correction log, the corrections are still visible in the records themselves as dated amendments.

Typical fixes. Redesign the top five fields only: pre-print values that never change, add units next to entry boxes, remove fields that duplicate an earlier entry, reword instructions that operators read two ways, and move second-person checks to the point where the checker is physically present. For an EBR, many of these are configuration changes; for paper, they are template changes. Both need change control. The EBR fields article goes into the design of specific field types in more depth.

What a good result looks like. Corrections from the top five fields fall sharply by weeks 9 and 10 and stay down through week 13, no new correction type appears in their place, and review cycle time for the product falls. The team can explain each remaining correction type and has a view on whether it is worth fixing next.

What to watch for. Do not let the pilot turn into an operator retraining exercise. If the fix is “retrain everyone on good documentation practice,” the corrections usually return within weeks. The goal is to change the record so the right entry is the easy entry.

Pilot Two: Equipment Identity Across Systems for One Suite

Scope. All major equipment in one manufacturing suite. The list should be short enough that someone can walk the room with it in an afternoon.

Data owner. The engineering or maintenance lead who owns the equipment master in the maintenance management system (CMMS). Other system owners (calibration, MES or batch record master data, logbooks) take part, but one person owns the equipment identification standard.

Why it matters. 21 CFR 211.105(b) requires major equipment to be identified by a distinctive identification number or code, recorded in the batch production record to show the specific equipment used for each batch.5 At most sites the same piece of equipment appears in several places: the asset tag on the equipment, the CMMS record, the calibration system record for its instruments, the equipment list in the MES or master batch record, and the equipment logbook. When these disagree, every investigation that asks “which equipment was used, and was it in a qualified and calibrated state?” starts with a manual reconstruction.

Measures.

  • Identity match rate: the share of equipment items where the ID on the physical tag, the CMMS, the calibration system, the MES or master batch record, and the logbook all agree (relational conformance).
  • Status match rate: the share where the operational status (in service, out of service, retired) agrees across the same systems (plausibility).
  • Orphans: equipment present in one system and absent from another, counted in both directions.

The baseline. Export the equipment lists for the suite from each system, then walk the room and record the physical tags. This walk is the step teams most want to skip and the one that finds the most surprises, such as replaced equipment still carrying the old tag or equipment moved between suites without a system update.

Typical fixes. Correct the mismatches under change control, then fix the process that created them. The most common cause we see is the absence of a rule saying which system an equipment ID is created in first. A simple rule (the ID is created in the CMMS, and no other system may create one) plus a check in the equipment introduction procedure prevents most new mismatches.

What a good result looks like. Full identity and status agreement for the suite by the end of the quarter, a written rule for where equipment IDs are created, and no new mismatches during the hold period. The suite is then a working model for the rest of the site.

This pilot pays back in two places. It shortens investigations directly, and it is a precondition for other pilots, especially Pilot Four, where calibration results must be tied to the batches that used the equipment. It is also one of the prerequisites for review by exception for batch records, which relies on the system knowing which equipment was used and in what state.

Pilot Three: Material Status Agreement in One Warehouse Zone

Scope. One storage zone or one class of material, for example incoming raw materials held in quarantine and released areas.

Data owner. The QA lead responsible for material disposition. Status is a quality decision, so QA owns it, even though the warehouse manager owns the physical location and the ERP or warehouse system key user owns the system record.

Why it matters. 21 CFR 211.82(b) requires components, containers, and closures to be stored under quarantine until tested or examined and released, and 211.89 requires rejected materials to be identified and controlled under a quarantine system that prevents their use.67 Chapter 5 of the EU GMP Guide says incoming materials should be physically or administratively quarantined until released, lists status (in quarantine, on test, released, rejected) as label information where appropriate, and notes that when fully computerized storage systems are used, that information need not all be legible on the label.8 That last point is why this pilot matters at sites with a warehouse management system: the system status is the control, so a mismatch between system status and physical reality is a control failure, not a clerical one.

FDA’s August 2026 warning letter to Safrel Pharmaceuticals shows how an inspector approaches this. Among other findings, FDA observed boxes of finished product labeled “Unlabeled, No Lot, No Exp” stored in the same general area as labeled product, and asked the firm for a current inventory including expiration dates and status (for example, released, quarantined, rejected).9 A site that cannot produce that list quickly and accurately has a material status data problem, whatever its procedures say.

Measures.

  • Status agreement rate: from a weekly blind sample of lots, the share where the system status, the physical location or label, and the QA disposition record all agree (relational conformance).
  • Critical mismatches: any lot shown as available in the system but rejected or still in quarantine in the QA record. This should be counted separately and treated as a deviation when found.
  • Expiry and retest date exceptions: lots past their retest or expiry date still showing as available (plausibility).
  • Quarantine age: lots sitting in quarantine longer than the site’s target time to disposition, which often points to paperwork that has stalled rather than a testing problem.

Typical fixes. The common causes are manual status updates done in the wrong order (the pallet moves before the system is updated, or the reverse), status changes made by people outside QA, and partial lots split across locations. Fixes include a system rule that only QA roles can change a status to released or rejected, a daily exception report of lots whose retest date has passed, and a single documented sequence for moving released material.

What a good result looks like. Zero critical mismatches found in the sample during the hold period, non-critical mismatches trending down, no lots past retest or expiry shown as available, and a status list the site could hand to an inspector the same day.

Pilots Four and Five: Equipment Condition Data

The next two pilots deal with data that tells you whether equipment was fit for use at the moment it was used. This data matters most when something goes wrong, which is why its gaps tend to be discovered during an investigation rather than before one.

Pilot Four: Calibration Out-of-Tolerance Records for One Instrument Class

Scope. All as-found out-of-tolerance (OOT) calibration results for one class of instrument (for example, temperature and pressure instruments in one building) for the past twelve months, plus any new ones during the quarter.

Data owner. The calibration or metrology lead owns the calibration record. QA owns the product impact assessment decision. The pilot needs both, and it needs the output of Pilot Two if the equipment identities are unreliable.

Why it matters. 21 CFR 211.68(a) requires automatic, mechanical, and electronic equipment to be routinely calibrated, inspected, or checked according to a written program, with written records kept. Chapter 3 of the EU GMP Guide says measuring, weighing, recording and control equipment should be calibrated and checked at defined intervals, with adequate records of those tests.10 For laboratory instruments, 211.160(b)(4) requires a calibration program with specific directions, schedules, limits for accuracy and precision, and provisions for remedial action when limits are not met, and says instruments not meeting specifications shall not be used.11 ISPE’s GAMP Good Practice Guide on calibration management describes a risk-based approach covering instrument risk assessment, program management, documentation, and corrective actions.12

The data quality problem is not usually the calibration itself. It is what happens after an OOT result. The instrument is adjusted and returned to service, and the question “which batches were made while it was reading wrong, and did it matter?” is answered slowly, partly, or from memory. That answer depends on three links in the data: the instrument to the equipment it serves, the equipment to the batches that used it, and the time window since the last good calibration.

Measures.

  • Share of OOT events with a completed product impact assessment (completeness).
  • Elapsed time from OOT result to completed assessment, as a median and a maximum.
  • Share of assessments where the list of potentially affected batches was generated from records (equipment logs, batch records, system data) rather than written from recollection. This is the hardest measure and the most telling.
  • Share of OOT records where the as-found value, the tolerance, and the date of the last good calibration are all recorded in a form that can be compared (conformance).

Typical fixes. A structured OOT record with mandatory fields for as-found value, tolerance, last good calibration date, and affected equipment ID; a standard query or report that lists batches processed on that equipment between the two dates; and a time limit in the procedure for the impact assessment. The report is often the step that turns a two-week assessment into a two-day one, and it only works if Pilot Two has made the equipment IDs agree.

What a good result looks like. Every OOT in scope has a completed assessment, every assessment lists affected batches from records, and the time from OOT to assessment is shorter and more consistent than at baseline. The team can produce the affected-batch list for any new OOT in a fixed, short time.

Pilot Five: Cleaning Status and Hold Time Data for One Equipment Train

Scope. One multi-product equipment train, and four timestamps for every cleaning cycle: end of use, start of cleaning, end of cleaning, and start of next use.

Data owner. The production supervisor for the train owns the cleaning records. Validation owns the approved dirty and clean hold time limits.

Why it matters. 21 CFR 211.67(b) requires written cleaning and maintenance procedures that include protection of clean equipment from contamination before use and inspection of equipment for cleanliness immediately before use.13 Section 211.182 requires equipment logs showing the date, time, product, and lot number of each batch processed, signed or initialed by the people who performed and checked the cleaning, with entries in chronological order.14 EU GMP Annex 15 is explicit about hold times: “The influence of the time between manufacture and cleaning and the time between cleaning and use should be taken into account to define dirty and clean hold times for the cleaning process.”15 Chapter 5 of the EU GMP Guide also lists cleaning status labels on equipment and manufacturing areas among the measures used to control cross-contamination.8

A validated hold time is only a control if the site can show, for each cycle, that it was met. That requires the timestamps to exist, to be in the right order, and to be recorded in a way that allows the hold time to be calculated. In its April 2026 warning letter to Fareva Amboise, FDA noted that the firm failed to maintain records of when equipment disinfection occurred and lacked cleaning validation data to support time limits for holding disinfected equipment before installation in the Grade A area.16 That case involved sterile manufacturing and disinfection rather than routine cleaning, but the data gap is the same one this pilot looks for: no timestamp, no hold time.

Measures.

  • Share of cleaning cycles with all four timestamps recorded (completeness).
  • Share of cycles where the timestamps are in a possible order: use ends before cleaning starts, cleaning ends before next use (plausibility).
  • Share of cycles where the calculated dirty and clean hold times fall within validated limits.
  • Share of spot checks where the equipment status label, the logbook, and any system status agree.

Typical fixes. Add explicit timestamp fields to the cleaning record where they were implied or combined, put the validated hold limits on the form itself so the person recording can see them, and if a system records the cleaning, have it calculate the hold times and flag an exceedance at the time rather than at review. The digital logbooks pilot design is a natural next step once this data is complete, because a digital log can enforce the timestamp order.

What a good result looks like. Near-complete timestamps on every cycle, no impossible orderings, hold times calculated for every cycle, and any exceedance raised and assessed when it happens rather than found weeks later during batch record review.

Pilots Six and Seven: Alarms and Deviation Coding

The last two pilots deal with data that people act on in the moment (alarms) and data that people use to see patterns over time (deviation categories). Both lose their value when there is too much noise or too little consistency.

Pilot Six: Alarm Data From One GMP Monitoring System

Scope. The alarm history from one monitoring system covering one area: for example, the environmental or building monitoring system for controlled-temperature storage or one cleanroom suite. Pull at least 30 days of history for the baseline.

Data owner. The system owner, usually in engineering or facilities. QA decides which alarms are GMP-relevant and what response each one requires.

Why it matters. An alarm that nobody reacts to is worse than no alarm, because it creates the belief that the condition is being watched. When a monitoring system generates many nuisance alarms, staff learn to acknowledge them without reading them, and the one alarm that matters is handled the same way. FDA’s February 2025 warning letter to ABR Laboratory is a plain example of the outcome: a 2–8°C refrigerator was out of range for more than 24 hours, reaching up to 17.4°C, and in FDA’s words, “This deviation was not identified or investigated.”17

The process industries have a mature standard for alarm system performance, ANSI/ISA-18.2, which requires alarm performance to be measured against goals in an alarm philosophy. It was written for control room operators in industries like chemicals and power, not for GMP monitoring systems, so its numbers are a reference point rather than a requirement. It is still the best public yardstick available. A Rockwell Automation white paper summarizes the ISA-18.2 performance metrics, which are based on at least 30 days of data:18

1 Average annunciated alarms per 10 minutes per operating position considered “very likely acceptable” (2 is “maximum manageable”)18
<1%–5% Target share of the total alarm load from the 10 most frequent alarms (5% at most), with action plans to address shortfalls18
<5 Stale alarms present on any day, with action plans to address them; the target for chattering and fleeting alarms is zero18

An overview of the standard by PAS, a member of the ISA committee, adds the caveat in the standard’s own text: the target metrics are approximate and depend on many factors, and “Alarm rate alone is not an indicator of acceptability.”19 For a GMP monitoring system, the more important question is whether every GMP-relevant alarm gets a documented response.

Measures.

  • Alarms per day for the area, and the share of the total from the ten most frequent alarms.
  • Count of chattering alarms (the same alarm repeatedly activating and clearing in a short time) and stale alarms (active for more than a day).
  • Share of GMP-relevant alarms with a documented response or assessment within the time the procedure requires (completeness).
  • Share of alarms whose setpoint, priority, and GMP classification are recorded in an approved alarm list (conformance).

Typical fixes. Rationalize the top ten alarms one at a time: adjust a deadband or delay, remove duplicates, correct a setpoint that no longer matches the approved range, or reclassify a non-GMP alarm so it no longer competes for attention. Each change goes through change control, and the approved alarm list becomes the reference for the system configuration.

What a good result looks like. A clear fall in alarms per day, the ten most frequent alarms either fixed or justified, no chattering alarms, and every GMP-relevant alarm in the hold period with a documented response. Staff should notice the difference without being told.

Pilot Seven: Deviation Coding Consistency in One Area

Scope. A sample of about 50 closed deviations from the past year in one area, recoded independently by two or three experienced reviewers using the site’s own category and root cause lists. Nobody sees the original codes or each other’s codes.

Data owner. The QA owner of the deviation process.

Why it matters. ICH Q10 expects a CAPA system that acts on deviations and on “trends from process performance and product quality monitoring,” and lists deviation, complaint, CAPA and change management processes among the performance indicators for management review.20 21 CFR 211.192 requires investigations to extend to other batches and products that may be associated with the same failure.3 Both depend on deviations of the same kind being coded the same way. If two reviewers would put the same event in different categories, the trend chart is showing reviewer habits, not process behavior.

Measures. Percent agreement and Cohen’s kappa between reviewers, calculated separately for the category field and the root cause field. Kappa adjusts for the agreement you would expect by chance. McHugh’s widely cited review of the statistic notes that many texts recommend 80% agreement as the minimum acceptable, and suggests that any kappa below 0.60 indicates inadequate agreement among raters.21 Those thresholds come from health research, not GMP, so treat them as a starting point and set your own target in the charter before you see the results.

Typical fixes. The disagreements point directly at the fixes. Categories that overlap get merged or given a decision rule (“if the operator followed the procedure and the procedure was wrong, code it as procedure, not human error”). Categories nobody uses get removed. Vague definitions get examples. Then a second sample of different deviations is recoded under the revised definitions and agreement is measured again.

Why this pilot is cheap and fast. It needs no system change and no new data. It needs a sample of closed records, a spreadsheet, and a few hours of reviewer time. It is also the pilot most likely to change what site leadership believes about its own trend reports. Once coding is consistent, the approach in our article on deviation trending analytics produces charts people can act on.

What a good result looks like. Agreement on the second sample meets the target set in the charter, the revised category list is approved and trained, and the next quarterly trend report is built on the revised coding. A clear finding that the old codes cannot support trending is also a good result, because it stops leadership from acting on misleading charts.

The Seven Side by Side, and How to Count the Payback

The table below summarizes the seven pilots. It is meant to help choose where to start, not to suggest running all seven at once. Most sites should run one or two per quarter.

PilotScopeData OwnerMain MeasureGood Result
1. Batch record correctionsOne product, one quarter of recordsProduction managerCorrections per record by field; review cycle timeTop five correction sources sharply down; faster disposition
2. Equipment identityOne suite’s major equipmentEngineering or maintenance leadIdentity and status match rate across systemsFull agreement; one rule for where IDs are created
3. Material statusOne warehouse zone or material classQA disposition leadStatus agreement in weekly blind sampleZero critical mismatches; no expired lots available
4. Calibration OOT recordsOne instrument class, twelve monthsCalibration lead, with QAAssessments complete; affected batches listed from recordsEvery OOT assessed quickly with a record-based batch list
5. Cleaning hold timesOne multi-product equipment trainProduction supervisorComplete, ordered timestamps; hold times within limitsHold time known for every cycle; exceedances raised at the time
6. Alarm dataOne monitoring system, one areaSystem owner, with QAAlarms per day; top ten share; documented responsesNuisance alarms removed; every GMP alarm answered
7. Deviation codingAbout 50 closed deviations, one areaQA deviation process ownerPercent agreement and kappa between reviewersAgreement meets charter target; trend reports rebuilt

How to Choose the First One

Three questions usually settle it. Where did the last inspection, internal audit, or deviation trend point? Which data set holds up release most often? And which data owner is ready to change something? A pilot with an eager owner and a medium-sized problem will do more in a quarter than a pilot with a large problem and a reluctant owner. If equipment identities are known to be unreliable, run Pilot Two before Pilot Four, since the calibration pilot depends on it. Pilot Seven is a good second pilot almost anywhere, because it runs alongside another without competing for the same people.

Counting the Payback Honestly

The return from these pilots is real but easy to overstate. We recommend counting it in the three forms introduced earlier, with the arithmetic shown so anyone can check it.

Hours given back. For each measure that reflects rework, the calculation is:

Hours per quarter = (baseline rate − new rate) × volume per quarter × hours per event

Baseline rate and new rate come from your pilot counts. Volume comes from your production or event records. Hours per event should be estimated by the people who do the work, by timing a handful of real events, not by assumption.

Here is a made-up example, only to show the arithmetic: a product runs 60 batches in a quarter; corrections fall from 12 to 5 per record; each correction takes about 20 minutes of combined operator and reviewer time. That is 7 × 60 × 20 minutes, or 140 hours per quarter. Your numbers will differ. The point is that the result is built from counts your site can verify, not from an industry benchmark.

Events avoided. For measures tied to deviations or discrepancies, count the events of that type in the baseline period and in the hold period. Only count events that did happen at baseline. Do not count “investigations we would have had” unless the baseline shows they were occurring.

Decisions made possible. Describe these in words, not numbers. For example: “The quarterly deviation trend report can now be used to select CAPA topics,” or “Affected-batch lists for calibration OOTs can be produced from records within a set time.” Leadership understands the value of these statements without a dollar figure attached, and attaching one usually weakens the case.

Four ways a pilot produces a result that is not real.

  • The baseline was taken after fixes started. Enthusiastic teams often begin fixing before measuring. The baseline then understates the problem and the improvement.
  • The measure definition changed mid-quarter. Redefining what counts as a correction or a mismatch can make any trend go the right way.
  • The period was unusual. A shutdown, a campaign change, or a product launch changes volumes and error rates on its own. Note these in the charter and in the decision paper.
  • Cleanup was counted as improvement. Correcting last year’s records improves the data set but not the process. The hold period, with no new fixes, is what shows whether the process changed.

After the Quarter: Expand, Hold, or Stop

Week 13 ends with a decision, and there are three honest options. Writing them down at the start, in the charter, makes the decision easier because the team agreed on the criteria before it saw the results. Our article on exit criteria for AI pilots makes the same argument in a different setting.

Expand

If the improvement held through the hold period and the fix was a change to how data is created, expand it. Expansion should be faster than the pilot: the charter, measures, and fixes already exist, and the next area mostly needs the same template change, the same system rule, or the same category definitions. Keep the same measure definitions so the results can be compared across areas. A site that expands Pilot Two suite by suite ends the year with a clean equipment master without ever running an equipment master data program.

Hold

If the improvement held but expansion is not the next priority, keep the measure running as a routine metric owned by the data owner, reported monthly. This is the step that most often gets dropped. Without a running measure, data quality drifts back as staff change, templates are revised, and new equipment arrives. A routine metric also gives the site something to show an inspector who asks how it knows its material status or cleaning data is reliable. Our manufacturing data quality scorecard article describes how such metrics can roll up for site leadership.

Stop

If the baseline showed the data was already in good shape, stop and record that finding. If the fixes did not work and the team cannot explain why, stop and write down what was learned before trying a different approach. A stopped pilot is not a failure. A pilot that runs on for three quarters because nobody wants to call it is.

Building a Rhythm

Sites that get the most from this approach run it as a rhythm rather than a one-off: one or two pilots per quarter, a short decision paper at the end of each, and a running list of candidate pilots fed by deviation trends, audit findings, and the people who create the data. Over a year that is four to eight data sets measured, fixed where needed, and handed to named owners with a routine metric. That is a data quality program, built from the bottom up, with evidence at every step.

The rhythm also changes how people at the site talk about data. Instead of a general sense that “our data is bad,” the site has specific statements: equipment identities in Suite 2 match across all systems; material status in the raw material zone has had no critical mismatches for two quarters; deviation coding agreement is above the target we set. Statements like these are what quality leaders need when they approve a new system, plan review by exception, or answer an inspector’s questions.

Conclusion

Manufacturing data quality improves one data set at a time, when someone who owns the process that creates the data decides to measure it, fix the cause, and keep measuring. The seven pilots here are chosen because each one has a clear owner, a baseline that can be pulled from existing records, and a result that shows up inside a quarter. None of them requires a new platform or a governance framework to start. Several of them, especially equipment identity and deviation coding, make the others easier, and together they build the foundation that review by exception, better trending, and any later use of AI in manufacturing all depend on.

Sakara Digital works with pharma and biotech organizations on manufacturing data quality and the systems that depend on it. If you are deciding which pilot to run first, or want an independent view on how to scope and measure one, we are happy to have that conversation.

For Further Reading