Why Most Data Quality Reports Go Unread

Ask a QA director what they do with the monthly data quality report and the honest answer is often “I look for red.” That is not a failure of the director. It is a sensible response to a report that was never designed for them. The typical report lists every rule the data team runs, every system it covers, and a completeness or validity percentage for each. It is accurate. It is also close to useless for the person who has to decide whether a site’s records can be trusted for release, whether a CAPA worked, or whether to ask for more reviewers next quarter.

This article is about the audience and the presentation, not about inventing new measures. Sakara Digital covered the core measures earlier this year in Data Quality Metrics That Matter. The question here is narrower and more practical: of everything you could measure, what are the data quality metrics pharma QA directors will read, understand in under a minute, and act on?

The Report Is Built for the Wrong Reader

Data quality reports usually start life as operational tools. A data steward needs to know which rule failed on which table last night. A system owner needs to know how many records are stuck in an interface queue. Those needs are real, and those reports should exist. The problem starts when the same report is forwarded upward with a cover note. The director now has the steward’s view of the world, with none of the steward’s context, and none of the decisions the director needs to make are visible in it.

Stephen Few, whose work on dashboard design is still the plainest guide to this problem, defined a dashboard this way: “A dashboard is a visual display of the most important information needed to achieve one or more objectives, consolidated and arranged on a single screen so the information can be monitored at a glance.”1 Two parts of that definition matter most for QA. The first is “most important information needed to achieve one or more objectives.” A report that does not name its objective cannot pick its most important information. The second is “a single screen.” Few lists exceeding the boundaries of a single screen as the first of his 13 common pitfalls, and supplying inadequate context for the data as the second.1

Numbers Without Context Do Not Get Acted On

Few’s second pitfall is the one we see most in quality reporting. A figure such as “record completeness 97.4%” invites the obvious questions he lists: compared to what? Is this good or bad? Is this better than before?1 Without a baseline, a trend, and a threshold, a QA director cannot tell whether 97.4% is normal, a slow slide, or the first sign of a system problem. So they do nothing, which is a reasonable choice when the report has not given them a reason to do anything else.

The ISPE Quality Metrics pilot reached a similar conclusion from a different direction. In summarizing what it learned from the first wave of its pilot with McKinsey and participating manufacturers, ISPE wrote that “Understanding context is crucial to interpreting results.”2 That finding came from quality metrics benchmarking across companies, but it applies just as much inside one site. A metric shown without its own history and its own limits is only half a metric.

Too Many Measures Compete for the Same Few Minutes

There is also a plain capacity limit. Research on working memory has argued for decades about whether people can hold about four items or about seven in mind at once, and recent work suggests the answer depends partly on how the task is set up.3 For our purposes the exact number does not matter. What matters is that it is small, and nowhere near the forty or sixty lines found in many monthly data quality packs. A director reviewing a page in a management meeting will hold a handful of points. The page should decide which handful, rather than leaving that choice to chance.

13Common dashboard design pitfalls identified by Stephen Few; the first two are exceeding one screen and giving inadequate context1
83Manufacturing sites from 28 companies across ISPE Quality Metrics pilot Waves 1 and 22
~3xISPE’s Wave 2 estimate of effort to collect FDA’s draft guidance metrics, compared with the figure in the Federal Register notice2

Start With the Decisions, Not the Data

The fastest way to shorten a data quality report is to stop asking “what can we measure?” and start asking “what does the QA director decide, and what would change that decision?” Every metric on the director’s page should answer to one of those decisions. If nobody can name the decision, the metric belongs on someone else’s page.

What the Regulations Already Ask Quality Leadership to Decide

The regulations give a useful starting list. Under 21 CFR 211.180(e), records must be kept so that the data “can be used for evaluating, at least annually, the quality standards of each drug product to determine the need for changes in drug product specifications or manufacturing or control procedures.”4 That is a decision: change or do not change. The same section requires written procedures for that evaluation, including a review of complaints, recalls, returned or salvaged products, and investigations.4

In the EU, Chapter 1 of the GMP Guide, which brought ICH Q10 concepts into EU GMP, says there should be periodic management review, with senior management involved, “to identify opportunities for continual improvement of products, processes and the system itself.”5 Its product quality review section also asks firms to highlight trends and to review the effectiveness of corrective and preventive actions taken after significant deviations.5

FDA’s guidance on a quality systems approach to CGMP is even more direct about what a management review should produce. It lists typical review outcomes as improvements to the quality system and related processes, improvements to manufacturing processes and products, and realignment of resources.6 That last one matters. A QA director’s page that never informs a resourcing decision is missing one of the three reasons the review exists.

A Working Catalog of QA Director Decisions

From those sources and from what we see in practice, most of the decisions a QA director makes that depend on data quality fall into five groups. We use this list as the filter for every metric that asks for space on the page.

Decision GroupTypical QuestionWhat Data Quality Has to Do With It
Release relianceCan we keep relying on this system’s records (and any review by exception) for batch disposition?If the records feeding release are wrong or incomplete, the release decision is exposed.
ResourcingDo we need more reviewers, a system fix, or a change in priorities next quarter?Backlogs and late reviews are usually a capacity signal before they become a compliance finding.
CAPA and effectivenessDid the fix work, or is the same data problem coming back?Repeat issues with the same root cause show a CAPA that did not hold.
EscalationDoes this need to go to site leadership, the quality council, or a notification procedure?Some data problems touch product quality directly and cannot wait for the next monthly meeting.
ChangeDo we need to change a specification, a procedure, or a system configuration?The 211.180(e) question, applied to the records themselves as well as the product.

The one-line test. Before a metric goes on the director’s page, the person proposing it should be able to finish this sentence: “If this number crosses this line, the QA director will ____.” If the blank is “note it” or “ask about it,” the metric belongs on a team-level report instead.

Who Owns the Metric Is Part of the Metric

A decision needs someone to bring it forward. For each metric on the page, name the owner who explains movement and proposes the action. That is usually a system owner, a lab lead, or the head of a review team, not the data team that calculates the number. The data team owns the calculation and the definition. The business owner owns the story and the recommendation. When those two roles are merged, the director tends to hear either a very technical explanation with no recommendation, or a recommendation with no evidence behind it.

The Short List: Six Metrics Worth a QA Director’s Time

The six metrics below are Sakara Digital’s recommended starting point for a site-level or product-family-level page. They are not a regulatory list, and a site with a very different profile may swap one or two. What they share is that each one ties to a decision group above, each one can be pulled from systems rather than assembled by hand each month, and each one tells the director something they cannot see from the standard quality metrics (deviations, CAPA, complaints) alone.

Several of them overlap with measures FDA itself has floated. In its 2022 Federal Register notice on the Quality Metrics Reporting Program, FDA listed examples under four practice areas: manufacturing process performance, pharmaceutical quality system (PQS) effectiveness, laboratory performance, and supply chain robustness. The examples included repeat deviation rate, CAPA effectiveness, laboratory right-first-time rate (which FDA said could be calculated from invalid assays or CGMP documentation errors found during review), and invalidated out-of-specification rate.7 That notice did not set any requirement, and it said so. It is still a useful signal of what regulators consider meaningful.

1. Record Right-First-Time Rate

What it is. The share of GMP records (batch records, logbooks, lab worksheets, electronic entries) that pass QA review without a correction, a query, or a return to the originator. Show it by record type, not as one blended number. An electronic batch record and a paper equipment logbook fail for different reasons.

Why the director cares. This is the most direct measure of whether records are right when they are made, which is the “contemporaneous” and “accurate” part of ALCOA+ (the attributes that make data trustworthy: attributable, legible, contemporaneous, original, accurate, plus complete, consistent, enduring, and available). It also tells the director how much hidden rework the review team is carrying.

The decision it drives. Release reliance and resourcing. A falling rate on a record type that feeds batch disposition is a reason to question any plan to reduce review on that record type. A steady high rate is part of the evidence a site needs before it moves to review by exception.

2. Audit Trail Review On-Time Rate and Finding Rate

What it is. Two numbers shown together. First, the share of scheduled audit trail reviews for GMP-critical systems completed by their due date. Second, the share of completed reviews that produced a finding needing follow-up (an unexplained change, a deletion, a data entry outside the expected window).

Why the director cares. The on-time rate is a compliance and capacity signal. The finding rate says whether the reviews are looking at the right things. A finding rate of zero across a year is not always good news. It can mean the review is a signature on a report nobody reads.

The decision it drives. Resourcing and escalation. Late reviews on a release-critical system are a hard line (see thresholds below). A persistent zero finding rate is a prompt to review what the audit trail review procedure asks reviewers to look for.

3. Data-Related Deviation Recurrence

What it is. The share of deviations whose assigned root cause is a data or documentation issue (a missing entry, a transcription error, an interface failure, an incorrect master data value) that repeat the same root cause at the same site within a set window, often twelve months.

Why the director cares. Recurrence is the clearest sign that a CAPA did not hold. The ISPE Wave 2 report identified an internal metric, deviations recurrence rate, as one that companies could use to help predict external quality outcomes.2 FDA’s 2022 notice also listed repeat deviation rate as an example PQS effectiveness metric, calculated from total deviations and the number with the same assignable root cause.7 Narrowing it to data-related root causes turns a general quality metric into a data quality metric the director can act on.

The decision it drives. CAPA effectiveness and change. A recurring data root cause usually means the fix was retraining when it should have been a system or procedure change. For a closer look at why closure-time targets can hide this, see our piece on why thirty days is the wrong investigation metric.

4. Invalidated Out-of-Specification Rate

What it is. The share of out-of-specification (OOS) results that the lab invalidated after investigation, usually shown per number of tests or lots. FDA’s quality metrics notices have used this measure under the name invalidated OOS rate.7

Why the director cares. An invalidated OOS means the lab concluded the original result was not a true measure of the sample. A high or rising rate points to problems with lab data reliability: method robustness, analyst practice, instrument performance, or, at worst, a habit of explaining away results. Few data quality measures get closer to the question an inspector will ask about the lab.

The decision it drives. Escalation and change. A signal on this metric is a reason for a focused review of invalidation rationales, not just a trend comment. ISPE found that the way this metric is calculated changes what it shows, so fix the definition before you trend it.2

5. Release-Critical Reconciliation Breaks

What it is. The count and age of open mismatches between systems for the specific fields that feed batch disposition. Typical pairs are the manufacturing execution system (MES) and the laboratory information management system (LIMS), or LIMS and the enterprise resource planning (ERP) system that holds batch status. Only include fields that matter for release: batch identifier, material status, result values, expiry dates, and similar.

Why the director cares. When two systems disagree about a release-critical fact, somebody resolves it by hand, usually under time pressure. Each manual fix is a point where the record of what happened can drift away from what happened. The age of open breaks matters as much as the count.

The decision it drives. Release reliance and change. If breaks are growing, the director should question any automated release check that depends on those interfaces. It is also one of the clearest ways to put an IT fix in front of the people who decide IT priorities. Our earlier article on the manufacturing data quality scorecard covers MES and LIMS reconciliation from the operations side.

6. Aged Data Quality Defects by Criticality

What it is. The number of open data quality defects (known errors in master data, reference data, system configuration, or stored records) that are past their target fix date, split by criticality. Show only high and medium criticality on the director’s page.

Why the director cares. This is the backlog view. It shows whether known problems are being fixed at the pace the site committed to, and it keeps high-criticality items visible as they get older.

The decision it drives. Resourcing and escalation. A growing count of aged high-criticality defects is a direct input to the “realignment of resources” outcome FDA describes for management review.6

MetricShow AsDecision GroupUsual Owner
Record right-first-time rateRate by record type, 13-month trendRelease reliance, resourcingHead of QA operations or batch record review
Audit trail review on-time and finding rateTwo rates side by side, by system tierResourcing, escalationSystem owners, with QA oversight
Data-related deviation recurrenceRate plus count of repeat root causesCAPA, changeDeviation and CAPA process owner
Invalidated OOS rateRate per tests or lots, with countEscalation, changeQC laboratory lead
Release-critical reconciliation breaksOpen count and oldest ageRelease reliance, changeBusiness system owner for MES, LIMS, or ERP
Aged data quality defectsPast-due count by criticalityResourcing, escalationData governance lead

Resist the seventh, eighth, and ninth metric. Every review cycle, someone will ask to add “just one more.” Keep a written rule: a new metric can join the director’s page only if one comes off, and only if its sponsor passes the one-line decision test. Without that rule, the page grows back into the pack it replaced within a year.

Setting Thresholds That Mean Something

Thresholds are where most quality dashboards go wrong. A single red-amber-green band, often set by gut feel at the start of the program and never revisited, is asked to do three different jobs at once: flag compliance breaches, flag unusual movement, and show progress toward a goal. It does none of them well.

Three Kinds of Line, Kept Separate

Hard Line

Procedure or Compliance Limits

Set by your own procedures or commitments. Example: every scheduled audit trail review for a release-critical system is done by its due date. A breach is a breach, whether it happens once or ten times. No statistics needed.

Signal Line

Process Behavior Limits

Calculated from the metric’s own history, usually 12 to 24 months. These tell you when a change is larger than normal month-to-month noise. They describe the process as it is, not as you want it.

Target

Improvement Goals

Where leadership wants the metric to be by a date. Targets belong on the page as a reference line, not as an alarm. Missing a stretch target in month three is not a signal.

Rule

Pre-Agreed Response

For each line, write down in advance what happens when it is crossed: who explains, by when, and which decision the director is being asked to make.

Use the Metric’s Own History for Signal Lines

The most reliable way to tell a real change from noise is a control chart or run chart built on the metric’s own history. This is the core of statistical process control, and it has been used in manufacturing for a long time. The practical point for a QA director’s page is that the signal line is calculated, not chosen. If the process is stable, the line will move only when the process changes.

There is a trade-off in how many signal rules you apply. The NIST Engineering Statistics Handbook notes that with the standard rule (signal only on a point beyond three standard deviations) a stable process will produce a false alarm every 371 points on average. Adding the Western Electric rules, a common set of extra pattern tests, raises the false alarm frequency to about once in every 91.75 points on average.8 The handbook leaves the choice to the user and notes that some users add the extra rules but give their signals less weight in troubleshooting.8

A 2018 study by Anhøj and Wentzel-Larsen compared the diagnostic value of common control chart rules and reached a similar conclusion: the more tests applied, the higher the risk of false positive results, and the choice of rules should be made deliberately, preferably before data collection begins.9 Their practical suggestion was to start with a run chart and a small set of rules, and add the three-sigma rule only once the process shows random variation.9

371Average points between false alarms for a stable process using only the three-sigma rule8
~92Average points between false alarms after adding the Western Electric rules (NIST gives 91.75)8
4Practice areas FDA proposed for quality metrics in 2022: process, PQS, laboratory, supply chain7

For a monthly page, that trade-off is concrete. Six metrics, each with a full set of pattern rules, will produce false alarms often enough that the director learns to ignore amber. Pick one or two rules per metric, write them down, and stick with them for a year before you revisit.

Small Numbers Need Special Handling

Many data quality metrics at site level have small denominators. A site might complete a modest number of audit trail reviews on its release-critical systems each month, and a single late review can move the rate by several points. Two habits help. First, always show the count next to the rate, so the director can see that “down 8 points” means one review. Second, for metrics with small monthly counts, trend on a rolling quarter or use a chart designed for counts rather than forcing a monthly percentage.

A Worked Example of a Threshold Set

Here is how the three kinds of line might look for the audit trail review metric. The structure matters more than the specific settings, which each site should derive from its own procedures and history.

LineHow It Is SetWhat Happens When Crossed
Hard lineFrom the audit trail review procedure: every review for a release-critical system done by its due dateAny breach is reported to the QA director within the week, with the system name and the plan to close it. Considered for escalation if the system fed a disposition while the review was late.
Signal lineRun chart of the monthly on-time rate across all GMP systems, using the prior 18 months as baselineSystem owner explains the shift at the monthly review and proposes whether it is capacity, scheduling, or a system issue.
TargetSet by quality leadership for the yearShown as a reference line. Discussed at quarterly management review, not monthly.

How Often to Show What

Cadence is part of presentation. The same metric can be useful weekly for one audience and noise for another. The goal is to match each view to the decision rhythm it supports, so the director sees exceptions quickly, trends monthly, and system-level questions quarterly.

Why Monthly Is the Anchor, and Annual Is Not Enough

The regulations set a floor, not a rhythm. 21 CFR 211.180(e) requires product quality standards to be evaluated at least annually.4 FDA’s quality systems guidance goes further: “Although the CGMP regulations (§ 211.180(e)) require product review on at least an annual basis, a quality systems approach calls for trending on a more frequent basis as determined by risk.”6 The same guidance notes that trending enables the detection of potential problems as early as possible.6

For data quality, monthly is the right anchor for most sites. It is frequent enough to catch a slide before it becomes an inspection finding, and slow enough that each point on the trend has a meaningful count behind it. Weekly views have their place, but for exceptions, not trends.

1

Weekly: Exceptions Only

A short note, sent only when a hard line is crossed. No charts. System, what happened, who owns it, and the date for the next update. If nothing crossed a hard line, nothing is sent. The director learns that a weekly note means something.

2

Monthly: The One-Page Review

The six metrics, each with 13 months of trend, signal lines, and a one-line owner comment. The page opens with decisions requested. This is the core document and the one most of this article describes.

3

Quarterly: Management Review Input

The same six metrics rolled into the quality management review, alongside deviations, CAPA, complaints, and audit results. The focus shifts from single-month signals to whether the data quality picture supports or undermines the quality system’s other claims.

4

Annually: Product Review and Metric Review

The data quality trends that touch each product feed its annual product review. Separately, quality leadership reviews the metric list itself: which metrics drove decisions, which never did, and which definitions need to change.

What Changes at the Quarterly Review

FDA’s quality systems guidance lists the typical inputs to a management review, including the analysis of data trending results, the status of actions to prevent a potential problem or a recurrence, and follow-up actions from previous management reviews.6 That is a useful structure for the quarterly view. Instead of repeating the monthly page three times over, the quarterly input should answer three questions: which signals appeared this quarter and what was decided, whether the actions from last quarter’s decisions have worked, and whether any metric needs a change in its lines or its owner.

The same guidance says reviews should take place more often when a quality system is new than when it has matured.6 The same logic applies to a new data quality page. For the first two quarters, consider reviewing it monthly with the full quality leadership team, not only the director. That builds shared understanding of what each line means before the page settles into routine.

The One-Page Layout, Section by Section

The monthly page is where the work above becomes visible. The layout below fits on one printed page or one screen without scrolling. It follows a fixed order, so a director who reads it every month knows exactly where to look.

The Layout in One Table

ZoneContentSpace on the Page
Header stripSite or product family, reporting month, data as-of date and time, source systems, definitions version, prepared byOne line
A. Decisions requestedNo more than three items. Each states the decision, the recommendation, the owner, and the date it is needed byTop quarter of the page
B. Six metric rowsFor each metric: current value with count, 13-month trend line with signal lines and target, status word, and a one-line owner commentMiddle half of the page
C. What changed since last monthShort bullets: signals that appeared or cleared, actions closed, lines or definitions that changedA few lines
D. Watch listUp to three items that are not yet signals but that an owner wants the director to know aboutA few lines
FooterLink to full definitions, link to the team-level reports, contact for questionsOne line

Zone A: Decisions Requested Goes First

The most important design choice is putting decisions at the top. Most reports put decisions (if they include them at all) at the end, after the charts. That order assumes the reader will work through everything. A director with ten minutes before a meeting will not. Leading with the requests means the page is useful even if the director reads only the top quarter.

Each request should be written so the director can say yes or no. “Approve an additional contract reviewer for batch record review in Q1 to bring record right-first-time review times back within procedure” is a request. “Record right-first-time rate is declining; monitoring continues” is not. If there are no decisions requested this month, say so in one line. That is useful information too.

Zone B: One Row per Metric, Same Shape Every Time

Each metric row has the same five parts in the same order: name, current value with its count, a small trend chart covering 13 months (so the same month last year is visible), a status word, and an owner comment of one sentence. Few warns against introducing meaningless variety, such as using a different chart type for each measure just to keep the display interesting.1 Six identical rows let the eye compare them quickly.

Use status words rather than colors alone. “Signal,” “Breach,” “Within limits,” and “Improving” carry meaning that a color does not, and they survive black-and-white printing and color-blind readers. Few also cautions against misusing or overusing color.1 Reserve color for the one or two rows that need attention this month.

Show precision that matches the decision. Few’s third pitfall is displaying excessive detail or precision.1 A rate shown to two decimal places on a base of forty records suggests an accuracy the number does not have. Whole percentages, with the count beside them, are enough.

An Illustrative Filled-In Page

The mock-up below shows how the zones come together. The site, the values, and the comments are invented for layout purposes only. They are not benchmarks, and they should not be used as targets.

Illustrative example only. All values are invented to show layout.

Header: Site B, Oral Solids | Month: example month | Data as of the first business day, 06:00 | Sources: eQMS, LIMS, MES | Definitions v2.1 | Prepared by: Data Governance Lead

A. Decisions requested (2)

  • Approve a temporary reviewer to clear the audit trail review backlog on the chromatography data system by month end. Owner: QC Lead. Needed by: next Friday.
  • Agree that the MES to LIMS interface fix moves ahead of the planned reporting upgrade in the IT queue. Owner: MES System Owner. Needed by: quarterly IT priorities meeting.

B. Metrics

  • Record right-first-time (batch records): 91% of 212 | Within limits | “Stable; packaging records improved after form change.”
  • Audit trail review on-time: 88% of 25 | Breach (release-critical) | “Three late reviews on chromatography data system; see decision 1.”
  • Data-related deviation recurrence: 3 of 14 | Signal | “Third repeat of manual transcription error on weighing step; CAPA under review.”
  • Invalidated OOS rate: 1 of 6 OOS | Within limits | “No change.”
  • Reconciliation breaks (release-critical): 9 open, oldest 23 days | Signal | “Interface drops status updates on weekends; see decision 2.”
  • Aged defects (high/medium): 2 / 7 | Within limits | “One high item closes this month.”

C. What changed: Recurrence signal new this month. Invalidated OOS signal from last quarter cleared after method update.

D. Watch list: New LIMS release scheduled next quarter; owner will confirm audit trail configuration before go-live.

Notice what is missing. There is no rule-by-rule list of data checks, no system uptime, no training completion percentage, and no single composite score. Those may all be tracked somewhere. They are not on this page because none of them passed the decision test at this level.

Zones C and D Keep the Story Honest

“What changed since last month” stops the page from looking the same every month. It is the section a returning reader looks at first. The watch list gives owners a sanctioned place to raise early concerns without triggering a formal signal. Both sections should be short. If either grows past a handful of lines, it is a sign that something belongs in Zone A as a decision.

A quick usability check. Hand a draft page to a quality leader who has not seen it before and give them 60 seconds. Then ask three questions: What is the director being asked to decide? Which metric needs attention? Is anything worse than last month? If they cannot answer all three, the layout needs another pass.

Metrics to Retire or Push Down a Level

Shortening a report is as much about removing as choosing. Most of the measures removed from the director’s page should not be deleted. They should move to the team that can act on them. The director’s page gets shorter, and the team-level reports become more useful because they are finally built for their real audience.

Composite Data Quality Scores

A single index that blends completeness, validity, timeliness, and consistency into one number looks tidy. It is almost always the wrong thing to put in front of a QA director. When it moves, nobody can say why without opening the pieces, and a large problem in one area can be hidden by small improvements elsewhere. If your organization reports a composite score to executives, keep it off the QA page and make sure the underlying measures are visible to the people who own them. Our article on why an executive KPI can have three different numbers covers what happens when roll-ups drift away from their sources.

Measures That Everyone Meets

A metric that every site hits every month tells the director nothing. The ISPE pilot offers a clear example from quality metrics. ISPE did not include the annual product review or product quality review on-time rate in its second wave, because findings from the first wave indicated that it was not differentiating.2 That does not mean on-time product reviews do not matter. It means that as a metric, it did not help tell sites apart. The same test applies to data quality measures. If a measure has sat at the same value for a year with no variation, move it to an annual check.

Raw Counts Without Denominators

“Data quality issues logged: 146” invites the wrong reaction. Is that high? Did the site process twice as many batches this month? Did a new detection rule start running? Counts belong on the page only beside a rate or an age, as in the reconciliation breaks metric above, where the oldest age carries most of the meaning.

Measures of Activity Rather Than Outcome

Number of data quality rules run, number of records scanned, and number of dashboards published are measures of the data team’s activity. They are fine for managing that team. They do not tell the QA director whether records are trustworthy. The same is true of training completion rates for data integrity courses. Completion shows that training happened, not that behavior changed. Record right-first-time rate and recurrence are closer to the outcome the training was meant to produce.

Commonly Reported MeasureWhere It BelongsWhy Not on the Director’s Page
Composite data quality indexExecutive or enterprise data report, with components visibleCannot be explained or acted on without breaking it apart
Field completeness for all fieldsData steward reportMixes critical and trivial fields; release-critical fields are covered by other metrics
Rules executed, records scannedData team operationsActivity measure, not outcome
System uptimeIT service reportRelevant only when it affects records, which reconciliation and defects already show
Data integrity training completionTraining compliance reportShows training happened, not that records improved
Product review on-time rateAnnual quality system checkISPE found this type of measure not differentiating2

Keeping the Page Honest Over Time

A good page in month one can be a misleading page by month eighteen. Definitions change without notice, a source system is replaced, and people learn how to make a number look better without making the underlying records better. Four habits keep the page trustworthy.

Version the Definitions and Say Where the Data Came From

Every metric needs a written definition: numerator, denominator, inclusions, exclusions, source system, and calculation timing. Put a version number on the definitions set and print it in the header strip. When a definition changes, note it in Zone C and mark the trend line so the director does not read a definition change as a performance change.

This matters more for data quality than for most metrics, because the words themselves are loosely used. In their review of methods for assessing electronic health record data quality, Weiskopf and Weng found five broad dimensions (completeness, correctness, concordance, plausibility, and currency) and noted that “There was a great deal of variability and overlap in the terms used to describe each of these dimensions.”10 Their setting was clinical research, not manufacturing, but the lesson carries over. Two teams can both report “accuracy” and mean different things. A written, versioned definition is the only defense.

The ISPE Wave 2 findings show how much a definition can change a result. When calculated using FDA’s draft definitions, the three FDA metrics ISPE evaluated did not show relationships with external quality outcomes or culture indicators. When ISPE’s alternative calculations were used, the same three metrics did show relationships with culture indicators.2 Same concept, different arithmetic, different conclusion.

Watch for Metrics That Get Managed Instead of Improved

Any metric that matters to people will be managed. That is the point. The risk is when the number gets better while the thing it stands for does not. Manheim and Garrabrant describe Goodhart’s Law as occurring when a metric that can be used to improve a system “is used to an extent that further optimization is ineffective or harmful.”11 In data quality, the common forms are easy to recognize: reclassifying deviation root causes so they no longer count as data-related, closing reconciliation breaks by adjusting one system without finding out which one was right, or completing audit trail reviews on time by narrowing what the review covers.

Pairing metrics is the most practical defense. The audit trail review metric already pairs on-time rate with finding rate, so a site that speeds up reviews by making them shallower will see its finding rate fall. Record right-first-time rate pairs naturally with data-related deviation recurrence. When one member of a pair improves sharply and the other does not move, look closer.

Collect From Systems, Not by Hand

A page that takes a week to assemble each month will not survive a busy quarter. The ISPE Wave 2 report estimated that the effort to collect the FDA draft guidance metrics was about three times the figure in FDA’s Federal Register notice, and called that probably an underestimate for some companies.2 Collection effort is a real constraint. Choose metrics whose inputs already live in the eQMS, LIMS, MES, or audit trail tools, and automate the extraction before you roll the page out. A hand-assembled number also carries its own data quality risk, which is an awkward thing to discover on a data quality report.

Review the List Itself Once a Year

Once a year, look back at every decision recorded from the page. Which metrics triggered decisions? Which ones never did? Which signal lines fired so often they were ignored, or never fired at all? Retire what did not earn its place, and consider whether a new risk (a new system, a new product, a new contract manufacturer) needs a metric.

Where the Regulatory Picture Stands

It is worth being clear about what is and is not required. FDA’s Quality Metrics Reporting Program has not become a reporting requirement. FDA’s program page describes a 2015 draft guidance proposing a mandatory program, a 2016 revised draft describing a voluntary phase, two pilot programs announced in 2018, and a public docket opened in March 2022 to take comments on a revised approach.12 The 2022 notice stated plainly that it was not intended to communicate regulatory expectations for reporting.7 One point from that notice is still worth taking to heart. Among the key lessons FDA said it drew from its two pilot programs, it wrote that “Any metric chosen to be reported should be meaningful to the practice area being measured, and the data collected on that metric should be able to influence decision making about process improvements and capital investments.”7 That is the one-line decision test, in FDA’s words.

Separately, FDA’s Quality Management Maturity (QMM) program aims to encourage manufacturers to adopt quality management practices that go beyond CGMP requirements. FDA announced a third cohort of its QMM prototype assessment protocol evaluation program in 2026.13 Neither program tells a site which data quality metrics to use. Both point in the same direction as this article: fewer measures, clearly defined, used to make decisions.

Enforcement shows the other side. In an October 2025 warning letter to an OTC topical drug manufacturer, FDA cited the firm’s quality unit for failing to exercise its responsibilities, including failing to ensure performance of periodic (at least annual) product review under 21 CFR 211.180(e).14 That is a basic failure, well short of the practices described here, but it is a reminder that product review and trending are expected of the quality unit, not optional extras.

Checklist: Is Your Page Ready?

  • Six metrics or fewer, each tied to a named decision group
  • Decisions requested at the top, written as yes or no questions
  • Every metric shows its count, 13 months of trend, and separate hard, signal, and target lines
  • One or two written signal rules per metric, agreed before the first report
  • A named owner for each metric who explains movement and recommends action
  • Versioned definitions and a data as-of date in the header
  • Inputs extracted from systems, not assembled by hand
  • An annual review of the metric list on the calendar

Conclusion

The data quality metrics pharma QA directors act on are few, stable, and tied to decisions they already own: whether to keep relying on a system’s records for release, where to put reviewers, whether a CAPA held, when to escalate, and when to change a procedure or system. A page built around those decisions, with each metric shown against its own history and a small set of pre-agreed rules, gets read. A page built around everything the data team can measure does not, however accurate it is. The regulations already expect trending and management review. What they leave open is presentation, and that is where most of the value is lost or gained.

Sakara Digital works with pharma and biotech quality and data leaders on exactly this kind of problem: choosing the handful of measures that matter, defining them so they hold up, and building the review rhythm around them. If you are rebuilding a data quality report for your quality leadership and want an independent view on what to keep and what to retire, we are happy to have that conversation.

For Further Reading