In This Article
- Executive Summary
- Three Reports, Three Numbers, One Meeting
- What Eleven Years of FDA Quality Metrics Work Actually Proved
- What ICH Q10 and Part 211 Ask of Management Review
- Quick Win One: Consolidate the Reports Carrying One KPI
- Quick Win Two: Run a Single-Number Audit on One Executive KPI
- Quick Win Three: Harmonize Site Comparison Before Anyone Compares Sites
- Presenting the Reconciled Number Without Undermining Trust
- Knowing the Pilot Worked, and What Comes Next
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Somewhere in your organization, the same quality KPI has three different values. Right-first-time is 94 percent in the site operations review, 91 percent in the quarterly quality report, and 88 percent in the deck that went to the board. Nobody is wrong. Each number came from a different system, applied a different definition, or cut the data on a different day. The result is that leadership stops trusting the reporting layer, and the conversation in the management review turns into an argument about arithmetic rather than a decision about the business.
This is not a governance failure that needs a five-year fix. It is a definition and lineage problem, and it responds well to small, bounded pilots. This article is part of Sakara Digital’s Data Quality Quick Wins series for life sciences, whose premise is that leadership confidence is earned by a few measurable pilots rather than a multi-year program. It sets out three quick wins, each scoped to two to four weeks with a before-and-after measure: consolidating the reports that carry a single KPI, running a single-number audit that traces one executive figure back to source records, and harmonizing definitions before anyone compares manufacturing sites.
The argument is grounded in evidence that most leaders have never seen. FDA has been working on pharmaceutical quality metrics since 2015 and has never finalized a guidance. The industry pilot work that ran alongside it, involving 83 sites across 28 companies, found that the same data, calculated under different definitions, produced different conclusions about the same sites. If the regulator and 28 companies could not agree on what a lot acceptance rate is, the three numbers in your management review deck are not an accident. They are the predictable result of never writing the definition down.
Three Reports, Three Numbers, One Meeting
The pattern is familiar enough that most quality and IT leaders recognize it from the first sentence. A KPI appears on the site operations dashboard. It appears again in the monthly quality report that goes to the leadership team. It appears a third time in a slide prepared for the board or for a corporate quality review. The three values do not match. Somebody notices, usually in the meeting rather than before it, and the next twenty minutes are spent trying to work out which figure is right.
It happens to every KPI that senior leaders actually care about: right-first-time, deviation closure rate, batch cycle time, complaint rate, on-time release. These are the metrics with the widest audience, which is exactly why they get rebuilt independently in several places. A metric that only one team uses tends to have one calculation. A metric that five teams use tends to have five.
There are only three underlying causes, and it helps to name them separately because each has a different fix.
Cause one: different systems
The site dashboard pulls deviation records from the quality management system. The corporate report pulls from a data warehouse that was loaded from the same quality management system, but through an extract written four years ago that filtered out a record type nobody remembers. The board slide was built in a spreadsheet by someone who exported a report and added the sites that were missing. All three numbers describe reality. They describe different extracts of it.
System divergence is the easiest cause to detect and often the easiest to fix, because the difference between two extracts can usually be reconciled record by record. It is also the cause people jump to first, which is a problem, because it is not the most common one.
Cause two: different definitions
This is the one that does the real damage. Right-first-time can mean batches released with no deviation. It can mean batches released with no critical or major deviation. It can mean batches that cleared quality review without a query being raised. It can be counted on batches manufactured in the period, batches released in the period, or batches that reached final disposition in the period. Each of those is a defensible definition. Each produces a different number from identical source records.
Definition divergence is hard to see because nothing looks broken. Two reports run correctly, against the same system, and disagree. Unless someone opens the calculation logic in both, the disagreement looks like a data quality problem when it is a semantics problem.
Cause three: different cut times
The site dashboard refreshes nightly. The monthly report is cut on the fifth working day. The board pack is assembled a week before the meeting from whatever was current at the time. In a quality system where investigations close continuously and records are back-dated to the event, a metric cut on the fifth working day and the same metric cut on the fifteenth will differ, and the later cut is usually the worse-looking one because late-closing investigations are late for a reason.
The tell: if the three numbers are close but not equal, and the older report always looks better than the newer one, you are looking at a cut-time problem rather than a definition problem. If the gap is wide and stable across periods, it is almost always a definition problem.
What all three causes share is that none of them is a data integrity failure. The underlying records are fine. The failure is in the layer between the record and the slide, and that layer has almost never been documented, reviewed, or subjected to change control the way the source systems have. A validated electronic batch record system carries full traceability. The spreadsheet that turns its output into a percentage on a board slide usually carries none.
What it does to leadership confidence
The practical damage is not the twenty minutes lost in the meeting. It is what happens afterward. Once a leadership team has been burned twice by conflicting numbers, three things follow. Executives start asking for the raw data rather than the summary, which pushes analytical effort back onto teams that were supposed to be relieved of it. Decisions get deferred pending reconciliation, which turns a monthly review cycle into a quarterly one. And the quality organization loses standing in exactly the forum where it most needs to be believed.
There is also a regulatory dimension that leaders tend to underestimate. Management review of the pharmaceutical quality system is an expectation under ICH Q10, and the annual product review is a requirement under Part 211. If an inspector asks how a performance indicator presented to senior management was derived, and three versions of that indicator exist with no documented reconciliation, the answer is uncomfortable. The issue is not that a number was wrong. It is that the company cannot demonstrate which number management actually reviewed.
What Eleven Years of FDA Quality Metrics Work Actually Proved
Before proposing fixes, it is worth understanding how hard this problem is, because the evidence is public and most people in industry have never read it.
On 28 July 2015, FDA published a draft guidance for industry titled Request for Quality Metrics, under docket FDA-2015-D-2537, together with a Federal Register notice and a public meeting.1 The draft proposed a mandatory program in which manufacturers would submit data at product level, from which FDA would calculate four primary metrics. The intent was stated plainly: use the data for risk-based inspection scheduling, to help predict and mitigate drug shortages, and to encourage modern quality management systems.2
The comments were extensive. In November 2016, FDA published a revised draft, Submission of Quality Metrics Data, which explicitly states in its own footnotes that in response to comments received in the public docket, FDA was replacing the 2015 draft with the revised version.3 The revised draft did two things: it dropped the mandatory framing in favor of a voluntary reporting phase, and it reduced the calculated metrics from four to three.4
The three surviving metrics, in FDA’s own words in that draft, were defined as follows:
| Metric | FDA definition in the 2016 revised draft | Stated as an indicator of |
|---|---|---|
| Lot Acceptance Rate (LAR) | The number of accepted lots in a timeframe divided by the number of lots started by the same covered establishment in the current reporting timeframe. | Manufacturing process performance |
| Product Quality Complaint Rate (PQCR) | The number of product quality complaints received for the product divided by the total number of dosage units distributed in the current reporting timeframe. | Patient or customer feedback |
| Invalidated Out-of-Specification Rate (IOOSR) | The number of OOS test results for lot release and long-term stability testing invalidated by the establishment due to an aberration of the measurement process, divided by the total number of lot release and long-term stability OOS test results in the current reporting timeframe. | Operation of a laboratory |
Read those definitions closely and the reason this is difficult becomes visible. Every one of them contains at least one term that a company has to interpret before it can produce a number. What counts as a lot started? What counts as an accepted lot? Is a lot released with an unexpectedly low yield accepted or not? FDA had to write a glossary to answer some of these, defining a lot as a batch or a specific identified portion of a batch having uniform character and quality within specified limits, and defining an accepted lot as a started lot that has been released for distribution or for the next stage of processing.3 The agency also published a separate technical conformance guide in 2016 to specify the data format.5
FDA’s own selection criteria for the data set are instructive. The agency said it wanted objective data to provide consistency in reporting, of the type contained in records subject to inspection, and valuable in assessing the overall effectiveness of a pharmaceutical quality system without an undue reporting burden.3 Consistency in reporting was the first criterion, and it is the one that proved hardest.
The industry pilot that tested the definitions
While FDA was drafting, ISPE ran a two-wave pilot program, working with McKinsey and Company, to test whether a standardized set of quality metrics could actually be collected and reported. Wave 1 was reported in June 2015 and covered 44 sites from 18 companies. Wave 2 started in July 2015 and reported in June 2016, with 21 participating companies; across both waves the pilot covered 83 sites from 28 companies.6
Wave 2 was deliberately redesigned once FDA’s 2015 draft appeared, so that it could test the proposed FDA metrics as defined and, at the same time, test alternative definitions applied to the same data. Three of the four FDA metrics were evaluated. The fourth, an annual product review or product quality review on-time rate, was dropped because Wave 1 had found it did not differentiate between sites.6
The findings deserve to be quoted in substance because they are the strongest available evidence that definition, not data, is the constraint:
What the ISPE Wave 2 pilot found
- Calculated using the FDA definitions, the three metrics did not show relationships with external quality outcomes or with culture indicators.
- Calculated using the alternative definitions the pilot defined, the same three metrics, on the same data, did show relationships with culture indicators.
- The effort to collect the FDA draft guidance metrics was approximately three times the estimate given in the Federal Register notice, and ISPE considered that an underestimate for over-the-counter companies and companies with complex supply chains.
- Using ISPE’s recommended calculations for the same three metrics was estimated at one-third the effort of collecting data according to the FDA calculations.
- ISPE recommended changing the lot acceptance rate denominator away from lots attempted, changing the complaint rate denominator from dosage units to packs, and dropping the double normalization in the invalidated out-of-specification rate.
One specific finding is worth carrying into your own organization. The pilot compared two candidate denominators for lot acceptance rate: lots attempted and lots dispositioned. It found the two counts highly related but different in magnitude, and the relationship differed by site type. Across 18 solid dose sites the fitted slope between the two counts was 1.57 with an R-squared of 95 percent. Across 18 sterile sites the fitted slope was 1.03 with an R-squared of 59 percent.6
Read that plainly. Choosing one denominator rather than the other changed the size of the number materially, and it changed it by a different amount depending on the type of site. Any cross-site comparison built on top of an unspecified denominator was comparing site types, not site performance.
Where the initiative stands today
In June 2018, FDA announced two voluntary pilot programs in the Federal Register: a quality metrics site visit program for CDER and CBER staff, and a quality metrics feedback program for establishments with existing metrics programs.78 In March 2022, FDA established a public docket describing considerations for refining what it then called the Quality Metrics Reporting Program, drawing on lessons from the 2018 pilots and on stakeholder feedback about the 2016 revised draft. Comments closed on 7 June 2022.9
As of this writing, the position is unchanged in one important respect. The 2016 revised draft guidance still carries the standard header stating that it is a draft, not for implementation, containing nonbinding recommendations.10 No quality metrics guidance has been finalized. No mandatory reporting requirement is in force. FDA’s public page on quality metrics for drug manufacturing still describes the program in these terms.4
Do not tell your leadership team that FDA quality metrics reporting is coming into force. It has been eleven years since the first draft and there is no final guidance, no compliance date, and no mandatory submission requirement. Building a business case on an imminent regulatory deadline that does not exist is the fastest way to lose credibility on this topic. The case for fixing your metrics is internal: leadership needs numbers it can act on, and management review needs a defensible record.
The useful lesson from this history is not about regulatory timing. It is about difficulty. A regulator with statutory authority, a technical conformance guide, an industry association, a consulting firm, and 28 volunteering companies spent years on this and could not converge on three definitions. Your organization has almost certainly spent less effort on the definition of right-first-time than that. The three numbers in your deck are not a sign of incompetence. They are the normal outcome of an unwritten definition.
What ICH Q10 and Part 211 Ask of Management Review
The regulatory framing that does apply gets less attention than a reporting mandate would, but it is more directly relevant to the meeting where the three numbers appear.
ICH Q10 describes a pharmaceutical quality system model across the product lifecycle and sets an expectation that management reviews the performance of the system. The model calls for monitoring of performance indicators, periodic management review of the quality system, and documented outcomes from that review, including escalation of appropriate issues to senior management.11 Q10 does not tell you which indicators to use or how to calculate them. That is deliberate. It leaves the choice to the company, which means the company owns the burden of showing that its chosen indicators are defined, consistent, and derived from records it can produce.
In the United States, 21 CFR 211.180(e) requires that written records be maintained so that data in them can be used for evaluating, at least annually, the quality standards of each drug product, to determine the need for changes in specifications or in manufacturing or control procedures. The regulation names what that evaluation must include: a review of a representative number of batches, whether approved or rejected, and a review of complaints, recalls, returned or salvaged product, and investigations.12 The annual product review, and its European counterpart the product quality review, is therefore a second place where the same underlying counts get aggregated, often by a different team, on a different calendar, with a different definition.
This is why organizations routinely end up with more copies of a KPI than they expect. The metric exists in operational reporting because sites need it weekly. It exists in management review because Q10 expects performance indicators. It exists in the annual product review because Part 211 requires an annual evaluation. It exists in the board pack because leadership asked for a trend. Four legitimate demands, four independent builds, one KPI.
The question an inspector can reasonably ask
Show me the performance indicator that management reviewed in the third quarter. Now show me how it was calculated, from which records, over which period, and who approved that calculation.
Most organizations can answer the first part immediately and the rest slowly or not at all. That gap, rather than the value of any particular metric, is the exposure.
The PDA has published useful practitioner work on this point, arguing that standardizing key performance indicators and metrics is a prerequisite for assessing the state of a pharmaceutical quality system, while acknowledging that good definitions are usually specific to an organization, particularly in how a KPI is calculated from raw data.13 Both halves of that statement matter. Standardization is necessary. Universal standardization across the industry has repeatedly failed. What works is internal standardization, written down, with a named owner.
That is the whole basis for the three quick wins below. None of them tries to solve metrics for the enterprise. Each of them takes one metric, or one report family, or one comparison, and makes it defensible inside a few weeks.
Quick Win One: Consolidate the Reports Carrying One KPI
Scope: two to three weeks. Team: one analyst, one quality subject matter expert, one report owner from each function that publishes the KPI. Before-and-after measure: the number of live reports, dashboards, and recurring decks that publish the chosen KPI, and the number of distinct calculation paths behind them.
Start with the inventory, because the count itself is usually the finding. Pick one KPI that senior leaders reference by name. Then find every place it is published.
Week one: inventory every artifact carrying the KPI
Search the business intelligence platform for reports and dashboards containing the metric name and its common variants. Ask each function to submit the recurring reports and decks they produce that include it. Include spreadsheets that are emailed on a schedule, because those are usually the ones nobody counts and often the ones leadership sees. Record the owner, the refresh frequency, the audience, and the source system for each.
Week one to two: open the calculation behind each one
For every artifact, write down the numerator, the denominator, the filters applied, the date field used for period assignment, and the refresh or cut schedule. Do this by reading the query or the formula, not by asking the owner what it does. Owners describe intent accurately and implementation approximately, and the gap between the two is where the divergence lives.
Week two: group into distinct calculation paths
Most inventories collapse to fewer paths than artifacts. Twelve reports may run on three calculations. Name each path, state what it measures, and identify which audience each path serves. This is the point at which the disagreement becomes explainable rather than embarrassing.
Week two to three: agree one source and retire duplicates
Choose one calculation path as the published version for leadership reporting. Where a second path serves a genuine operational need, keep it and rename it so the two are never mistaken for each other. Retire everything else, with the owner’s agreement, and redirect its audience. Record the decision and the rationale in a short memo signed by the KPI owner.
Two practical notes carry most of the difficulty. First, renaming matters as much as retiring. If a site keeps a local right-first-time calculation because it needs a weekly operational signal, that is fine, but it should be called something else, because the moment two things share a name they will be compared. Second, resist the urge to build a new report. The consolidation quick win is about subtraction. A new dashboard added to an inventory of eleven produces twelve.
What good looks like at the end of three weeks: a one-page register listing every artifact that published the KPI, the calculation path behind each, the single path now designated for leadership reporting, the artifacts retired, and the named owner of the surviving definition. That page is the deliverable. It is also the artifact you hand an inspector who asks how the indicator in management review was derived.
Quick Win Two: Run a Single-Number Audit on One Executive KPI
Scope: three to four weeks. Team: one analyst, one data engineer with access to the source systems, one quality subject matter expert. Before-and-after measure: the percentage difference between the highest and lowest published value of the KPI for the same period, before the audit and after the reconciliation.
The consolidation quick win tells you how many versions exist. The single-number audit tells you why they differ, by tracing one figure from the executive slide back to the source records and documenting every transformation on the way.
The lineage trace, worked
Take one KPI, one period, one reported value. Then walk backward. At each step, record what the value is at that point, what operation was applied, and where the operation is defined. Here is what a completed trace looks like for a deviation closure rate reported to leadership as 91 percent for a quarter.
| Step | Artifact or system | Value at this point | Operation applied here | Where the rule is defined |
|---|---|---|---|---|
| 0 | Executive slide, quarterly business review | 91% | Rounded from source; prior-quarter figure restated in footnote | Nowhere. Built in the slide. |
| 1 | Quarterly quality report workbook | 90.6% | Weighted average of four site figures by deviation volume | Formula in the workbook; no written procedure |
| 2 | Site extract, four sites | Four values, 84% to 96% | Each site applies its own on-time target (30 or 45 days) | Local site procedures, three versions |
| 3 | Warehouse view dev_closure_v2 | Record-level | Excludes deviations with status “on hold pending supplier response” | View definition; change made 14 months ago, no ticket |
| 4 | Warehouse load from QMS | Record-level | Nightly incremental load; records amended after load are not re-read | Integration specification, section 4.2 |
| 5 | Quality management system | Source records | Closure date is the date of final quality approval | QMS configuration and SOP |
A trace like this normally takes an analyst and an engineer three to five working days per KPI, and it usually turns up two or three findings that nobody knew about. In the example above, step 3 is the finding: an exclusion added over a year ago with no record of who approved it, which removes the slowest population of deviations from the denominator and therefore improves every downstream number. Step 4 is the second finding, and a subtler one: a nightly incremental load that never re-reads amended records means late corrections never reach the report at all.
Expect to find an undocumented exclusion. In practice this is the single most common finding of a lineage trace. A filter was added to solve a real problem at the time, often to remove records the business genuinely considered out of scope, and the reason was never written down. It is not misconduct. It becomes a problem only when the exclusion outlives the person who understood it and the metric is presented as if no exclusion exists.
Running the audit
The sequence that works is deliberately narrow. Pick the KPI leadership questions most often, not the one that is easiest to trace. Pick a closed period so the underlying records are stable. Then work backward one layer at a time, and at each layer, reproduce the value independently before moving on. If you cannot reproduce the value at a layer, you have found the divergence and you stop and investigate rather than continuing the trace.
Two rules keep this from expanding. First, one KPI, one period. The temptation to trace three at once is strong and it is what turns a four-week pilot into a six-month project. Second, document as you go, in the table format above, rather than writing the report at the end. The table is the deliverable.
The definition template
The audit produces a reconciled number. What makes the reconciliation stick is writing the definition down in a form that is specific enough to be implemented identically by two people who have never met. The following template is short on purpose. Anything longer does not get filled in.
| Field | What to record | Worked example: deviation closure rate |
|---|---|---|
| Metric name | The exact name used in reporting, with no abbreviation | Deviation closure rate, on time |
| Plain statement | One sentence a non-specialist can read | The share of deviations closed within the target period, out of deviations due for closure in the period |
| Numerator | Exact population and the field that identifies it | Deviations with quality approval date on or before target closure date |
| Denominator | Exact population and the field that identifies it | Deviations with target closure date falling within the reporting period |
| Inclusions | Record types, sites, product families in scope | All GMP deviations, all four commercial sites, all product families |
| Exclusions | Every exclusion, with the reason and the approver | None. The prior supplier-hold exclusion was removed on reconciliation. |
| Date field | Which date assigns a record to a period | Target closure date, not event date and not approval date |
| Target | The threshold and where it is defined | 30 calendar days from event date, per the harmonized deviation procedure |
| Source system | System of record and the specific object | QMS deviation module, production instance |
| Cut rule | When the period closes for reporting, and the restatement rule | Fifth working day after period end; restated once at 30 days, then frozen |
| Owner | Named individual accountable for the definition | Head of quality systems |
| Change control | How a change to this definition gets approved | Quality systems change record; effective from the next reporting period only |
Three fields on that template do more work than the rest. The date field, because period assignment is the least visible source of divergence and almost nobody records it. The cut rule, because a metric that gets restated every month without announcement can never be reconciled against an older report. And change control, because the whole point of the exercise is to prevent the definition drifting again the moment someone reasonable makes a reasonable change.
The restatement rule is a decision, not a detail
Deciding that a figure is cut on the fifth working day, restated once at 30 days, and then frozen forever is a policy choice. It means the number in the board pack will sometimes be superseded. Everyone accepts that in financial reporting. It is worth saying out loud in quality reporting, because the alternative, which is a figure that keeps improving after publication, is what makes leaders stop trusting the trend.
Quick Win Three: Harmonize Site Comparison Before Anyone Compares Sites
Scope: three to four weeks. Team: the quality lead from each site, one analyst, one corporate quality owner to arbitrate. Before-and-after measure: the number of sites using an identical definition, calculation, and reporting calendar for the chosen KPI, before and after.
This is the quick win with the highest political content and the highest payoff. The moment a network of manufacturing sites is ranked on a single KPI, the ranking drives attention, investment, and sometimes performance conversations about named individuals. If the underlying definitions differ, the ranking is measuring definitional choices rather than performance, and the sites usually know it before corporate does.
The ISPE pilot finding described earlier is the evidence to bring into this conversation. Two defensible denominators for the same lot acceptance rate produced counts that differed in magnitude, and differed by different amounts at solid dose sites than at sterile sites.6 A network with both site types, comparing lot acceptance rate without a common denominator, is comparing dosage form.
What divergence looks like in practice
A typical starting position across four sites, for right-first-time, looks like this. None of these sites is doing anything improper. Each definition was set locally for a local reason.
| Element | Site A | Site B | Site C | Site D |
|---|---|---|---|---|
| Numerator | Batches with zero deviations | Batches with no critical or major deviation | Batches released with no quality review query | Batches with zero deviations |
| Denominator | Batches released in period | Batches released in period | Batches manufactured in period | Batches reaching final disposition in period |
| Rework treatment | Counted as a failure | Counted as a failure | Excluded entirely | Counted as a failure |
| Reporting calendar | Calendar month | Calendar month | 4-4-5 fiscal month | Calendar month |
| Cut day | 3rd working day | 5th working day | 2nd working day | 10th working day |
Every cell that differs is a reason two sites cannot be compared. Site C’s exclusion of rework and its use of a fiscal calendar make its figure structurally higher and structurally offset in time. Site D’s tenth working day cut captures late closures that Site A’s third working day cut does not, which makes Site D look worse for a reason that has nothing to do with how Site D operates.
The harmonization sequence
One definition
Get the site quality leads in one room and agree a single numerator, denominator, inclusion set, and exclusion set. Expect the argument to be about rework and about deviation severity. Resolve it by asking what decision the metric is meant to inform, then choosing the definition that best serves that decision, and writing down what was rejected and why.
One calculation
Implement the agreed definition once, centrally, from source data. Do not ask each site to implement the agreed definition locally and submit a percentage. Sites submitting percentages is how definitions diverge again within two quarters. Sites should submit records or counts; the calculation happens in one place.
One calendar
Agree the period boundary, the cut day, and the restatement rule, and apply them to every site identically. If one site genuinely needs a fiscal calendar for other reasons, produce its fiscal view separately and keep the comparison view on the common calendar.
One back-calculation
Recalculate the last four to eight periods under the harmonized definition before publishing anything. Without a restated history, the first harmonized report shows every site moving at once and nobody can tell improvement from redefinition.
Step 4 is the one most often skipped and the one that most often causes the harmonization to fail. If the first comparable report is also the first report under the new definition, every site can plausibly claim its number moved because of the change, and the exercise produces a new argument rather than ending an old one.
Sequence the announcement carefully. Harmonize the definition, back-calculate the history, brief the site leads on their own restated numbers privately, and only then publish the comparison. A site lead who first sees a restated figure in a group forum will spend the meeting defending it rather than acting on it.
Presenting the Reconciled Number Without Undermining Trust
This section addresses the part of the work that has nothing to do with data and determines whether the pilot succeeds. You have reconciled a KPI. The reconciled value differs from what leadership has been shown for the past several quarters, and it is usually worse, because undocumented exclusions and early cuts tend to flatter. Now you have to present it.
The failure mode is to present the reconciliation as a discovery of error. That framing invites two damaging questions: what else has been wrong, and who is responsible. Neither question has a good answer, and both distract from the finding that actually matters, which is that the organization now has a number it can defend.
What to say
Lead with the decision, not the discrepancy. Open by stating the reconciled figure, the definition behind it, and what it means for the business. Then explain the reconciliation as the reason the figure can now be relied on, rather than as a correction of past reporting.
Be specific about why previous figures differed, in neutral terms. An exclusion was applied at the warehouse layer and was not carried into the definition. Two sites used different closure targets. The figure was cut on the fifth working day and late closures were never re-read. These are mechanical explanations, and mechanical explanations are easier to accept than judgments about people.
Say clearly what has not changed. The underlying operational performance did not move. The records were always correct. What changed is the calculation applied to them and, critically, the fact that the calculation is now written down and under change control.
A structure that works for the reconciliation briefing
- The number. The reconciled figure for the current period, and the restated series for the prior four to eight periods so the trend is readable.
- The definition. One sentence in plain language, plus the completed definition template as an appendix.
- The reason for the difference. Two or three mechanical causes, quantified where possible: the exclusion accounted for roughly this much, the cut timing for roughly that much.
- What did not change. Source records, site behavior, the underlying trend direction.
- What is now different. A named owner, a written definition, a change control route, and a stated restatement rule.
- What you want from them. Endorsement of the definition as the version used in management review, and agreement not to accept an unreconciled version of the same KPI from another source.
What not to say
Do not say the old numbers were wrong. In almost every case they were not wrong, they were differently defined, and the distinction is both accurate and considerably easier for an audience to absorb. Do not present a long list of every discrepancy found during the trace; present the two or three that account for most of the difference and keep the rest in the appendix. Do not commit, in that meeting, to reconciling every other KPI. The value of a quick win is that it is bounded, and a leadership team energized by one good result will happily extend the scope until it stops being a quick win.
Finally, do not let the reconciled number be presented by an analyst alone. The person who owns the definition should present it, because the durable message is that this metric now has an owner. That is the change that prevents the three numbers coming back.
Knowing the Pilot Worked, and What Comes Next
Every quick win in this series carries a before-and-after measure that can be stated in one line and verified by someone who was not involved. For these three, the measures are as follows.
| Quick win | Measure before | Measure after | Evidence produced |
|---|---|---|---|
| Dashboard consolidation | Count of live artifacts publishing the KPI, and count of distinct calculation paths | Same two counts, after retirement and renaming | One-page register with owners and the designated leadership version |
| Single-number audit | Percentage spread between highest and lowest published value for the same period | Spread after reconciliation, target zero for the leadership version | Completed lineage trace table and completed definition template |
| Site harmonization | Number of sites sharing an identical definition, calculation, and calendar | All sites in scope | Agreed definition, restated history, and the record of rejected alternatives |
Those measures are deliberately unglamorous. That is the point. A leadership team that has been promised data transformation before is more persuaded by a stated count that moved from eleven to two than by a maturity assessment.
What to do next, and what not to
The natural next step after one successful reconciliation is the second KPI, then the third. Do them one at a time, and use the same definition template, because a set of five completed templates is the beginning of a real definition library and five differently formatted documents is not.
The step to avoid is converting the quick wins into a program. The moment this becomes a data governance initiative with a charter, a steering committee, and a two-year roadmap, it acquires overhead that exceeds the value of the next few reconciliations, and it starts competing for attention with everything else on the portfolio. The organizations that get furthest with this work keep it small deliberately, run two or three reconciliations a year, and let the definition library accumulate.
The maturity signal worth watching for: somebody outside the original team asks for the definition template before building a new report, without being told to. That is the point at which the practice has taken hold, and it usually happens after the third or fourth reconciliation rather than after a policy announcement.
Where these pilots point
There is a longer-term payoff that is worth stating even though it should not be the justification for the pilot. Every one of these quick wins produces exactly the artifacts that an AI or advanced analytics initiative later requires: documented metric definitions, traced lineage from source record to reported figure, named owners, and change control on the calculation layer. Organizations that attempt to apply machine learning to quality data before doing this work generally discover that the model has learned the definitional inconsistency rather than the process.
But the immediate justification is simpler and stronger. Leadership needs to be able to look at a number and act on it. Three numbers means no decision. One number, defined, owned, and traceable, means the management review can be about the business again.
Conclusion
The three-numbers problem is not evidence of a broken quality system or a failed data strategy. It is evidence that the layer between the record and the slide was never treated with the same discipline as the record itself. The regulatory history makes the point better than any argument could: FDA has worked on pharmaceutical quality metrics since 2015, published a draft, replaced it, run pilot programs, opened a docket, and still has no final guidance, largely because agreeing what a metric means turned out to be harder than collecting the data behind it. When 28 companies and a regulator cannot converge on the definition of a lot acceptance rate, an unwritten definition inside a single company is going to produce more than one number every time.
The fix does not require a program. It requires picking one KPI that leadership actually uses, counting how many places publish it, tracing one published figure back to source records, writing the definition down with an owner and a change control route, and then presenting the reconciled figure as the beginning of something reliable rather than the correction of something wrong. Two to four weeks per pilot, a measure that moves, and an artifact you can hand to an inspector. Do that three times and the reporting layer starts to earn the confidence the source systems already have.
Sakara Digital works with pharma and biotech organizations building exactly this kind of defensible reporting layer, from metric definition and lineage tracing through to the governance that keeps definitions from drifting again. If you are looking at a KPI with more than one value and want an independent perspective on where to start, we are happy to have that conversation.
For Further Reading
For Further Reading
- The Quality Metrics Dashboard That Actually Drives Investigations to Closure
- The Manufacturing Data Quality Scorecard: KPIs Beyond Regulatory Submissions
- Master Data Management for Life Sciences: Creating a Single Source of Truth Across Global Operations
- Right-First-Time Manufacturing KPIs for Cell Therapy Sites
- Quality Metrics That Actually Drive Improvement in Pharma
- Deviation Trending Analytics: From Excel to Real-Time Dashboards
References & Sources
- Food and Drug Administration. “Request for Quality Metrics; Notice of Draft Guidance Availability and Public Meeting; Request for Comments.” Federal Register, 28 July 2015. https://www.federalregister.gov/documents/2015/07/28/2015-18448/
- Food and Drug Administration. “Request for Quality Metrics: Guidance for Industry (Draft).” Docket FDA-2015-D-2537, July 2015. https://downloads.regulations.gov/FDA-2015-D-2537-0027/attachment_1.pdf
- Food and Drug Administration, CDER and CBER. “Submission of Quality Metrics Data: Guidance for Industry (Revised Draft).” November 2016. https://www.fda.gov/media/93012/download
- Food and Drug Administration. “Quality Metrics for Drug Manufacturing.” Pharmaceutical Quality Resources. https://www.fda.gov/drugs/pharmaceutical-quality-resources/quality-metrics-drug-manufacturing
- Food and Drug Administration. “Quality Metrics Technical Conformance Guide: Technical Specifications Document, Version 1.0.” June 2016. https://www.fda.gov/media/98939/download
- ISPE. “ISPE Quality Metrics Initiative: Quality Metrics Pilot Program Wave 2.” June 2016. https://ispewebassets.org/files/regulatory/2023/QMWAVE2DL.pdf
- Food and Drug Administration. “Quality Metrics Site Visit Program for Center for Drug Evaluation and Research and Center for Biologics Evaluation and Research Staff; Information Available to Industry.” Federal Register, 29 June 2018. https://www.federalregister.gov/documents/2018/06/29/2018-14006/
- Food and Drug Administration. “Modernizing Pharmaceutical Quality Systems; Studying Quality Metrics and Quality Culture; Quality Metrics Feedback Program.” Federal Register, 29 June 2018. https://www.federalregister.gov/documents/2018/06/29/2018-14005/
- Food and Drug Administration. “Food and Drug Administration Quality Metrics Reporting Program; Establishment of a Public Docket; Request for Comments.” Federal Register, 9 March 2022. https://www.federalregister.gov/documents/2022/03/09/2022-04972/
- Food and Drug Administration. “Submission of Quality Metrics Data: Guidance for Industry.” Guidance document record, docket FDA-2015-D-2537. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/submission-quality-metrics-data-guidance-industry
- European Medicines Agency. “ICH Q10 Pharmaceutical quality system: Scientific guideline.” https://www.ema.europa.eu/en/ich-q10-pharmaceutical-quality-system-scientific-guideline
- Electronic Code of Federal Regulations. “21 CFR 211.180: General requirements.” Part 211, Subpart J, Records and Reports. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-J/section-211.180
- Parenteral Drug Association. “Points to Consider when using KPIs/Metrics.” PDA Journal of Pharmaceutical Science and Technology, 15 June 2021. https://journal.pda.org/content/early/2021/06/15/pdajpst.2020.012401
- ISPE. “FDA Quality Metrics.” ISPE Quality Metrics Initiative. https://ispe.org/initiatives/quality-metrics/fda-quality-metrics
- ISPE. “ISPE Quality Metrics Initiative: A Report from the Pilot Project Wave 1.” June 2015. https://ispewebassets.org/files/regulatory/2023/QMWAVE1DL.pdf
- International Pharmaceutical Quality. “Wave 2 of ISPE’s Quality Metrics Pilot Will Explore Metric Relationships and How to Deepen Industry’s Role in Advancing the Regulatory Process.” https://ipq.org/wave-2-of-ispes-quality-metrics-pilot-will-explore-metric-relationships-and-how-to-deepen-industrys-role-in-advancing-the-regulatory-process/
- Regulations.gov. “Docket FDA-2015-D-2537: Request for Quality Metrics; Draft Guidance for Industry.” https://www.regulations.gov/docket/FDA-2015-D-2537








Your perspective matters—join the conversation.