In This Article
- Executive Summary
- What Counts as Master Data in Pharma and Biotech
- Where Bad Master Data Shows Up in Operations
- How to Put a Price on Bad Master Data
- Pilot 1: Product Master Source-of-Truth Check
- Pilot 2: Customer and Site Master Reconciliation
- Pilot 3: Reference Data Inventory
- Pilot 4: Specification Cleanup
- Turning Pilot Results Into a Business Case
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Every pharma and biotech company runs on a small set of shared records: the product, the material, the specification, the site, the customer, and the code lists that describe them. When those records disagree across systems, people fix the problem by hand, deviations get opened against the wrong limits, and submissions and shipments stall. The spending is real, but it is spread across so many teams that it rarely appears as a line in anyone’s budget.
That is why master data programs struggle to get funded. Leaders hear industry averages and vendor benchmarks, but they approve money when they see their own numbers. The fastest way to get those numbers is to run small, bounded pilots that measure error rates, rework hours, and downstream events in your own records.
This article describes four such pilots: a product master source-of-truth check, a customer and site master reconciliation, a reference data inventory, and a specification cleanup. Each can be run in a few weeks with a small team and existing system access. For each one we cover the question it answers, how to run it, what to measure, and how to turn the result into a dollar figure. We close with a simple model for combining the four results into a business case that a CFO and a head of quality can both sign.
What Counts as Master Data in Pharma and Biotech
Gartner describes master data as the smallest consistent set of identifiers and attributes that describe the core entities of an enterprise and that many business processes share1. The key word is shared. A batch record, a lab result, or a sales order is a transaction. It happens once. The product, the material, the site, and the specification it refers to are master data. They get created once and then read thousands of times by other systems and other people.
That reuse is what makes master data so valuable and so risky. A correct material record helps every purchase order, every goods receipt, and every batch that uses the material. An incorrect one damages all of them, and it keeps doing so until someone finds and fixes the source.
Five Domains That Matter Most
In a regulated drug company, five domains carry most of the operational and compliance weight. The table below lists them, where they usually live, and what tends to go wrong. The systems named are the usual ones: regulatory information management (RIM), enterprise resource planning (ERP), manufacturing execution systems (MES), laboratory information management systems (LIMS), and the quality management system (QMS) that holds controlled documents, deviations, and CAPA (corrective and preventive action) records.
| Domain | Typical Systems | Common Failure | Where It Shows Up |
|---|---|---|---|
| Product master | RIM, ERP, labeling, serialization | Name, strength, pack size, or product code differs between systems | Labeling changes, serialization exceptions, submission rework |
| Material master | ERP, MES, LIMS | Wrong unit of measure, shelf life, storage condition, or status | Batch record errors, planning errors, expired material in use |
| Specification | QMS documents, LIMS, ERP inspection plans | Limits, units, or method versions in LIMS do not match the approved specification | False out-of-specification results, or real failures that pass unnoticed |
| Customer and site master | ERP, RIM, QMS supplier list, trading partner data | Duplicate sites, outdated addresses, missing or wrong identifiers | Rejected shipment data, facility information gaps in applications |
| Reference data | Every system, often as local code lists | Different lists for units, dosage forms, countries, or deviation categories | Reports that cannot be combined, manual mapping, trending that misleads |
Regulation already treats several of these as controlled records. Under 21 CFR 211.186, the master production and control record for each drug product must include the product name, strength, and dosage form, a complete list of components designated by names or codes specific enough to show any special quality characteristic, and the specifications to be followed2. In other words, the regulation assumes that product, material, and specification data are correct and consistent before a batch is ever made.
Regulators are also building their own master data services. The European Medicines Agency runs four master data services, known as SPOR for substance, product, organization, and referential data, that follow the ISO standards for the identification of medicinal products (IDMP). The referentials service holds controlled lists of terms such as dosage forms and units of measurement, and the organizations service holds names and location addresses for marketing authorization holders, manufacturers, and other entities3. EMA’s stated goal is for companies to supply regulatory data once and reuse it across procedures. That goal only works if the company’s own internal data agrees with itself first.
Master Data Versus Reference Data
People often use these terms loosely, so it helps to separate them. Master data describes specific things the company deals with: this product, this material, this contract manufacturer, this hospital. Reference data is the set of allowed values used to describe those things: the list of dosage forms, the list of units, the list of countries, the list of deviation categories. Reference data changes less often, but when it drifts, it damages every master record that uses it. That is why one of the four pilots below focuses on reference data on its own.
Why the Errors Stay Out of Sight
Master data errors are rarely dramatic on the day they are created. A pack size entered as 30 instead of 28, a site address that still shows the previous owner’s name, or a LIMS limit that was never updated after a variation all look like small clerical issues. The damage appears later and somewhere else: in a warehouse, in a lab investigation, in a distributor’s rejected file, or in a reviewer’s information request. By then, the people dealing with the consequence are usually not the people who created the record, and nobody connects the two.
This distance between cause and effect is the core reason master data quality is underfunded. The business case has to rebuild that connection with evidence.
Where Bad Master Data Shows Up in Operations
Before designing pilots, it helps to name the three kinds of damage you are trying to measure. Each one has a different owner and a different way to count it.
Rework: The Correction Work Nobody Budgets
Thomas Redman, writing in Harvard Business Review, described how people who receive bad data usually fix it themselves to meet a deadline instead of going back to the person who created it. He called these added correction steps the “hidden data factory.” He also cited IBM’s estimate that poor quality data cost the United States $3.1 trillion in 2016, and put the share of knowledge workers’ time wasted on hunting for data, correcting errors, and looking for confirming sources for data they do not trust at 50%4.
Those national figures are too broad to use in a pharma business case, but the pattern behind them is familiar in any drug company. Regulatory operations staff check product data against three systems before a submission. Supply chain planners keep a private spreadsheet of “real” lead times because the ERP values are wrong. Quality reviewers correct material codes in batch records. Each of these is a person doing work that exists only because a master record was wrong at the source.
Redman’s point is that this correction work creates no value for the customer and can be reduced sharply. He reported that organizations that go back to the data creator, share requirements, and remove root causes almost always reduce the associated costs by two thirds, and often by 90% or more4.
Deviations and Investigations
In a good manufacturing practice (GMP) setting, bad master data does not only waste time. It can generate quality events. The clearest example is the specification. FDA’s guidance on out-of-specification (OOS) results defines OOS results as all test results that fall outside the specifications or acceptance criteria established in drug applications, drug master files, official compendia, or by the manufacturer. The same guidance notes that FDA regulations require an investigation whenever an OOS result is obtained (21 CFR 211.192), and that a failure investigation extending to other batches or products that may be associated with the failure must be completed5.
Now picture a LIMS record where a limit was entered one decimal place tighter than the approved specification. Every result in that gap becomes an OOS result that must be investigated, documented, and closed, even though the product was within its approved limits all along. The reverse case is more serious: a LIMS limit looser than the approved specification lets a real failure pass without an investigation. Either way, the root cause is a master data record, not the manufacturing process or the lab.
Material master errors produce a similar pattern. A wrong shelf life or retest period on a material record can allow expired material to be issued, or can block good material from use. A wrong unit of measure can produce a weighing or dispensing error that becomes a deviation. The investigation usually ends with “data entry error” as the root cause, and the CAPA corrects the one record. Few companies go on to ask how many other records have the same problem.
Delays in Submissions and Supply
The third kind of damage is time. FDA’s question-and-answer guidance on identifying manufacturing establishments in applications notes that a lack of clarity about what facility information to include has led to applications with extraneous, misplaced, or missing information, and that these issues cause delays in the assessment process and, in some cases, unnecessary information requests, refuse-to-file, and refuse-to-receive actions. The same guidance explains that FDA needs the FDA Establishment Identifier (FEI) number to proceed with the facility evaluation, and that the agency’s preferred unique facility identifier for registration is the Data Universal Numbering System (DUNS) number6. Site master data that is incomplete or inconsistent between the regulatory, quality, and supply systems feeds directly into these problems.
The supply chain has its own version. The Healthcare Distribution Alliance’s exceptions handling guidelines for the Drug Supply Chain Security Act (DSCSA) describe what happens when a distributor receives electronic product data and cannot find the manufacturer’s Global Location Number (GLN) or the product’s Global Trade Item Number (GTIN) in its own master data. The system flags an error. For a missing GTIN, part or all of the file is rejected depending on how the distributor’s systems are set up. In both cases, the manufacturer has to correct its master data and resend7. At an HDA webinar reported by RAPS, a manufacturer’s traceability lead said manufacturers need “meticulous” oversight of GTINs and GLNs and should understand that these numbers are “moving targets,” since a GTIN can change when packaging changes or when the company goes through an organizational change. A distributor executive described a company using two sets of location codes for the same hospital, one under its abbreviated name and one under its full name8.
Large projects also slow down. McKinsey reports that around 65 percent of advanced planning system programs fail to achieve their expected return on investment, and that poor data management is one of the top five reasons. The same article gives a specific example from supply chain planning: inconsistencies in lot-sizing data between production versions and the material master. In McKinsey’s experience, discovering problems like this late can double testing timelines and add 15 to 20 percent to project budgets9.
The 47% figure comes from a study by Nagle, Redman, and Sammon, who asked executives to check the last 100 records their own departments had created. We return to their method below, because it is the basis for all four pilots10.
How to Put a Price on Bad Master Data
Most master data proposals fail at the funding stage, not the design stage. Understanding why helps you design pilots that produce the right evidence.
Why Industry Benchmarks Do Not Win Funding
There is no shortage of large numbers. Gartner reports that poor data quality costs organizations at least $12.9 million a year on average, based on its research from 202011. Redman, writing in MIT Sloan Management Review, estimates the cost of bad data at 15% to 25% of revenue for most companies12. These figures are useful for getting attention. They are not useful for getting a budget approved, for three reasons.
- They are not about your company. A CFO can always argue that the average does not apply to a company with your products, systems, and sites.
- They do not point to a fix. A figure spread across the whole enterprise gives no guidance on which records, which systems, or which teams to start with.
- They cannot be checked later. If the program is funded on an industry average, there is no baseline to measure improvement against.
A useful business case needs the opposite: numbers that come from your own records, that point to specific fixes, and that can be measured again after the fix.
Why Master Data Programs Get Framed Wrong
A second reason proposals stall is that master data management is often presented as a software purchase. A full MDM platform is a large, multi-year commitment, and the business case for it tends to rest on promised future benefits. When a leadership team is asked to approve a platform before anyone has shown what the data problems are and what they cost, the answer is often “not this year.”
The pilots in this article reverse the order. They measure the problem first, using the systems you already have. If the numbers justify a platform, the pilots also give you the requirements and the baseline. If the numbers point to a smaller fix, such as a new ownership model or a change to one data entry process, you have saved the platform spend.
A common mistake: launching an enterprise-wide data quality assessment as the first step. Broad assessments take months, produce long lists of issues with no dollar value, and lose sponsor attention before they finish. A pilot scoped to a few dozen records in one domain produces a defensible number in weeks.
The Rule of Ten
The simplest model for pricing bad data comes from the same HBR study. The authors describe the “rule of ten,” which states that “it costs ten times as much to complete a unit of work when the data are flawed in any way as it does when they are perfect.” Their example: if 100 units of work each cost $1 with perfect data, and 11 of them have flawed data, the total rises from $100 to $199. They also point out that the rule does not account for costs such as lost customers, bad decisions, or damage to reputation10.
The rule of ten is a starting assumption, not a law. In a GMP environment, the multiplier can be much higher when a data error triggers a deviation, an OOS investigation, or a batch hold. The pilots below replace the assumed multiplier with measured hours and measured events wherever possible.
Mapping Master Data Errors to Cost of Quality
Quality leaders already have a language for this. ASQ defines cost of quality as a method for finding out how much of an organization’s resources go to preventing poor quality, to appraising quality, and to internal and external failures. Internal failure costs are those for defects found before the product or service reaches the customer, and external failure costs are those for defects found after13. Master data errors fit this model well, and using it makes the business case familiar to finance and quality at the same time.
| Cost of Quality Category | Master Data Example | How to Count It |
|---|---|---|
| Prevention | Data ownership, entry standards, validation rules at creation | Current spend (often near zero) and proposed spend |
| Appraisal | Manual cross-checks before submissions, periodic reconciliations | Hours per check, times checks per year |
| Internal failure | Corrections, deviations and OOS investigations caused by data, rejected electronic files | Hours per event, times events per year, plus any scrapped material |
| External failure | Information requests, delayed approvals, distributor exceptions, labeling corrections after release | Events per year, days of delay, and the revenue or penalty value of that delay |
ASQ also reports that, according to its 2025 cost of quality research, only 31% of respondents feel they fully understand the impact of quality costs on their organization’s financial performance13. Master data is one of the least understood parts of that picture, which is why measured results from a pilot are persuasive.
The Measurement Method Behind All Four Pilots
Each pilot uses a version of the Friday Afternoon Measurement described by Nagle, Redman, and Sammon. In their method, managers assemble 10 to 15 critical data attributes for the last 100 units of work their department completed, mark the obvious errors in each record, and count the error-free records. The count, from 0 to 100, is the Data Quality score: the share of records created correctly the first time10.
The method works for master data because it is fast, it uses real records, and the people doing the checking are the people who use the data. It also produces a number that a non-specialist can understand without a briefing. The four pilots adapt it in one important way: instead of checking records against the judgment of the team, they check records against each other across systems, and against the approved regulatory source where one exists.
Pilot 1: Product Master Source-of-Truth Check
The Question It Answers
For our marketed products, do the core product attributes agree across the regulatory, supply, labeling, and serialization systems, and when they disagree, which system is right?
This is the pilot to start with if your organization has had labeling corrections, serialization exceptions, or last-minute submission rework in the past year. It is also the pilot that most often surprises leadership, because most executives assume the product list is the one dataset everyone agrees on.
How to Run It
Pick the Sample
Select 30 to 50 marketed product presentations (a presentation is one strength in one pack). Include a mix of long-established products and recent launches or packaging changes, since changes are where drift starts.
Choose 10 to 15 Critical Attributes
Typical choices: product name, strength, dosage form, pack size, national product code, GTIN, marketing authorization or application number, holder, manufacturing and packaging sites, shelf life, and storage condition.
Pull Each Attribute From Each System
Extract the values from the regulatory information management system, the ERP, the labeling system or artwork records, and the serialization system. Put them side by side in one sheet, one row per presentation.
Mark Every Disagreement
A presentation passes only if every critical attribute agrees across every system. Record which attribute failed and, after checking the approved source, which system held the wrong value.
Trace Consequences
For each failed presentation, search the past 12 to 24 months of deviations, change controls, complaints, and distributor exceptions for events linked to the wrong attribute.
What to Measure
- Agreement score: the share of presentations where every critical attribute matches across every system.
- Error pattern: which attributes and which systems account for most failures. In practice, a small number of attributes usually explains most of the disagreement.
- Resolution effort: the hours it took the pilot team to find the correct value for each disagreement. This is a direct measure of the appraisal work people do every time they need the data.
- Linked events: the number of deviations, labeling corrections, and electronic data exceptions in the lookback period that trace to a product master error, with the hours or spend recorded for each.
What Good Output Looks Like
A strong result from this pilot is a one-page summary that says, in effect: “Of 40 presentations checked, this many had at least one disagreement. The most common failures were in these three attributes. The ERP was the wrong system in most cases. Resolving each disagreement took this many hours on average. In the past 18 months, this many events traced back to these errors.” Every number in that sentence came from your own records, and every one of them can be measured again after a fix.
The HDA guidance offers one reason this pilot matters for supply. It advises manufacturers to check their change process for packaging changes and new product launches so that changes reach distributors before product is sent, and to send master data to downstream trading partners before any product is sent7. A product master pilot that finds disagreement between the ERP and the serialization system is finding the exact conditions that lead to rejected files.
Pilot 2: Customer and Site Master Reconciliation
The Question It Answers
Do we have one correct record for each manufacturing site, testing lab, supplier location, and customer delivery location, with the right identifiers, and do our systems agree about them?
Site data is more complex in pharma and biotech than in most industries because the same physical location carries different identifiers for different purposes. A contract manufacturer might have an FEI number and DUNS number for FDA, organization and location identifiers in EMA’s OMS, a vendor number in the ERP, a supplier qualification record in the QMS, and a GLN for product traceability. Mergers, acquisitions, renamed sites, and changed addresses all break these links.
Two Versions of the Pilot
Regulated Site Reconciliation
Start from the list of manufacturing, packaging, and testing sites named in your applications. For each, compare the legal name, address, FEI, DUNS, and EMA OMS organization and location identifiers across the regulatory system, the QMS approved supplier list, and the ERP vendor master. Best for companies with frequent CMO changes or site transfers.
Customer and Trading Partner Reconciliation
Start from the top 50 to 100 customer delivery locations by volume. Look for duplicate records, conflicting GLNs, outdated addresses, and accounts linked to the wrong parent. Best for companies with commercial products in distribution and a history of shipment or data exceptions.
How to Run It
For Version A, the source list is the set of establishments named in your active applications. FDA’s question-and-answer guidance on this topic describes the facility information expected in Form FDA 356h and Module 3 and explains why FEI and DUNS numbers matter for the facility evaluation6. EMA’s Organisation Management Service (OMS) holds organization names and location addresses with unique identifiers, and EMA asks companies to submit change requests when their data needs updating14. Your pilot compares your internal records to these external sources as well as to each other.
For Version B, pull the customer master from the ERP and group records by normalized address. Two records at the same street address with different names or different GLNs are a candidate duplicate. The RAPS report of the HDA webinar describes this exact pattern: two sets of codes for one hospital, one under its abbreviated name and one under its full name8. The same report quotes a distributor’s operations leader noting that his company handles tens of thousands of stock-keeping units (SKUs), and that one SKU carrying multiple GTINs becomes a problem.
What to Measure
- Duplicate rate: the share of physical locations represented by more than one active record.
- Identifier completeness: the share of regulated sites with FEI, DUNS, and (where relevant) EMA OMS identifiers recorded and matching.
- Cross-system agreement: the share of sites where name and address agree across the regulatory, quality, and ERP systems.
- Linked events: information requests about facility data, supplier qualification gaps found at audit, rejected electronic shipment files, and misdirected deliveries in the lookback period.
Why this pilot often earns back its effort quickly: a single delayed application or a single held shipment can carry more financial weight than a year of data stewardship. When the pilot can connect even one such event to a site master error, the business case becomes much easier to make.
Ownership Findings Matter as Much as Error Counts
Site master pilots almost always reveal an ownership gap. Regulatory affairs owns the application, quality owns supplier qualification, procurement owns the vendor record, and supply chain owns trading partner data. Nobody owns the fact that all four describe the same place. McKinsey’s example of a global pharmaceutical company that cut data errors by more than 90 percent is relevant here: that company set up a central data management organization, defined clear master data management processes, and assigned ownership across data management and data owner functions9. Record who currently maintains each field during the pilot. The gaps you find become part of the proposal.
Pilot 3: Reference Data Inventory
The Question It Answers
How many versions of the same code lists are in use across our systems, who owns each one, and how much manual work goes into mapping between them?
This pilot is different from the other three. It does not check individual records. It checks the lists of allowed values that records depend on. Reference data problems rarely cause a single dramatic failure. Instead, they make it impossible to combine data from different sites or systems without a person translating codes by hand, and they make trend analysis misleading because the same event is classified differently in different places.
Where to Look
Start with the lists that most often diverge in pharma and biotech companies:
- Units of measure (for example, whether micrograms appear as “mcg”, “ug”, or “µg” in different systems)
- Dosage forms and routes of administration
- Storage conditions and temperature ranges
- Country and region codes
- Deviation, complaint, and CAPA categories and root cause codes
- Material types, container types, and status codes
- Test names and method codes in LIMS
EMA’s referentials service is a useful benchmark here. It maintains controlled lists of terms, such as dosage forms and units of measurement, that are used to describe product attributes in EU regulatory processes3. If your internal lists for the same concepts do not map cleanly to these, every regulatory data submission will require translation work.
How to Run It
Build a simple inventory with one row per code list per system. For each row, record the list name, the system, the number of values, the owner (if any), when it was last changed, and whether it is mapped to an external standard. Then compare lists that describe the same concept. Count the values that exist in one list but not the other, the values that mean the same thing but are spelled differently, and the values that look the same but mean different things.
Next, find the mapping work. Ask each reporting, analytics, and regulatory team where they translate one set of codes into another. The answers are often spreadsheets or lookup tables that one or two people maintain by hand. Record how many exist, who maintains them, and roughly how many hours per month they take.
What to Measure
- List count: number of distinct lists in use for each concept (for example, four different unit of measure lists across ERP, LIMS, MES, and RIM).
- Ownership gap: share of lists with no named owner.
- Divergence: share of values in each list that have no exact match in the other lists for the same concept.
- Manual mapping effort: hours per month spent maintaining translation tables, and the number of reports that depend on them.
Why reference data matters for AI and analytics: any model or dashboard that combines deviation, batch, or product data across sites depends on consistent categories. If one site classifies an event as “equipment” and another as “facility,” trending will split one real problem into two small ones. A reference data inventory is often the fastest way to show leadership why an analytics or AI investment is not yet delivering what was expected.
Turning the Inventory Into a Dollar Figure
Reference data is the hardest domain to price directly, because its damage is spread thinly. Two measures work well. First, the manual mapping hours, converted to money at a loaded labor rate, give a recurring appraisal cost. Second, pick one recent decision that relied on cross-site trending (for example, a CAPA effectiveness review or an annual product review) and check whether inconsistent categories changed the result. One documented case of a trend that was wrong because of reference data is often more persuasive than any hourly estimate.
Pilot 4: Specification Cleanup
The Question It Answers
Do the specifications configured in our LIMS and ERP match the approved specifications, and how many investigations and holds in the past year were caused by the difference?
This is the most GMP-critical of the four pilots, and the one that most directly connects master data to deviations. ICH Q6A gives this definition of a specification: “A list of tests, references to analytical procedures, and appropriate acceptance criteria that are numerical limits, ranges, or other criteria for the tests described”15. Each part of that definition is master data: the test, the method reference, the limit, the unit, and the reporting rule. Each part can drift between the approved source and the system that applies it.
Where Specifications Drift
In most companies, a specification exists in at least three places. The approved version is in the regulatory dossier. The controlled document is in the QMS. The version that decides whether a result passes or fails is the one configured in the LIMS, and sometimes a fourth version exists in the ERP as an inspection plan or material specification. Changes approved through a variation or supplement must flow through change control to every one of these. When one step is missed, the systems disagree.
Common drift points include:
- A limit changed after a post-approval change, updated in the QMS document but not in LIMS
- Rounding or reporting rules configured differently from the approved specification
- Units that differ between the specification and the LIMS configuration
- An analytical method version in LIMS that no longer matches the method referenced in the specification
- Specifications for a site or market that were copied from another and never fully adjusted
How to Run It
Select Products With Recent Changes
Choose 10 to 20 drug substance and drug product specifications, weighted toward products with post-approval changes, site transfers, or new markets in the past three years.
Line Up Every Version
For each test, place the approved limit, the QMS document limit, the LIMS limit, and any ERP value side by side, along with units, reporting rules, and method references.
Classify Each Mismatch
Mark each difference as tighter than approved (risk of false failures), looser than approved (risk of missed failures), or administrative (method version, unit label, or format with no effect on the pass or fail decision).
Search Investigation History
For each tighter-than-approved mismatch, count OOS results and investigations in the lookback period that fell between the LIMS limit and the approved limit. For each looser-than-approved mismatch, escalate immediately through your quality system.
Correct Through Change Control
Any correction to a GMP specification record goes through the normal change control and validation process. The pilot finds and prices the problem. It does not bypass the controls that govern the fix.
Treat looser-than-approved findings as quality events, not pilot data. If the pilot finds a LIMS limit that would pass a result the approved specification would fail, that finding goes straight to the quality unit for assessment, including whether released batches are affected. The pilot should have a pre-agreed escalation route before it starts, so that nobody hesitates when this happens.
What to Measure
- Specification agreement rate: the share of tests where every system matches the approved specification.
- Mismatch direction: counts of tighter, looser, and administrative mismatches.
- Avoidable investigations: OOS investigations in the lookback period that were caused by a tighter-than-approved configuration, with hours and any batch hold time recorded.
- Change control leakage: the share of approved specification changes in the lookback period that did not reach every system.
The last measure is often the most useful for the business case, because it describes a process failure rather than a set of bad records. FDA’s OOS guidance expects investigations to determine whether a failure is associated with other batches or products5. A specification pilot applies the same logic to data: if one specification failed to update, how many others followed the same path?
Turning Pilot Results Into a Business Case
Four pilots give you four sets of measured numbers. The final step is to combine them into a proposal that finance can check and quality can support.
Four Numbers per Pilot
For each pilot, report the same four numbers so leadership can compare them:
- Error rate: the share of records that failed the check (100 minus the Data Quality score).
- Volume: how many records of this type are created or changed per year, and how many exist in total.
- Unit effort: the measured hours to find and fix one error, plus the hours each downstream event consumed.
- Consequence events: the number of deviations, investigations, exceptions, information requests, or delays traced to this domain in the lookback period, with the documented financial impact of each where one exists.
A Simple Annual Estimate
The annual figure for each domain has two parts. The first is recurring correction work: records created per year, times the error rate, times the hours to fix each error, times a loaded labor rate. The second is consequence events: the number of events per year, times the average hours or spend per event. Add the appraisal work you measured, such as the manual cross-checks before each submission, and you have a defensible annual figure for that domain.
Worked example (hypothetical numbers, for illustration only): suppose a company creates or changes 2,000 material master records a year, and its pilot shows that 15% have at least one critical error. If each error takes an average of three hours to find and fix at a loaded rate of $100 per hour, recurring correction work is 2,000 × 0.15 × 3 × $100 = $90,000 a year. If the lookback also found six deviations caused by material master errors, averaging 40 hours each, that adds 6 × 40 × $100 = $24,000. The domain total is $114,000 a year before counting any batch holds, scrap, or delay. Replace every number here with your own pilot results.
Present each total as a range, not a single number. Use the measured values as the lower bound and apply the rule of ten from the HBR study as a sensitivity check for the upper bound10. Finance teams trust a range they can see the basis for more than a precise figure they cannot.
What to Ask For
The pilots also tell you what to ask for, and it is often less than a platform. Match the request to what the pilots found:
| If the Pilots Show | Ask For |
|---|---|
| Errors concentrated in a few attributes or one system | Validation rules at the point of entry and a named owner for those attributes |
| Disagreement between systems with no clear source of truth | A decision on the authoritative system for each domain, and interfaces or controlled processes that copy from it |
| Change control leakage for specifications or products | A change control checklist that names every system a change must reach, with verification before closure |
| Many local code lists and heavy manual mapping | A small reference data function with ownership of shared lists and mapping to external standards |
| High volume, many domains, many systems | A phased master data management program, using the pilot results as the baseline and requirements |
McKinsey’s work on planning system deployments points to the value of timing: companies that start structured data preparation three months before detailed design often cut remediation effort by half and speed up go-live9. If a major ERP, LIMS, or planning project is on your roadmap, the pilots are best run before it starts, and their findings belong in the project’s scope and budget.
GxP Controls Still Apply
Master data in a drug company is not ordinary business data. It falls under GxP, the family of good practice regulations that includes GMP. Product, material, and specification records feed directly into the master production and control records that 21 CFR 211.186 requires2. Any correction the pilots recommend must go through change control, and any system change must follow your computer system validation approach. Build this into the business case. The effort of validated correction is part of the investment, and hiding it only creates a second budget request later.
What success looks like after 90 days: four one-page pilot summaries with measured error rates and linked events, a combined annual estimate presented as a range, a named owner for each domain, a short list of fixes matched to the findings, and a date to run the same measurements again. That last item is what turns a one-time finding into a program leadership can track.
Measure Again
The strongest feature of this approach is that it can be repeated. Each pilot is a measurement you can run again six months after a fix, using the same sample design and the same attributes. When the second measurement shows the error rate falling and the linked events declining, the business case for the next phase is much easier to make. Nagle, Redman, and Sammon make a similar point: eliminating a single root cause can prevent thousands of future errors10.
Conclusion
Bad master data causes real, measurable damage in pharma and biotech: correction work spread across every function, deviations and investigations that trace back to a wrong record, and delays in submissions and supply. The reason this rarely gets funded is not that leaders do not care. It is that the evidence usually arrives as an industry average, a vendor benchmark, or an anecdote, and none of those can be checked against the company’s own books. Four small pilots, run on your own records with the systems you already have, change that. They give you error rates, hours, and events you can defend, and they point directly to the fixes that matter most.
Sakara Digital works with pharma and biotech organizations on data quality and governance in GxP settings, including the measurement work that comes before a master data investment. If you are trying to build a business case for master data quality and want an independent perspective on where to start and how to scope the first pilot, we are happy to have that conversation.
For Further Reading
For Further Reading
- Master Data Management for Life Sciences: Creating a Single Source of Truth Across Global Operations
- The Cost of Poor Data Quality in Pharma Manufacturing: A 2026 Benchmark
- Reference Data Governance in Pharma: UNII, MedDRA, SNOMED, and Version Drift
- Beyond the SOP Index: Four Data Quality Pilots for Document Management
- The Data Quality Debt Audit Pattern We Run With Clients
References & Sources
- Gartner. “Master Data Management.” Gartner Data and Analytics Topics, accessed September 2026. https://www.gartner.com/en/data-analytics/topics/master-data-management
- U.S. Code of Federal Regulations. “21 CFR 211.186: Master Production and Control Records.” Legal Information Institute, Cornell Law School. https://www.law.cornell.edu/cfr/text/21/211.186
- European Medicines Agency. “Substance, Product, Organisation and Referential (SPOR) Master Data.” EMA, accessed September 2026. https://www.ema.europa.eu/en/human-regulatory-overview/research-development/data-medicines-iso-idmp-standards-overview/substance-product-organisation-referential-spor-master-data
- Redman, Thomas C. “Bad Data Costs the U.S. $3 Trillion Per Year.” Harvard Business Review, September 22, 2016. https://hbr.org/2016/09/bad-data-costs-the-u-s-3-trillion-per-year
- U.S. Food and Drug Administration. “Investigating Out-of-Specification (OOS) Test Results for Pharmaceutical Production: Guidance for Industry.” Revision 1, May 2022. https://www.fda.gov/media/158416/download
- U.S. Food and Drug Administration. “Identification of Manufacturing Establishments in Applications Submitted to CBER and CDER: Questions and Answers.” Guidance for Industry, Revision 1, October 2019. https://www.fda.gov/media/131911/download
- Healthcare Distribution Alliance. “Exceptions Handling Guidelines for the DSCSA.” HDA, April 2022. https://hda.org/getmedia/bffcc1e6-8d5d-4fe5-b1c0-cd011ebd675b/HDA-Exceptions-Handling-Guidelines-for-DSCSA.pdf
- Regulatory Affairs Professionals Society. “Expert: Industry Needs ‘Meticulous’ Control Over Master Data Under DSCSA.” RAPS Regulatory Focus, June 2024. https://www.raps.org/resource/expert-industry-needs-meticulous-control-over-mas.html
- McKinsey & Company. “The Quiet Enabler: Data Management Best Practices for APS Deployments.” McKinsey Operations Insights, April 7, 2026. https://www.mckinsey.com/capabilities/operations/our-insights/the-quiet-enabler-data-management-best-practices-for-aps-deployments
- Nagle, Tadhg, Thomas C. Redman, and David Sammon. “Only 3% of Companies’ Data Meets Basic Quality Standards.” Harvard Business Review, September 11, 2017. https://hbr.org/2017/09/only-3-of-companies-data-meets-basic-quality-standards
- Gartner. “Data Quality: Why It Matters and How to Achieve It.” Gartner Data and Analytics Topics, accessed September 2026. https://www.gartner.com/en/data-analytics/topics/data-quality
- Redman, Thomas C. “Seizing Opportunity in Data Quality.” MIT Sloan Management Review, November 27, 2017. https://sloanreview.mit.edu/article/seizing-opportunity-in-data-quality/
- ASQ. “What Is Cost of Quality (COQ)?” ASQ Quality Resources, accessed September 2026. https://asq.org/quality-resources/cost-of-quality
- European Medicines Agency. “Organisation Management Service (OMS).” EMA, accessed September 2026. https://www.ema.europa.eu/en/human-regulatory-overview/research-development/data-medicines-iso-idmp-standards-overview/substance-product-organisation-referential-spor-master-data/organisation-management-service-oms
- U.S. Food and Drug Administration. “International Conference on Harmonisation; Guidance on Q6A Specifications: Test Procedures and Acceptance Criteria for New Drug Substances and New Drug Products: Chemical Substances.” Federal Register 65, no. 251, December 29, 2000. https://www.govinfo.gov/content/pkg/FR-2000-12-29/html/00-33369.htm








Your perspective matters—join the conversation.