The Emerging Consensus, and Where It Runs Out

Industry groups have converged on a reasonable way to begin. The ISPE Pharma 4.0 community has been advancing a version of it through the “ISPE Good Practice Guide: Pharma 4.0 Holistic Digital Enablement,” which treats digital enablement as a cross-functional operating model connecting technology with knowledge management, data governance, interoperability, and workforce culture rather than as a set of tools.3

The practical advice that comes out of that work is good advice. Secure senior sponsorship. Name data owners and data stewards, and make sure they are empowered to decide things, because a tool will not hold without accountability behind it. Start with a baseline assessment. Pick one small thing rather than attempting the whole estate. Build the catalog and the lineage. Choose a few policies. Define measures.

None of that is wrong. The gap is in sequence rather than content.

Consider what an organization holds at day ninety if it follows that order faithfully: a chartered council, a populated catalog, a documented access model, three approved policies, and a measure set. Every one of those is a real asset. Not one of them yet answers the question a finance or operations leader will ask, which is what changed as a result.

The failure mode this article is trying to prevent

Governance programs rarely fail because the framework was wrong. They fail at the funding boundary between the first phase and the second, when the team has infrastructure to show and no outcome. Reordering the first thirty days fixes that without discarding anything.

Why the Quality System Is the Right Starting Point

Three reasons, and the third is the one that gets budget approved.

The Data Is Already Governed, So You Are Correcting Rather Than Creating

Quality system data already has retention rules, access controls, defined record owners, and change control around the systems that hold it. Compare that to commercial or research data, where a governance program often has to establish those things from nothing. Starting in the quality system means extending and correcting controls that exist, which is faster and far less disruptive.

The Failures Are Visible to People Who Do Not Work in Data

A quality director can see that two sites classify the same deviation differently. A document owner can see that a procedure has been past its review date for fourteen months. These are legible problems. An argument about data architecture is not, and the difference matters when you need a decision from someone outside the program.

The Inspection Record Carries the Argument

This is the part that changes the funding conversation, and it is worth going through the numbers carefully.

What the Inspection Record Supports, and What It Does Not

FDA publishes an annual inspection observations dataset counting Form 483 observations by citation.1 In fiscal year 2025, covering inspections that ended between October 2024 and September 2025, FDA recorded 2,837 drug CGMP observations across 316 cited provisions, drawn from 713 drug Form 483s.

243 21 CFR 211.22(d), the most cited provision of the year: quality control unit responsibilities and procedures not in writing, or not fully followed
164 21 CFR 211.192, failure to thoroughly review any unexplained discrepancy or batch failure
162 21 CFR 211.100(a), absence of written procedures for production and process control

Rolled up to the section level, 211.22 accounted for 309 observations, roughly 11 percent of every drug observation recorded that year, which makes the quality unit the most cited section in the dataset. Section 211.192 accounted for 236, and within it a further 32 observations cited an incomplete written record of the investigation specifically. Laboratory controls under 211.160(b) drew 121. Training under 211.25(a) drew 67.

Read those together. The most common findings in pharmaceutical manufacturing are not about equipment, facilities, or chemistry. They are about whether procedures exist, whether they are followed, and whether the record demonstrates it. That is a data problem carrying a compliance label, and it is the argument for starting data work inside the quality system rather than beside it.

The same pattern holds outside drug manufacturing, which is useful corroboration that this is structural rather than a quirk of one program area. In the device program area, across 2,660 observations from 791 Form 483s, the most cited provision was 21 CFR 820.100(a), corrective and preventive action procedures lacking or inadequate, at 279, with CAPA documentation under 820.100(b) drawing a further 63.1

Three things this dataset does not say, and you will hear all three

FDA’s citation database contains no citation for “data integrity,” no citation for “audit trail,” and no citation for “quality risk management.” None of those phrases appears in the reference text.

Data integrity is a concept enforced through specific CGMP provisions rather than a regulation you can be cited against. ICH Q9 is guidance rather than codified United States regulation, so it cannot appear on a Form 483 at all. The closest provision to the audit trail concept is 21 CFR 211.68, computerized systems, which accounted for 153 observations across the year, ranking seventh by section. Within it, 211.68(b) accounted for 110, and its largest single line item, at 87, reads that appropriate controls are not exercised over computers or related systems to assure that changes to records are instituted only by authorized personnel.

If you put inspection data into a business case, cite the provision and the count rather than the industry shorthand. The specifics are more persuasive, and they survive a challenge from someone who knows the dataset. One further caveat worth carrying: FDA notes the dataset excludes manually prepared Form 483s, so these are counts from the system rather than a complete census.

Days 1 to 30: Find the Failure, Not the Framework

The conventional sequence builds governance first and applies it later. This version inverts the first month.

1

Choose One Dataset Somebody Already Wants to Use

In most quality systems the two best candidates are deviation and CAPA classification, or SOP metadata. Deviation data is the better choice when there is a trending or prediction ambition behind the request. SOP metadata is better when a system migration is coming, because the findings convert directly into migration scope.

2

Scope It Hard

One site, or one product, or one department. Not the enterprise. The purpose of the first thirty days is a finding, and a narrow scope produces a precise finding faster than a wide scope produces a vague one.

3

Measure the Actual Condition

For deviation and CAPA data, pull twelve months of records and count how categories are applied, how many root causes exist only in free text, and how consistently the same event type is classified. For SOP metadata, pull the full document register and count orphan documents with no current owner, documents past their review date, and documents referencing retired equipment or discontinued products.

4

Count. Do Not Estimate

A precise number from a bounded scope is more useful, and more defensible, than a rough figure from a wide one. An estimate invites a debate about the estimate. A count invites a debate about what to do, which is the conversation you want.

5

State the Finding in One Sentence

By day thirty the finding should fit in a sentence a non-specialist understands. Something in the shape of: across this site, 40 percent of deviation records classified as human error carry no structured root cause, which makes trending across sites impossible.

Notice what the first month does not include. No council charter. No catalog. No policy set. Those are month three, and they will be easier to write once there is a concrete problem to write them about.

Choosing Between the Two Candidate Datasets

Deviation and CAPA classification, or SOP metadata. Both are defensible starting points and the choice should follow from what the organization is already trying to do rather than from which looks worse.

Choose deviation and CAPA classification when someone has already asked for trending, prediction, or a cross-site comparison. The failure is easy to demonstrate, the business interest already exists, and the fix produces a dataset people were asking for. The tradeoff is that the remediation touches a field many people use daily, so the change control and retraining effort is larger.

Choose SOP metadata when a system migration or a document system consolidation is on the roadmap within a year. The findings convert directly into migration scope, which means the assessment pays for itself in a decision that is already scheduled. The remediation is also gentler: assigning owners and clearing overdue reviews does not change how anyone enters data. The tradeoff is that the result is less visible to operations, so the case for phase two rests more on the migration than on analytics.

If both apply, start with SOP metadata. It produces a result faster, carries less change control weight, and builds the credibility you will want before asking people to change how they classify deviations.

What the Baseline Assessment Should Actually Contain

Most organizations already own the method for this and do not recognize that they do. The data integrity assessment described in the MHRA, PIC/S, and World Health Organization guidance is a structured gap exercise, and its shape transfers directly to a data quality assessment.789 FDA’s data integrity questions and answers covers the same territory for drug CGMP.6

The sequence those documents describe runs: assess system by system, identify the gaps, evaluate each gap for criticality, put a remediation plan in place, execute it, and check effectiveness. That is exactly the sequence proposed here. The difference is only what you are assessing against. A data integrity assessment asks whether records can be trusted. A data quality assessment asks whether they can be used together.

Run both on the same dataset and keep the results separate. In a mature organization the integrity assessment usually comes back largely clean, and that clean result is what makes the second assessment persuasive rather than merely negative. It preempts the objection you will otherwise meet, which is that the data is already compliant and therefore already fine.

Three things belong in the assessment record regardless of which dataset you choose:

  • The population you actually examined. Which site, which date range, how many records, and what you excluded. A finding without a stated population cannot be defended or repeated.
  • The distinct values in use. For any field intended to hold a controlled list, the count of distinct values actually present. This single number carries more weight in a governance conversation than any narrative description of the problem.
  • The criticality grade and its rationale. Not just critical, major, or minor, but why. The rationale is what survives staff turnover.

Days 31 to 60: Grade It, Then Fix the Smallest Useful Piece

Prioritize With a Method Regulators Already Accept

ICH Q9(R1) gives you quality risk management as a structure, and using it means the prioritization is defensible rather than a matter of preference.2 For each gap found in month one, assess the likelihood of it causing a problem and the impact on product quality and patient safety. Grade the results as critical, major, or minor.

That grading does two things. It determines sequence, and it gives you the language to explain why one gap was fixed and another deferred. The second is more valuable than it sounds, because deferral is what a governance program spends most of its time doing and the reason is rarely recorded.

A note on scoring honestly

Resist the temptation to grade everything critical to secure funding. A risk register where every item is critical carries no information and will not survive review by anyone who does risk assessment for a living. The credibility of the exercise depends on some things being graded minor.

Remediate the Smallest Piece That Produces a Visible Result

For deviation data, this is usually harmonizing the category list and adding structured root cause capture to the intake form. For SOP metadata, it is assigning owners to the orphan documents and clearing the overdue reviews.

Both of these are change control exercises with training implications, which is worth planning for rather than discovering. Changing a controlled vocabulary in a validated system is not a configuration change you make on a Tuesday afternoon. Build the change control and the training update into the sixty-day window from the start.

Design the Fix Wider Than the Assessment

This is the step that prevents the most common criticism of a narrow pilot, which is that it creates another silo. Scope the assessment narrowly and design the remedy broadly. Harmonize the category list with the other sites in mind even if you only apply it at one. Write the metadata standard so it is adoptable elsewhere without rework. The assessment is local; the standard should not be.

The Change Control Problem Nobody Plans For

Harmonizing a controlled vocabulary in a validated quality system is a change to a validated system. That sentence causes more ninety-day plans to slip than any other single factor, because it is usually discovered in week seven.

Changing the picklist behind a deviation classification field means a change request, an assessment of what the change affects, a decision about whether the change needs testing and how much, updated procedures, and retraining of everyone who uses the field. None of that is optional and none of it is fast.

The useful move is to size the validation effort by risk rather than by habit. FDA’s computer software assurance guidance sets out a risk-based approach to exactly this question: establish the intended use, assess whether a failure would affect product quality or patient safety, and then apply assurance activities proportionate to that risk rather than applying the same scripted testing to everything.10

Applied to a vocabulary change, that reasoning usually produces a modest answer. Adding structured values to a classification field does not change how the system computes anything or how records are retained. The risk is in the data migration decision, meaning how existing records mapped to old categories are treated, and in whether users apply the new list consistently. Those are the two things worth testing. Re-validating the whole module is the habit, not the requirement.

Plan the mapping decision explicitly, and write it down

When a category list changes, every historical record carries an old value. You have three options: leave history on the old vocabulary and analyze the two separately, map history forward using documented rules, or map history forward using judgment. The third is the one that happens by default when nobody decides, and it is the one that makes the resulting dataset untrustworthy for trending. Decide deliberately, record the rule, and state the boundary date in the assessment record.

Naming Owners and Stewards So the Roles Hold

Every version of this advice says to name data owners and stewards. Fewer say what the difference is, and the distinction matters because the two roles fail in different ways.

The data management body of knowledge published by DAMA International draws the line roughly where practice does: accountability for a data asset and its business definitions is one role, and the day-to-day work of maintaining quality within those definitions is another.12 In a quality system that translates cleanly. The owner decides what a deviation category means and approves changes to the list. The steward monitors whether the category is being applied correctly and raises it when it is not.

Two practical tests of whether the appointment is real. First, can the named owner approve a change to the controlled vocabulary without escalating? If not, they are a contact rather than an owner, and the role will not function. Second, does the steward have a scheduled activity and a place to record findings? A steward without a recurring task is a title.

This is the point in the ISPE guidance that deserves the most emphasis, because it is the one most often implemented in name only: a tool will not hold unless the people named around it are accountable and empowered to decide things.3

Days 61 to 90: Build Governance Around What Worked

Now stand up the machinery, with a working example behind it.

Name Owners

Data owner and steward for the remediated dataset, with written accountabilities

Extend

Same ownership model applied to the two or three adjacent datasets now known to be affected

Write Policy

The policy that would have prevented the problem you just fixed

Define Measures

A small set tracked on the dataset, not an enterprise scorecard

Verify

Effectiveness check on the remediation, recorded

On measures, keep the set small and specific to what was fixed. Category consistency across sites. Percentage of records with a structured root cause. Documents past review date. Documents without an owner. Four numbers that move are worth more than twenty that nobody updates, and they are the numbers that make the case for phase two.

On policy, there is a real advantage to writing it last. A data governance policy written before anyone has examined the data tends to be generic, because it has to cover situations nobody has looked at. A policy written after a remediation can be specific about what good looks like, and specific policies are easier to audit against.

Who Needs to Be in the Room, and for How Long

A ninety-day sprint of this shape does not need a large team. It needs a small one with the right authority, available in short bursts rather than continuously.

RoleWhat they doRealistic time
AnalystPulls the records, counts the distinct values, produces the assessmentMost of month one, then part time
Quality lead for the chosen areaAnswers what a category was supposed to mean, decides the harmonized listA few hours a week throughout
Quality assurance representativeConfirms the change control path and what validation the change requiresConcentrated in month two
System administratorMakes the configuration change, supports the mapping decisionDays, in month two
Executive sponsorApproves the scope, receives the finding, decides on phase twoThree short meetings

The sponsor’s three meetings are worth naming, because the timing is the mechanism that keeps the work moving. One at the start to approve the scope and the dataset. One at day thirty to receive the finding, which is the meeting where the program either earns attention or does not. One at day ninety to see the effectiveness check and decide on the second phase.

What this team does not need is a chartered council. A council is the right structure for ongoing decisions across many datasets, which is a month-three-onward problem. Forming one in week one produces a group with nothing yet to decide.

The Four Measures Worth Tracking From Day One

Measure sets tend to grow until nobody maintains them. Four numbers, tracked on the chosen dataset only, are enough to demonstrate progress and to make the case for extending the work.

CONSISTENCY

Distinct Values in a Controlled Field

Count the distinct values actually present where one controlled list was intended. This is the single most persuasive number in the whole exercise, because the gap between intended and actual is usually large and requires no interpretation.

STRUCTURE

Share of Records With a Structured Root Cause

The percentage of deviation or CAPA records where the root cause exists in a structured field rather than only in narrative text. This is the number that determines whether trending is possible at all.

OWNERSHIP

Documents Without a Current Owner

A count, not a percentage, because the absolute number is what drives the remediation task list. Track it alongside how many were cleared this period so the trend is visible.

TIMELINESS

Documents Past Their Review Date

Already tracked in most quality systems, which makes it the easiest of the four to establish and a useful control: if this number is well managed and the others are not, that itself tells you something about where attention has gone.

Two rules about these measures. Report them against a stated population, so a change in scope cannot be mistaken for a change in performance. And report the new records separately from the historical ones once a fix is in place, because mixing them hides whether the process actually changed.

What Ninety Days Produces, Compared Side by Side

At day 90Conventional sequenceQuality-system-first sequence
Governance structureCouncil chartered, sponsor namedOwners and stewards named for one dataset and its adjacents
DocumentationCatalog populated, access model defined, three policies approvedOne policy, written from an observed failure
Evidence of conditionBaseline assessment, usually qualitativeCounted finding from a bounded scope
Risk positionNot usually addressed in phase oneFindings graded critical, major, minor under ICH Q9(R1)
Something actually fixedNot yetOne dataset remediated
Proof it workedNot yetEffectiveness check completed and recorded
Inspection valueIndirectAssessment, graded findings, remediation, effectiveness check

Both columns contain real work and the left column is not wasted. The point of the comparison is which column is easier to take into a conversation about funding a second ninety days, and that is the whole argument of this article.

The Effectiveness Check That Makes It Inspection-Ready

This is the step most data programs skip and every quality professional already knows how to do. Return to the remediated dataset, measure it again, and record whether the fix held.

Run it the way a CAPA effectiveness review runs. Define in advance what success looks like as a number. Wait long enough for new records to accumulate under the changed process. Then measure the new records, not the old ones, because the question is whether the process changed rather than whether the historical data was cleaned.

What that produces is a chain: an assessment, graded findings, a remediation, and evidence the remediation worked. That is the same objective evidence an investigator would ask for when testing whether an organization is in a state of control, and it was produced as a by-product of getting the data in order.

Why this matters more than it appears

The effectiveness check is what separates a data cleanup from a governance capability. A cleanup fixes records. A capability demonstrates that the organization can detect a data condition, act on it, and verify the action. The second is what a regulator is interested in, and it is also what makes the program defensible internally when someone asks whether the first ninety days accomplished anything.

Five Objections Worth Answering

“We Need Enterprise Scope or We Will Just Create Another Silo”

A real risk, and the mitigation is in how the fix is designed rather than how wide the assessment is scoped. Harmonize with other sites in mind, design the standard to be adoptable, and record the decisions so the next site inherits them. Scope the assessment narrowly and design the remedy broadly.

“Our Data Is Too Bad to Start”

That is the finding, and it is worth having in writing. A documented assessment showing the actual condition of deviation data is a more useful artifact than a governance charter written in the hope that the data turns out to be workable.

“We Should Wait for the New System”

Migrating unresolved data quality problems into a new platform moves them without fixing them, and the move has been paid for. Assessment before migration is one of the highest-return sequences available, because the findings become migration scope and cleanup criteria rather than problems discovered after go-live. Migration is also the point at which structure is most often lost, since fields without a clean destination get concatenated into notes under schedule pressure.

“This Is IT’s Job”

IT builds and runs the systems. The definitions, the categories, the controlled vocabularies, and the ownership have to come from the people accountable for the process. The FY2025 citation pattern points the same way: the most cited findings concern whether procedures exist and are followed, which is quality’s territory rather than IT’s.1

“Ninety Days Is Not Long Enough to Do This Properly”

Correct, and that is not the claim. Ninety days is long enough to produce a counted finding, a graded risk assessment, one remediation, and an effectiveness check on a bounded scope. It is not long enough to govern an enterprise. The purpose of the first ninety days is to earn the next ninety.

What Belongs in the Second Ninety Days

Assuming the first phase produced what it should, the second phase is where the conventional advice becomes exactly right. With a demonstrated result behind it, the catalog, the lineage work, and the broader policy set are much easier to fund and much easier to scope, because you now know what the data actually looks like.

EXTEND

Apply the Standard to the Next Two Sites

The harmonized vocabulary already exists and the change control path is already proven. This is the cheapest expansion available and it is where the aggregate analysis becomes possible.

CATALOG

Build the Catalog Around Known Datasets

A catalog populated after an assessment records what the data is and what condition it is in. A catalog populated before one records only what exists.

LINEAGE

Trace the Datasets You Now Care About

Prioritize lineage for the systems that anchor regulated workflows: the manufacturing execution system, the laboratory information management system, the enterprise resource planning system, and the learning management system.

GOVERN

Charter the Council With Real Decisions in Front of It

A council formed to approve a vocabulary change for three more sites has a purpose. A council formed before there is anything to decide tends to meet twice and lapse.

One forward-looking note on why the sequence matters now rather than eventually. The draft EU Annex 22 on artificial intelligence, still in draft as of this writing and not in force, with EMA’s inspectors working group having convened a multistakeholder workshop on 30 June and 1 July 2026 to gather expert input before finalization,11 requires that test data be representative of and expand the full sample space of the intended use, that it be stratified and include all subgroups, and that the criteria and rationale for its selection be documented.45 An organization whose deviation categories differ across sites cannot demonstrate that a dataset includes all subgroups, because it cannot define the subgroups consistently. The data work and the regulatory work are the same work.

Conclusion

The ninety-day structure that industry groups are converging on is sound, and nothing here argues against its components. Sponsorship, ownership, catalogs, lineage, policies, and measures all belong in a mature program.

The change worth making is the order. Lead with a bounded assessment of a quality system dataset that has a visible failure. Grade it with a risk method regulators already accept. Fix the smallest piece that produces a result. Verify the fix held. Then build governance around the thing that worked.

The destination is the same. The difference is that you are considerably more likely to still be funded when you arrive.

For Further Reading