The Eighteen-Month Deadlock

The pattern is consistent enough that you can predict the meeting before you walk into it.

A commercial or manufacturing leader brings a proposal. There is a decision made hundreds of times a year, by people, slowly, with uneven judgment, and a model could support it. The proposal has a number attached. Everyone in the room agrees the number is plausible.

Then the head of data speaks. The systems that hold the relevant records were implemented at different times by different teams. Site names are entered three ways. Two of the four plants moved to a new execution system in 2023 and the older records were converted with known gaps. Nobody owns the reference lists. Building a model on this would produce answers that are confidently wrong, and in a regulated setting a confidently wrong answer is worse than no answer. The recommendation is to fix the foundation first.

Then finance speaks. The foundation program has a seven-figure estimate, a three-year horizon, and no attributable return. It cannot be approved on its own merits. Bring back a use case that produces value and the enabling work can ride along with it.

Every person in that room is being reasonable. The data leader is describing a real risk. The business leader is describing a real opportunity. Finance is applying the same test it applies to everything else. And the meeting ends with an action item to do more analysis, which is how the next eighteen months go.

What makes this expensive is not the delay by itself. It is that the delay is invisible on any dashboard. No project failed. No budget was overspent. Nothing was canceled, because nothing was started. The company simply arrives at the following year with the same conversation, a slightly worse data situation because two more systems were added in the meantime, and a leadership team that has grown skeptical of both the AI proposal and the data proposal.

92% of the 53 AI practitioners interviewed in a study of high-stakes domains reported data cascades, meaning compounding downstream effects traced back to data problems7
5 data quality dimensions defined in the EMA and HMA Data Quality Framework: reliability, extensiveness, coherence, timeliness, and relevance1
112% the upper bound of bias introduced in a published simulation when record linkage error varied between hospitals rather than occurring at random9

Research on why AI projects fail tends to get quoted in these meetings by both sides. RAND interviewed 65 data scientists and engineers with at least five years of experience building models in industry or academia, and reported five leading root causes of failure. The first was that stakeholders misunderstand or miscommunicate the problem to be solved. The second was that the organization lacks the data needed to train an effective model.6 The data team hears the second finding and reads it as proof that starting without a foundation fails. The business team hears the first and reads it as proof that data programs disconnected from a real problem go nowhere. Both readings are supportable, which is exactly the problem.

What is worth noticing is that those two findings are not in tension. They describe one failure, which is a data effort and a problem definition that were never connected to each other. That connection is what the rest of this article is about.

Why Both Positions Are Wrong

“Fix the data first” is wrong because “the data” is not a thing

The fix-first position treats data quality as a property of a company. It is not. It is a property of a dataset relative to a question, and the regulators have already said so plainly.

The Data Quality Framework for EU medicines regulation, adopted by the CHMP in October 2023, defines data quality as fitness for purpose for users’ needs in relation to health research, policy making, and regulation, together with the requirement that the data reflect the reality they aim to represent.1 The framework then builds five dimensions on top of that definition: reliability, extensiveness, coherence, timeliness, and relevance. Two of those are explicitly question-dependent. Relevance is defined as the extent to which a dataset holds the elements useful to answer a given research question. And the framework is direct about reliability as well, noting that while reliability itself is independent of a specific question, each question sets its own threshold for acceptable reliability.1

Read that again, because it dismantles the fix-first position from a source no one in the room will argue with. There is no such thing as data that is clean. There is only data that is sufficient for a stated purpose. A company cannot finish cleaning its data any more than it can finish being ready. It can only make a specific dataset sufficient for a specific decision.

FDA’s guidance on real-world data and real-world evidence reaches the same place through different language, assessing fitness for use through relevance and reliability, where relevance covers whether the key data elements are available and whether there are enough representative subjects, and reliability covers accuracy, completeness, provenance, and traceability.4 Again the standard is set by the use, not by the source. The EFPIA commentary on implementing the European framework makes the practical version of the same point: assessment has to be anchored to the regulatory question being asked, or it becomes an unbounded exercise.2

When a data team says the data is not ready, the correct response is not agreement or disagreement. It is a question: not ready for what?

“Just start with AI” is wrong because data dependencies are the expensive kind

The other position fails for a different reason. Starting a use case on data you have not characterized does not avoid the data work. It defers the data work to the worst possible moment, which is after a model is in production and people have started relying on it.

The clearest research statement of this is the work on data cascades by Sambasivan and colleagues, who interviewed 53 AI practitioners working in high-stakes domains including medical applications. They define a data cascade as a compounding event that produces negative downstream effects from an unaddressed data problem, and they report that these cascades affected 92 percent of the practitioners studied. Their central observation is that the effects are delayed and opaque: the data problem is created early, the damage appears late, and by then it is expensive to trace.7 A model built on a feed whose meaning changes when a source system is upgraded will keep producing numbers after the change. It will just produce different numbers, for reasons nobody in the room can explain, at a moment when a quality investigation is open and someone is asking why the output moved.

In a regulated environment that failure mode is not merely embarrassing. It is a change control problem, a data integrity problem, and an inspection problem at the same time. The work you skipped comes back with interest and with an auditor attached.

The framing that resolves the argument. “Fix the data first” and “just start with AI” are both answers to the wrong question. The question is not whether data work comes before AI work. It is: for this specific decision, what is the minimum condition the data has to be in before a model supporting it can be trusted, and how much of the distance to that condition is already covered?

The Use Case Is the Unit of Planning

Once you accept that data quality is defined relative to a purpose, the planning unit follows. Neither a data program nor an AI program is the right container. The use case is, because it is the smallest object that carries both a value estimate and a data requirement.

This has three consequences that change how the portfolio is managed.

CONSEQUENCE 1

Data requirements become finite and arguable

“Our master data is a mess” cannot be scoped, estimated, or finished. “Site identifiers must resolve across the trial management system and two vendor feeds for the 41 sites active since January 2023” can be scoped, estimated, and finished, and someone can disagree with it in a way that produces a better answer.

CONSEQUENCE 2

Foundational work gets a beneficiary

Reference data governance and identifier resolution are hard to fund on their own. Tied to three named use cases with dated commitments, they stop being infrastructure and become a dependency with a schedule attached, which is a thing organizations know how to approve.

CONSEQUENCE 3

Sequencing replaces prioritization

Ranking use cases by value alone produces a list where the top item is blocked. Ranking by value against the effort to reach the minimum data condition produces an order in which each item is startable when its turn arrives.

CONSEQUENCE 4

Deferral becomes a normal outcome

When the analysis is per use case, saying that one of them is not feasible this year is a routine finding rather than a verdict on the whole program. That is what makes the honest answer sayable, which is the only way it ever gets said.

None of this requires a new methodology. It requires that the conversation move down one level of specificity, from programs to individual decisions, and stay there.

How to State a Minimum Data Condition

A minimum data condition is a written statement of what has to be true about the data before a specific use case can be built and trusted. It is not a data quality assessment. It is much narrower, and it is written before any remediation is scoped, because its purpose is to bound the remediation.

Five components. Each one answers a question that otherwise gets settled by whoever speaks with the most confidence.

1

Which fields, from which system of record

Name the specific fields, not the systems. “Deviation short description, deviation classification at closure, product code, equipment identifier, investigation outcome category, closure date.” For each one, name the system that is authoritative when two systems disagree. This step alone usually shrinks the perceived scope by a large margin, because most of the fields people worry about are not needed for the decision at hand.

2

At what completeness, per field

Completeness is measured per field, not per dataset, and the threshold differs by field. The European framework separates completeness (how much of what could have been captured is present) from coverage (how much of the real world is represented at all).1 Both matter. A field the model keys on may need to be near-complete. A field used as a weak signal may be usable at sixty percent if the missingness is not related to the outcome. State the threshold and state whether missingness is believed to be random, because that second part is where the real risk is.

3

Over what history, and how many events

Depth of history is not the same as number of events, and for most regulated use cases the event count is what binds. Sample size guidance for clinical prediction models makes the general principle explicit: adequacy depends on the number of outcome events relative to the number of candidate predictors, not on a fixed record count.8 The same logic applies outside clinical prediction. Five years of equipment history holding nine relevant failures is five years of history and nine events, and nine is the number that decides whether the use case is feasible.

4

With what identifier resolution

State which entities have to be resolvable across which systems, and to what accuracy. Product, site, batch, equipment, supplier, investigator, patient. This is the component teams most often leave implicit, and it is the one that most often turns out to be the blocker. Published work on record linkage found that when linkage error varied between hospitals rather than occurring randomly, the resulting bias in estimates reached far beyond what the raw match rate would suggest.9 Non-random matching failure does not just add noise. It moves the answer in one direction, and it moves it for the group whose identifiers were worst.

5

To what accuracy the decision actually needs

This is the component that gets skipped, and skipping it is why data requirements inflate. The question is not how accurate the data can be made. It is how accurate it has to be for the decision the model supports, given what happens when the model is wrong. A model that ranks a review queue and is checked by a human before anything is actioned tolerates error that a model feeding a release decision does not. Write down the decision, the consequence of a wrong answer in each direction, and the review step that catches it. The accuracy requirement follows from those three facts and cannot be set without them.

The template

Keep it to one page. If it takes more than a page, the use case is not defined tightly enough to plan.

LineWhat to write
Use caseOne sentence, naming the decision and who makes it today.
Decision supportedThe specific choice the output informs, and whether the model recommends, ranks, drafts, or decides.
Consequence of errorWhat happens on a false positive. What happens on a false negative. Which is worse and by how much.
Human review stepWho checks the output before anything happens, and what they can see when they check it.
Required fieldsField name, source system of record, and the tie-breaking rule when sources disagree.
Completeness thresholdPer field. Plus a statement of whether missingness is believed random, and how that was checked.
History and event countTime window required, and the minimum number of outcome events in that window.
Identifier resolutionWhich entities, across which systems, at what match rate, with what evidence that failures are not systematic.
Accuracy requirementThe threshold the decision needs, derived from the consequence and the review step, not from what is achievable.
Refresh and latencyHow current the data has to be for the decision to be useful. Daily, weekly, monthly, at batch close.
Current stateMeasured, not estimated. The date it was measured and by whom.
Gap classificationEach gap marked blocking or parallel, with an owner and a date.
VerdictStartable now, startable after named blocking work, or defer with the condition that would change the answer.

One rule about the current state line. It has to be measured against real records, not asserted from experience. A senior person’s recollection of how bad a dataset is will be wrong in both directions and there is no way to predict which. In practice teams find that the field everyone complains about is fine, and the field nobody mentioned is populated by free text in four languages.

Expect some dimensions to resist direct measurement. Published work applying the European framework to registry-based safety studies found that several dimensions could only be assessed through proxies rather than measured outright, which is a real limitation and not a reason to skip the exercise.3 Record what you measured directly, record what you inferred, and say which is which.

None of this is unfamiliar territory for a regulated organization. Characterizing a data source before relying on it, in writing, against stated dimensions, is already established practice on the regulatory side. The European good practice guidance for describing real-world data sources asks for exactly this kind of structured description so that a reader can judge whether a source suits a given study.5 A minimum data condition applies the same habit to an internal use case.

Blocking Work Versus Parallel Work

With a minimum data condition written, every gap in it falls into one of two categories, and getting this classification right is most of the value of the exercise.

Blocking work is data work that must be finished before the use case starts, because without it the build produces something that cannot be evaluated. The test is simple: if this gap remains, will we be unable to tell whether the model is working? If the answer is yes, it is blocking. Identifier resolution is usually blocking. So is the presence of the outcome field the model is trained to predict, and so is any transformation that changes what a field means.

Parallel work is data work that improves the use case but does not prevent an honest evaluation of it. Broadening history beyond the minimum. Improving completeness on secondary fields. Automating a control that is currently run manually. Extending a fix from two plants to all six. Parallel work can start at the same time as the build and land in later versions.

Data workUsual classificationWhy
Resolving entity identifiers across the systems the use case joinsBlockingNon-random match failure biases the result in one direction and the bias is invisible in aggregate metrics.
Defining and populating the outcome field the model predictsBlockingWithout a trustworthy label there is nothing to evaluate against.
Documenting field-level lineage from source to the model inputBlockingTraceability sits inside the reliability dimension and you will be asked for it.1
Freezing the reference lists the model depends on and version-controlling themBlockingAn unversioned reference list is a standing source of delayed, hard-to-trace downstream effects.7
Backfilling history beyond the minimum event countParallelImproves precision in later versions. Does not prevent an honest first evaluation.
Raising completeness on fields used as weak signalsParallelMarginal effect on output, measurable later, safe to sequence behind the build.
Migrating a source system to a modern platformParallelAlmost never genuinely blocking. Extract what you need from the current system and migrate on its own schedule.
Adopting an enterprise-wide standard such as ISO IDMP product identificationParallel, with a blocking subsetThe full adoption program runs for years.13 The subset of products in scope for this use case can be resolved in weeks.
Standing up a data governance council and a stewardship modelParallelNeeded for durability. Not needed to evaluate one model. Treat it as a program the first use cases fund.

The two ways teams get this wrong

The data team’s characteristic error is classifying too much as blocking. Every gap is real, so every gap feels like a prerequisite, and the blocking list grows into the same multi-year program the business already refused to fund. The discipline that fixes it is the evaluation test above. If the model can be honestly assessed with the gap present, the gap is not blocking, however much you want it closed.

The business team’s characteristic error is classifying too little as blocking, particularly around identifier resolution, because matching failures are invisible in a demonstration. A pilot on a hand-picked subset where somebody quietly reconciled the identifiers looks convincing and proves nothing about what happens at full scale. When a pilot works and the production version does not, unexamined identifier resolution is one of the first things to check.

A test worth applying to every “blocking” claim. Ask the person making it to describe the evaluation that becomes impossible if the gap stays open. If they can describe it, the gap is blocking and now everyone understands why. If the answer is that the results would be worse, that is a parallel improvement and it belongs in the backlog with a date, not in front of the start line.

Shared Milestones That Make Both Teams Succeed Together

Classification alone does not end the deadlock, because the two teams are still measured separately. The data team is judged on remediation completed. The business team is judged on the model going live. Under separate measures, each side has a rational reason to protect its own position, and the argument reappears at the next planning cycle wearing different clothes.

The structural fix is milestones that neither team can reach alone and neither team can claim alone. Three of them, defined at the start of the use case and reviewed by the same people.

Milestone 1: The data condition verdict. A written, measured answer to whether the minimum data condition is met, partially met, or not met, with numbers per line of the template and a date. The data team produces the measurement. The business owner signs that the condition as written matches the decision they actually make. Neither signature alone completes the milestone. Typical elapsed time is four to six weeks, and it is the cheapest milestone in the plan.

Milestone 2: The first decision-supporting output. Not a demonstration and not a model metric. An output produced on current production data, delivered to the person who makes the decision, in the form they would use, with the review step in place. The business team owns the delivery. The data team owns the feed behind it and confirms it will still be correct next month.

Milestone 3: The reuse commitment. A named element built for this use case that a later use case will depend on, with the later use case named and the date it takes the dependency. This is what converts one-off work into a foundation. Without it, every use case rebuilds its own extract, and after three use cases the company has three pipelines and no platform.

What makes these work is that they are jointly claimable. A data leader who blocks the start of a use case now blocks their own milestone. A business leader who wants to skip the identifier work is skipping the verdict their own delivery depends on. That is a better mechanism than a governance forum, because it changes what each party is trying to achieve rather than adding a body to adjudicate between them.

Two practical notes. First, keep the verdict milestone short and hard-dated. If the data condition assessment is allowed to run for a quarter, it becomes the assessment phase everyone was trying to avoid. Four to six weeks with a written verdict at the end, even a partial one, beats a thorough answer three months later. Second, publish the verdict whatever it says. A verdict of “not met, defer” that is circulated as widely as a favorable one is what makes the next verdict believable.

A Worked Example: Three Use Cases, One Sequenced Plan

The following is a composite drawn from the way this analysis usually goes at a mid-size company with commercial manufacturing and an active clinical portfolio. The names and numbers are illustrative. The shape of the finding is not.

Three use cases arrive in the same planning cycle, each with a sponsor and a value estimate.

  • Use case A: deviation triage support. Classify incoming manufacturing deviations by likely category and probable investigation depth, and draft the initial summary for the investigator. Sponsor is the quality operations lead. Value is in cycle time and in consistency of classification across four plants.
  • Use case B: lyophilizer failure forecasting. Predict component failures on freeze dryers far enough ahead to schedule maintenance into planned downtime rather than losing a batch. Sponsor is the engineering director. Value is in avoided batch loss and unplanned downtime.
  • Use case C: site enrollment forecasting. Forecast enrollment by site by month for active studies to support reallocation decisions. Sponsor is clinical operations. Value is in schedule protection.

Running the minimum data condition on each

Use case A. The decision is which queue a deviation enters and how much investigation effort it gets, made today by a quality reviewer within a working day of the deviation being raised. A wrong classification in one direction wastes investigator time. In the other direction it under-investigates something that should have been escalated, which is the serious error. There is a human review step: the reviewer sees the recommendation and the deviation text together and confirms or changes it before anything is routed.

Required fields are the deviation narrative, the classification recorded at closure, product and equipment references, and the investigation outcome category. All of it comes from one quality management system, which the company has run for six years. History depth is six years. Event count at the closed-deviation level is in the thousands, which is comfortably above what the number of candidate predictors requires.8 Identifier resolution is needed only within one system. The measured finding is that classification at closure is populated on ninety-six percent of records, the narrative field is present on all of them, and the equipment reference is free text on records before the 2023 upgrade.

Verdict: startable now. One blocking item, which is that the pre-2023 free-text equipment references have to be either mapped or excluded, and the exclusion has to be shown not to correlate with the outcome. Two weeks of work. Everything else is parallel.

Use case B. The decision is whether to pull a unit for maintenance during a scheduled window. A wrong answer in one direction spends maintenance hours unnecessarily. In the other direction it means an unplanned failure during a run, which is the outcome the use case exists to prevent. Required data is sensor history from the process historian, maintenance work orders with cause codes from the maintenance system, and a resolvable link between the two describing the same physical asset.

The measured finding is the one that decides everything. Across all lyophilizers and five years of maintenance records, the number of unplanned failures with a usable cause code is eleven. Eleven events. The historian holds high-frequency sensor data, so there is no shortage of rows, but rows are not events, and the thing being predicted has occurred eleven times. Separately, historian tags and maintenance asset identifiers were established independently and do not resolve without manual mapping, which nobody has done.

Verdict: defer, with a stated condition. The condition is not a date. It is a fact that would have to become true: enough labeled failure events to support a model, or a redefinition of the use case to predict a more frequent intermediate event, such as an out-of-range excursion, rather than a failure. That redefinition is worth exploring and it is a different use case with its own minimum data condition.

Use case C. The decision is whether to add sites, reallocate recruitment spend, or adjust a timeline, made monthly by a study team. Errors in either direction are absorbed by a human decision process with other inputs, so the accuracy requirement is moderate, and the value comes from being roughly right early rather than precisely right late.

Required data is enrollment events by site by month, site attributes, and protocol characteristics, drawn from the trial management system plus two vendor feeds. History covers the study portfolio back to 2019. The measured finding: enrollment records are complete, but site identity does not resolve cleanly across the three sources. Manual review of a sample of one hundred sites showed that automated matching on name and address agreed with the manual answer on seventy-eight of them, and that the failures were concentrated in multi-site networks and in non-US sites where naming conventions differ. That concentration is the important part. It is exactly the non-random pattern the linkage literature warns about, where the error is associated with a characteristic that also relates to the outcome, and the resulting bias is directional rather than merely noisy.9

Verdict: startable after named blocking work. The blocking item is site identifier resolution across the three sources for the studies in scope, to a stated match rate with evidence that residual failures are not concentrated in one region or network type. Estimated at eight weeks. And this is the item worth noticing, because a resolved site master serves the enrollment forecast, the vendor performance analysis someone else has been asking for, and any future study feasibility work. It is foundational work with a specific first beneficiary.

Use caseBinding constraintBlocking workParallel workVerdict
A. Deviation triage None material. Single system, deep history, adequate events. Map or exclude pre-2023 free-text equipment references (2 weeks). Extend to the two smaller plants; improve outcome category consistency. Start now. First delivery targeted at week 14.
B. Failure forecasting Eleven labeled failure events in five years. Not scopeable. The events do not exist to be remediated. Historian and maintenance asset mapping, useful regardless. Defer. Revisit if redefined around a more frequent event.
C. Enrollment forecasting Site identity resolves on 78 of 100 sampled, with failures concentrated by region and network. Site identifier resolution across three sources with a bias check (8 weeks). Backfill pre-2019 studies; enrich site attributes. Start after blocking work. First delivery targeted at week 26.

The sequenced plan that comes out of it

The plan writes itself once the three verdicts are on one page, and it looks nothing like either opening position.

1

Weeks 1 to 6: verdicts on all three

Measured minimum data conditions, written, signed by both the data owner and the business sponsor. This is the whole assessment. It replaces the enterprise data quality assessment that would otherwise have been proposed, and it is bounded because each question is bounded.

2

Weeks 3 to 14: build use case A while site resolution starts

Use case A’s two-week blocking item runs immediately and the build follows. In the same window, the site identifier work for use case C begins, funded as a dependency of a use case with a sponsor and a date rather than as a master data initiative.

3

Week 14: first decision-supporting output

Deviation triage recommendations reach quality reviewers in the tool they already use, with the review step in place. This is the first evidence the program produces, and it arrives roughly three months in rather than in year two.

4

Weeks 14 to 26: use case C build on resolved identifiers

The enrollment forecast is built on a site master that now exists, with a documented match rate and a documented check that residual failures are not concentrated. That documentation is reusable evidence, not overhead.

5

Week 26 onward: reuse and reassessment

The site master supports the next two candidates without a new build. Use case B is reassessed against its stated condition, either because the redefinition around excursions proved workable or because the event count has grown. Nothing about B was abandoned. It was scheduled against a fact.

Notice what this plan does to the original argument. The data team got its blocking work funded and done before the models were built. The business got a working output in fourteen weeks. Finance funded named work with named beneficiaries. Nobody had to win the argument about which comes first, because after the verdicts there was no argument left to win.

Sequencing the Portfolio So Early Work Funds Later Work

The single-use-case analysis is the building block. Portfolio sequencing is where the compounding happens, and it turns on one question asked of every candidate: what does this use case build that a later one can reuse?

Rank by value against distance to condition, not by value alone

A ranked list by value alone puts the most attractive use case first, which is frequently the one with the deepest data gap, which is how programs stall at the starting line. Ranking by value against distance to the minimum data condition produces a different first item, usually a less exciting one, and it is startable. Getting something working changes the political economy of the program more than a better idea does.

This is where the staged framing of data readiness is useful. The data readiness levels proposal describes readiness as a series of bands from data that is merely accessible, through data that is verified and characterized, to data that is fit for a specific stated purpose.10 Two use cases can both be described as “blocked on data” while sitting in entirely different bands. One needs a field populated. The other needs a source system that does not yet capture the concept. Those are not the same distance and they should not be sequenced as though they were.

Map the shared dependencies before ordering anything

Lay every candidate use case against the data elements it needs and look for the elements that appear three or more times. Those are your foundational items, identified by demand rather than by architecture preference. In most pharma and biotech portfolios the same few appear: product and material identity, site and organization identity, equipment and asset identity, batch genealogy, and the reference lists that classify everything else.

Sequence so that the first use case to need a shared element pays for building it properly rather than building a private version. This is the practical difference between a program that accumulates a platform and one that accumulates pipelines. It also gives you a defensible answer when someone asks why the first use case took fourteen weeks instead of nine: five of those weeks built something three later use cases will not have to build.

Let the standards programs run on their own clock

Large standardization efforts do not fit inside a use case and should not be forced to. ISO IDMP product identification in Europe runs on a multi-year implementation path with its own milestones and its own regulatory drivers.13 The FAIR data principles set expectations for findable, accessible, interoperable, and reusable research data that no single project completes.14 Industry roadmaps for digital manufacturing describe maturity progressions measured in years.1112

These belong in the plan as long-running programs, and the connection to the use case portfolio runs in one direction only: a use case may need a specific subset of a standard delivered early, and that subset is a legitimate blocking item. The reverse claim, that a use case must wait for a standards program to finish, is almost never true and should be tested every time it is made.

Four sequencing rules

  • The first use case is chosen for startability, not for value. It buys the credibility that funds the rest.
  • Shared data elements are built properly by the first use case that needs them, and that use case’s schedule carries the difference openly.
  • No use case starts before its blocking work has an owner and a date. Starting anyway is how the data condition becomes a production problem.
  • Deferred use cases keep a written condition for reconsideration. A deferral with no condition attached is a cancellation that nobody has admitted to.

When the Honest Answer Is That the Data Is Not There

Some use cases fail the analysis. Saying so is the part that makes the whole method credible, and it is the part most programs avoid because deferring anything feels like admitting the program is not working.

Four patterns account for most genuine deferrals.

The outcome was never recorded

The most common one. The company wants to predict something it has never systematically written down. Which investigations later required a corrective action. Which supplier lots caused downstream trouble. Which deviations were reopened. The operational knowledge exists in people’s heads and in narrative text, and there is no field holding the answer. This is not a data quality problem and no remediation project fixes it retroactively. The fix is to start recording the outcome now, which is a worthwhile decision with a two-year payoff, and to defer the use case until the record exists.

The events are too rare

Use case B above. The thing being predicted has happened a handful of times. No amount of sensor data changes that, because the constraint is on the label side. The sample size literature is unambiguous that adequacy is governed by the number of outcome events relative to the number of candidate predictors, not by the number of rows.8 The productive response is usually to look for a more frequent intermediate event that is genuinely related to the rare one, and to treat that as a new use case rather than a reframing of the old one.

The identifier problem is structural

Sometimes entities cannot be resolved because the systems never captured anything that identifies them consistently, and the manual reconciliation required is larger than the use case is worth. Where residual matching failures concentrate in a particular group, the effect on the answer is directional and cannot be corrected by scaling up.9 Defer, fix the capture at the source, and revisit.

The accuracy the decision needs is not reachable

Occasionally the honest finding is that the decision requires a level of confidence the available data cannot support, and no review step is adequate to catch the errors in time. This is the deferral most worth making explicit in a regulated setting, because the alternative is a system that produces plausible outputs into a process where being wrong has consequences that the review step does not catch.

How to defer without killing the idea. A deferral is only credible if it comes with the condition that would reverse it. Write the condition as a fact, not a date: “revisit when there are at least forty labeled failure events, or when the use case is redefined around excursions.” Put a review date on the calendar. Tell the sponsor what would have to change and who owns the change. A deferral written this way keeps the idea alive and keeps the analysis trusted. A deferral with no condition teaches everyone that the exercise is a way of saying no.

What to Tell an Executive Who Wants a Date

None of the above survives contact with an executive committee unless you have an answer to the question that will actually be asked, which is when will this be done.

The answer is not to explain that it depends. It is to give three dates instead of one, and to be precise about what each one commits to.

Date you commit toWhat it meansWhat it does not mean
The verdict date
Typically 4 to 6 weeks out
A measured, written answer on whether each candidate use case can start, what blocking work stands in the way, and what it takes. It does not promise that any use case will be startable. It promises that the ambiguity ends on that date.
The first output date
Set only for use cases that passed
A decision-supporting output in the hands of the person who makes the decision, on production data, with review in place. It does not promise a measured business result yet, and it does not promise enterprise deployment.
The portfolio review date
Typically two quarters out
A reassessment of every deferred use case against its written condition, and a re-sequencing based on what the first deliveries showed. It does not promise that the deferred items will have become feasible.

Four things to say, and one to avoid

Say what the condition is, not how confident you feel. “We can start deviation triage in two weeks because the classification field is populated on ninety-six percent of six years of records” is a sentence an executive can act on. “We think the quality data is in reasonable shape” is not, and it will be remembered as a commitment anyway.

Give a date for the next decision, not for the outcome. Programs lose credibility by promising results whose timing depends on facts not yet established. They keep credibility by promising decisions on dates that are entirely within their control. The verdict date is fully controllable. Use it.

Name the cancel condition up front. Executives are more willing to fund work when they know what would end it. State what result at the verdict milestone would stop each use case. This is also the single most effective protection against a program that continues past the point of usefulness because nobody defined what failure looked like.

Show what the eighteen months of nothing has already consumed. Deadlocks feel free because their expense is not booked anywhere. It is worth putting a number on the assessments, the workshops, the vendor evaluations, and the leadership hours already spent on the argument. That number is usually larger than the first blocking item, and it changes the framing from “should we spend money” to “should we keep spending it this way.”

What to avoid: the composite readiness score. A single number claiming the organization is, say, sixty-two percent AI-ready is worse than no number, because it invites a target and a program to move it, and moving it does not make any specific use case startable. Maturity models are useful for describing a trajectory over years1112, and they are the wrong instrument for deciding what to do next quarter. Per use case verdicts are the right instrument, and they are harder to argue with because they are specific.

The sentence that ends the deadlock. “We are not going to decide whether data comes before AI. In six weeks you will have a written verdict on each of these three use cases telling you which can start now, which can start after named work with an owner and a date, and which should wait and for what specific reason. Then you will decide what to fund.” Nobody in the room can reasonably object to that, which is the point.

Conclusion

The deadlock between data readiness and AI delivery persists because both sides are arguing about programs, and programs at that level of abstraction cannot be resolved against each other. Neither one has a boundary. Neither one can be finished. The argument is unwinnable by design, which is why eighteen months of it produces nothing.

Moving the conversation down to individual use cases changes what is being discussed. A use case has a decision behind it, and a decision sets a standard for the data, which is precisely how the regulators define data quality in the first place: fitness for a stated purpose rather than an absolute property of a dataset.14 Once the standard is stated for a particular decision, the gap to it is measurable, the work to close it is scopeable, and the split between what must happen first and what can happen alongside becomes a technical judgment rather than a negotiating position. Some use cases turn out to be startable next month. Some need eight weeks of identifier work that three later use cases will also depend on. Some should wait, and saying so with a written condition attached is what makes the rest of the analysis believable.

Sakara Digital works with pharma and biotech organizations sequencing data work against an AI roadmap without stalling either one. If you are holding a portfolio of AI candidates and an argument about what has to come first, and you would like an independent read on which of them is actually startable and what the minimum data condition for each one really is, we are happy to have that conversation.

For Further Reading