What the Completion Evidence Actually Says

Any article on this subject has an obligation to be careful with numbers, because this is a subject where bad numbers travel further than good ones. The figure most often quoted to executives is that more than 80 percent of AI projects fail. It appears in the opening of RAND’s 2024 report on the root causes of AI project failure, which is a serious piece of work by James Ryseff, Brandon De Bruhl, and Sydne Newberry.1 Follow the footnote, though, and the 80 percent figure is not RAND’s measurement. It traces to a 2022 Fortune article reporting one software company chief executive’s estimate of how often AI projects disappoint.2 That is a professional opinion, published in a reputable magazine, and it is not a study. RAND is transparent about this. Almost everyone who repeats the number is not.

The same discipline is worth applying to the figure that dominated boardroom conversation through late 2025, that 95 percent of generative AI pilots produce no measurable return. That claim comes from a report issued by a research group at MIT, based on a review of publicly disclosed initiatives, a set of structured interviews, and a modest number of survey responses.6 It was not peer reviewed, its sampling approach drew immediate criticism from industry analysts, and the number describes a specific and narrow definition of return.7 None of that makes it worthless. It makes it a data point with a method attached, which is a different thing from a fact.

Here is what survives that filter and can be cited with the sample and the method stated alongside it.

42% of surveyed organizations abandoned most AI initiatives before production, up from 17% the prior year (S&P Global Market Intelligence, more than 1,000 IT and line-of-business respondents in North America and Europe)3
46% the average share of AI proofs of concept scrapped between demonstration and production in the same survey3
65 experienced data scientists and engineers interviewed in RAND’s study of why AI projects fail, all with at least five years building models1

Four sources worth knowing, and what each one can and cannot support

RAND, 2024. Ryseff and colleagues interviewed 65 data scientists and engineers with at least five years of experience building machine learning models across a range of company sizes, industries, and academia. They identified five recurring root causes: the problem to be solved was misunderstood or miscommunicated; the organization lacked the data needed to train an effective model; the team chased the technology rather than the problem; the infrastructure to manage data and deploy models was not there; and the technology was applied to problems beyond what it can currently do. The recommendation that applies most directly to this article’s subject is the second one: choose enduring problems, and be prepared to commit a team to a specific problem for at least a year before starting. If a problem is not worth that commitment, it is probably not worth starting.1

S&P Global Market Intelligence, 2025. The Voice of the Enterprise survey covered more than a thousand mid-level and senior IT and line-of-business professionals across North America and Europe. The abandonment rate for the majority of an organization’s AI initiatives rose from 17 percent to 42 percent in a single year, and organizations reported scrapping an average of 46 percent of proofs of concept before production. Respondents named expense, data privacy, and security risk as the leading obstacles.3 Note what this measures: it is a self-reported count of abandonment, not a verdict on whether the abandonment was correct. Some of that 42 percent is organizations finally making decisions they should have made a year earlier.

Gartner, July 2024. Gartner predicted that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating expense, and unclear business value.4 This is a forecast rather than a measurement, and it should be quoted that way.

BCG, October 2024. A survey of 1,000 senior executives across 59 countries placed 74 percent of companies as not yet showing tangible value from AI, with 49 percent still focused on proofs of concept and 4 percent operating at scale.5 The category boundaries are the survey author’s own, which is worth remembering, but the shape of the distribution is consistent with everything else.

Alongside the industry surveys, there is a small academic literature built the same way RAND built its study. Schlegel, Schuler, and Westenberger interviewed AI practitioners and grouped twelve failure factors into five categories: unrealistic expectations, use case issues, organizational constraints, missing key resources, and technological issues.8 Read next to RAND, the overlap is striking. Neither study finds that models underperform as the leading problem. Both find that the organization never established what the project was supposed to establish.

The practical takeaway from the evidence. Across every credible source, the dominant failure is not technical. It is that nobody defined the question the pilot was supposed to answer, so nobody could tell when it had been answered. Every recommendation in the rest of this article follows from that single finding.

What a Pilot Is Actually For

A pilot is an instrument for reducing one specific, named uncertainty at a scale small enough that being wrong is affordable. That is the whole definition. It is not a demonstration of enthusiasm, it is not a way to keep a vendor relationship warm, and it is not a signal to the board that the organization is engaged with AI. Those things may happen as side effects. They are not the purpose, and when they become the purpose, the pilot loses the property that made it useful: the ability to end.

The discipline that makes this concrete is writing the uncertainty down as a question with a possible negative answer. “We want to explore how generative AI could support deviation triage” is not a question. “Can a model classify incoming deviations into our existing four severity bands with agreement of 90 percent or better against two qualified reviewers, on 300 historical records drawn from the last 18 months across all three sites?” is a question. It can be answered no. A pilot built around the second sentence can finish. A pilot built around the first cannot, because there is no state of the world in which exploration is complete.

The uncertainties worth spending a pilot on

Most AI proposals in a regulated environment carry several uncertainties at once, and the useful first step is separating them, because they need different evidence and they do not all require a pilot. Some can be settled by a two-day desk review, which is a much cheaper way to get an answer.

UncertaintyWhat resolves itDoes it need a pilot?
Is the technical performance achievable on our data?A measured evaluation against a held-out set and a defined baselineYes. This is the classic case.
Do we have enough data of adequate quality, with usable provenance?A data readiness assessment against the specific fields the model needsUsually no. Assess first, and if the answer is no, the pilot has already failed at a fraction of the effort.
Will the output fit the actual workflow the reviewer performs?Observation of the current task plus a structured walkthrough with reviewersPartly. Workflow fit is best tested in a pilot, but only after a desk review confirms the workflow is understood.
Is there a validation and compliance path we can defend?A written path reviewed by quality and regulatory before the pilot startsNo. This is prerequisite work. A pilot cannot create a path that does not exist.
What will it take to run this every day for three years?A run estimate covering compute, monitoring, retraining, and named ownersPartly. The pilot should produce the inputs to this estimate, not discover the question at the end.
Will the people who have to use it accept it?Adoption measured during the pilot with the real users, not a demonstration audienceYes, and this is the one most often skipped.

Separating the uncertainties this way has an immediate benefit that has nothing to do with governance. It shrinks the pilot. If the data readiness assessment and the validation path review happen first, the pilot itself is testing one or two things rather than six, which means it can be smaller, shorter, and easier to read. It also means the pilot that does go ahead is not carrying a hidden dependency that will surface in month four and become the reason it cannot end.

This is also where the relationship between a data program and an AI roadmap matters. Sequencing the underlying data work against the model work is a separate discipline with its own logic, and getting the order wrong is one of the more common ways an otherwise sound pilot becomes unfinishable. The point for this article is narrower: if the data work is the real uncertainty, name it as the pilot, and stop calling the exercise a model pilot.

Writing Stop Criteria and Scale Criteria on Day One

The single most useful thing a sponsor can do is write both criteria before any work starts, when nobody is invested and the evidence does not yet exist. Written at the start, criteria are an analytical exercise. Written at month five, when the team is tired and the sponsor has already described the pilot as promising in two steering meetings, they are a negotiation. The whole value of pre-writing is that it moves the decision out of a moment when the decision is hard.

Two criteria are needed, not one. A scale criterion alone leaves the pilot in an undefined state when it is not met, and undefined states persist. A stop criterion alone gives the team nothing to aim at. Both must be written so that a reasonable person reading the pilot report six months later, who was not in the room, could apply them and reach the same conclusion the sponsor did.

SCALE CRITERIA

What must be true to invest further

Performance thresholds against a named baseline, an accepted validation path, a run estimate within a stated ceiling, a named operating owner who has agreed in writing, and evidence of adoption by the users who will actually do the work.

STOP CRITERIA

What makes further investment unjustified

Performance below a floor that no realistic iteration closes, a data gap that cannot be filled within the roadmap, a validation path quality cannot accept, a run estimate above the ceiling, or a use case whose value has been overtaken by a change elsewhere.

TIME BOX

The date the decision is made

A fixed calendar date, set at the start, that does not move when results are ambiguous. Ambiguity is a result. The date is when the decision is made, not when a good result is expected to arrive.

EFFORT CEILING

The budget and staffing that will not be exceeded

A stated limit on spend and on the number of people committed. Passing the ceiling before the date is itself a stop trigger, and it should be treated as one rather than as a routine request for more.

What a real criterion looks like

The test of a criterion is whether it can be evaluated by someone who does not want a particular answer. Vague criteria are not neutral. They reliably resolve in favor of continuation, because the person applying them is the person who wants to continue.

Not a criterionA criterion
The model shows promising accuracy.Agreement with two qualified reviewers is at or above 90 percent on a held-out set of 300 records, with no severity band below 80 percent.
Users find it helpful.At least 60 percent of the eight pilot reviewers use the output in their routine work in weeks six through ten without being prompted, measured from system logs.
Quality is comfortable with the approach.Quality assurance has signed a written validation approach covering intended use, risk classification, acceptance criteria, and periodic review frequency.
The running expense seems manageable.Estimated annual run expense including compute, monitoring, retraining, and 0.5 FTE of support is at or below the stated ceiling, with the estimate reviewed by finance.
We will decide when we have more data.The decision occurs on 14 November, using whatever evidence exists on that date.

The criterion that does the most work is the baseline. A model that agrees with reviewers 88 percent of the time sounds like a near miss until someone establishes that two qualified human reviewers agree with each other 84 percent of the time, or that the existing rule-based screen already reaches 86 percent. Without a measured baseline, every performance number is unreadable, and an unreadable number is exactly the kind of evidence that keeps a pilot alive. Measure the current process before the pilot starts, not after.

Who writes them

The criteria should be drafted by the sponsor and reviewed by at least one person with no stake in the outcome. In practice, that reviewer is often quality assurance or an internal audit function, both of which are practiced at asking whether a written acceptance criterion can actually be evaluated. In pharma and biotech there is an additional benefit to routing them through quality: the habit of writing acceptance criteria before running a study is already the professional norm. Nobody would run a process validation without pre-defined acceptance criteria. The argument for writing them for an AI pilot is the same argument, applied to a newer object.

The Decision Date and Who Owns It

A criterion without a date is a wish. The date should be set at the same time as the criteria, recorded in whatever governance minute the organization actually keeps, and communicated to the pilot team on day one. The team should know from the beginning that a decision meeting exists, that it is on the calendar, and that it will happen whether or not the results are tidy.

Three details make the date hold.

1

Name a single accountable decision owner

One executive makes the call. Committees do not stop things, because stopping requires someone to accept the discomfort of being the person who stopped it, and a committee distributes that discomfort until it disappears. Name the individual in the charter.

2

Separate the decision owner from the pilot champion where you can

The person whose reputation is attached to the pilot’s success is the worst-placed person to end it. Where the sponsor and champion are necessarily the same person, compensate by requiring the decision to be presented to a governance body that did not build the thing, with the pre-written criteria displayed alongside the results.

3

Fix the agenda of the decision meeting in advance

Results against each pre-written criterion, in order, with a stated pass or fail for each. Then the decision. Not a demonstration, not a roadmap for the next phase, and not a discussion of what could be achieved with more time. Those conversations belong after the decision, if the decision was to continue.

4

Allow exactly three outcomes

Scale, stop, or one bounded extension of fixed length with a new date and a criterion that was not previously testable. The extension is permitted once. Naming that limit in the charter is what prevents the extension from becoming the operating mode.

The third outcome deserves care, because it is the one that gets abused. A legitimate extension has a specific shape: a criterion could not be evaluated for a reason outside the team’s control, the reason has since been removed, and evaluating it will take a known and short amount of time. A site was unavailable for user testing and is now available. A data extract was delayed by a system migration that has now completed. Those are extensions. “The results are close and the team believes another round of tuning will close the gap” is not an extension. It is the beginning of the pattern this article exists to describe.

The Evidence a Scale Decision Genuinely Needs

Scaling is a much larger commitment than piloting, and the evidence needed to justify it is correspondingly broader than a performance number. A pilot that produced only a performance number has not produced enough to scale on, however good the number is. Four categories of evidence are needed, and in a regulated environment all four are load-bearing.

1. Measured benefit against the baseline

Not model performance. Benefit. The distinction matters because a model can be accurate and change nothing. If a classification model agrees with reviewers 92 percent of the time but a qualified person still reviews every record, because the risk classification requires it, then the benefit is whatever time the reviewer saves by starting from a suggestion rather than a blank form. That may be substantial or it may be minutes. It has to be measured, against the same task performed the old way, by the same people, on comparable work.

The baseline needs to be established before the pilot, because retrospective baselines are unreliable and everybody knows it. This is also the point at which claims about the technology’s general capability should be set aside in favor of what happened here, on this data, in this workflow. The broader record of what AI has and has not delivered in drug development is a useful corrective to optimistic framing, and it is a subject worth reading on its own, but it does not substitute for a measurement of your own process.

2. The validation and compliance path

In pharma and biotech, a model that cannot be validated cannot be scaled, regardless of performance. The path has to be written and accepted before scaling, not discovered afterward. When FDA’s Center for Drug Evaluation and Research collected public feedback on its 2023 discussion paper on artificial intelligence in drug manufacturing9 and at the associated public workshop, agency staff summarized what interested parties said in a 2025 commentary in AAPS Open. The recurring themes were revealing: respondents valued good data management practices, wanted clearer best practices for AI model development, reported real uncertainty about how to manage models supplied by third parties, and found implementing AI within the pharmaceutical quality system framework genuinely difficult.10 Lifecycle management of models was one of the explicit discussion topics.

The practical reading for a scale decision is that the hard questions are not about the model. They are about what happens to the model over the following three years: who reviews its performance, what triggers retraining, how a retrained model gets through change control, what the pharmaceutical quality system says about it, and what happens when the vendor updates something you did not ask them to update. If the pilot did not produce answers to those questions, it did not produce a scale case. European regulators have moved along similar lines, with the European Medicines Agency’s reflection paper on artificial intelligence in the medicinal product lifecycle setting out a risk-based expectation for development, deployment, and performance monitoring across the product lifecycle.

3. The run expense

The expense of running a model in production is routinely underestimated because the pilot does not incur most of it. Pilots run on borrowed infrastructure, with the people who built the model also operating it, with no monitoring beyond someone looking at results, and with no retraining cycle because the pilot is too short to need one. The interview study by Shankar and colleagues with machine learning engineers running production systems describes the actual shape of the work: a continual loop of data collection and labeling, experimentation, staged evaluation, and monitoring for performance degradation.11 None of that loop exists during a twelve-week pilot, and all of it has to be funded for a production system.

A defensible run estimate covers compute and licensing, monitoring and alerting, the periodic performance review the validation approach commits to, retraining and the change control that accompanies it, the support capacity to answer user questions, and the documentation maintenance that keeps the validated state current. Expressing it as an annual figure with a named budget owner is what turns it from an assumption into a decision input.

4. The ownership question

This is the one most often left unanswered, and it is the one that most often turns a scaled pilot into an orphan. The pilot team is usually a mix of a data science group, a vendor, and a few enthusiastic subject matter experts borrowed from their day jobs. None of those three is the long-term operator. So the question has to be asked directly and answered by name: which line function owns this system after the pilot team disperses, who in that function is accountable for its performance, what capacity has been allocated to them, and have they agreed?

The four questions that decide a scale case

  • Benefit: How much better is this than the measured baseline, on the same task, with the same people?
  • Path: Has quality assurance accepted a written validation and lifecycle approach, including retraining and change control?
  • Run: What does a year of operation require in money and in staff time, and who owns that budget line?
  • Owner: Which named function operates it after the pilot team leaves, and have they agreed in writing?

A pilot that answers all four supports a scale decision. A pilot that answers only the first supports another conversation, not a scale decision.

Why Pilots Stay Alive Past Their Usefulness

Knowing the criteria and missing them does not, by itself, end a pilot. Four forces reliably keep work going after the evidence has stopped supporting it, and all four are ordinary human behavior rather than misconduct. Naming them in advance is most of the defense.

Sunk effort

The tendency to continue an endeavor because of what has already been invested in it is one of the most consistently replicated findings in the behavioral literature, and it is well documented in project settings specifically. Reviews of behavioral biases in project management place escalation of commitment among the recurring patterns that distort decisions about whether to continue.12 The mechanism is not stupidity. It is that stopping requires acknowledging that the prior investment produced no continuing asset, and people are strongly averse to that acknowledgment, particularly in public.

The defense is structural rather than motivational. Pre-written criteria, evaluated by someone who did not make the original investment decision, remove the moment where sunk effort gets to argue. It is also worth stating explicitly in the charter that money already spent is not evidence about the future, because saying it out loud at the start makes it easier to say at the decision meeting.

The sponsor’s reputation is attached to it

When an executive has described a pilot as strategically important in three consecutive leadership meetings, stopping it reads as a personal reversal. This is the reason the decision owner and the champion should be different people wherever the organization can arrange it, and the reason the decision should be presented against pre-written criteria rather than argued on merits. The criteria give the sponsor something to point at that is not their own judgment on the day.

There is a cultural remedy that requires no budget and works. Leadership should say, out loud and more than once, that stopping a pilot on its criteria is a successful outcome and will be described that way. The first time a senior person stops their own pilot and is publicly credited for the decision, the organization learns something that no policy document teaches.

Redefining success mid-pilot

This one is subtle because it usually happens with good intentions. The model does not reach the classification threshold, but the team notices it is very good at flagging a narrower category of records, and the pilot is redescribed around that. Sometimes this is a genuine and valuable discovery. Often it is the criteria being rewritten to match the result, which destroys the pilot’s ability to answer anything.

The way to keep the discovery without losing the discipline is to treat the new finding as a new pilot proposal. Record the original criterion as not met, stop the original pilot, and let the narrower use case compete for funding against everything else in the portfolio on its own merits. If it is genuinely valuable it will win that competition. If it only looks valuable because it is the thing that survived, the competition will show that too.

One more iteration

The final pattern is the most persistent, because each individual request is reasonable. Another two weeks of tuning. A slightly larger training set. A different prompt structure. A newer model that was released last month. Each request has a plausible technical rationale, each is cheap on its own, and together they extend indefinitely, because there is always a next thing to try. This pattern is what the effort ceiling exists to catch. When the ceiling is reached, the pilot ends, and the question of whether one more iteration would have worked is left permanently open. Leaving it open is the point. An organization that will not tolerate an open question of that kind will not be able to stop anything.

A useful diagnostic. Ask the pilot team: what result, if you saw it next week, would make you recommend stopping? If the team cannot answer, or the answer is a result so catastrophic that it was never plausible, the pilot has no stop condition in practice regardless of what the charter says. Ask this question at the midpoint, not at the end.

Closing a Pilot Without Discrediting the Team or the Technology

Everything above is about reaching the decision. This section is about the fortnight after it, which is where the damage usually happens. A pilot that is stopped carelessly does lasting harm in three directions at once. The team that ran it is understood by everyone else to have failed. The technology is understood to have been tried and found wanting, which makes the next proposal harder. And the organization learns that AI work carries reputational risk, which pushes the next round of proposals toward the safest and least valuable use cases.

None of that is inevitable. It follows from how the stop is executed and described, and the execution is largely a documentation and communication exercise.

The closure record

Write a closure record within two weeks of the decision, while the people involved are still available and still remember the details. It should be short, five to eight pages, and it should be filed where the next person evaluating a similar use case will find it. In practice that means the same repository as the rest of the governance documentation for the portfolio, not a folder belonging to the team that ran the pilot.

SectionWhat it contains
The questionThe uncertainty the pilot was designed to reduce, stated as it was written on day one.
The criteriaScale and stop criteria exactly as written at the start, unedited.
The result against eachA pass or fail for every criterion, with the measurement and the method used to produce it.
The decision and the deciderWhat was decided, by whom, on what date, and in which forum.
What was learnedFindings about the data, the workflow, the validation path, and the operating model. Separate what is specific to this use case from what applies to the next one.
Assets retainedData pipelines, labeled sets, evaluation harnesses, and governance documents that survive the pilot, with their new owner named.
Revisit conditionThe named change in circumstances that would justify looking again, the owner who watches for it, and the review date.
DispositionWhat happens to the environment, the data, the model artifacts, and the vendor arrangement, with dates.

The disposition section deserves attention in a GxP context. Even a pilot that never reached validated status generates records, and those records have retention implications that should be settled deliberately rather than by whoever eventually cleans up the environment. Decide what is retained, for how long, and under whose ownership, and write it down. The same discipline that applies to decommissioning a system applies in miniature to closing a pilot environment.

How to describe a stop

The language used in the announcement determines how the organization reads the event, and it is worth drafting deliberately rather than improvising in a status meeting. The frame that works is that a decision was completed, not that an effort failed.

A description that closes cleanly: “The deviation classification pilot completed on 14 November. It was designed to answer whether a model could reach 90 percent agreement with qualified reviewers on our historical records. It reached 81 percent, with weaker performance on the two lower-volume severity bands, and the analysis points to inconsistent historical coding rather than model capability. On that evidence we decided not to scale. The evaluation method and the cleaned record set are being retained and transferred to the data governance team, and the coding inconsistency is being handled as a separate piece of work. We will look at this use case again once the historical coding remediation is complete, which is scheduled for the second quarter.”

Four things are happening in that paragraph. The pilot is described as complete rather than cancelled. The criterion is restated so the audience can see the decision was rule-based rather than a matter of opinion. The reason is attributed to a specific, factual, fixable condition rather than to the team or the technology. And a future is left open with a named condition attached to it. None of that is spin, because every sentence is true. It is simply the accurate version, told in the order that makes it legible.

What not to say

Avoid describing the pilot as a failure, avoid attributing the outcome to the team’s execution unless that is genuinely and specifically what happened, and avoid the broader generalization that “AI did not work for this.” The pilot tested one model on one dataset for one task under one set of constraints. Extending that to a claim about the technology is not supported by the evidence the pilot produced, and it makes the organization worse at evaluating the next proposal.

Equally, avoid overcorrecting into language so soft that nobody can tell a decision was made. “We are pausing to reassess” leaves the pilot in exactly the undefined state this whole article is about. Say it ended. Say why. Say what happens next.

What to do for the team

The people who ran the pilot have spent months on something that will not continue, and how they are treated in the following weeks is watched closely by everyone who might run the next one. Three things help. Credit them publicly for the quality of the answer rather than the direction of it, because a well-run pilot that produces a clean negative is genuinely valuable work. Move them onto the next thing quickly, since ambiguity about their next assignment reads as a verdict. And ask them to present the lessons to the group that governs the portfolio, which converts them from people associated with a stopped project into people who supplied evidence the portfolio needed.

What to Harvest From a Pilot You Stop

A stopped pilot leaves assets behind, and most organizations lose them because nobody was assigned to collect them. The harvest is worth planning before the decision meeting, so that if the decision is to stop, the collection starts the same week rather than after the environment has been torn down.

The data work

This is almost always the most valuable thing the pilot produced, and it is the least specific to the model. Getting a model to a testable state usually requires locating records across systems, resolving format inconsistencies, agreeing definitions that were previously ambiguous, and producing a labeled set that a subject matter expert reviewed. All of that has value independent of the model, and much of it supports the next use case in the same domain. Identify the pipelines, the cleaned extracts, the labeled sets, and the definitional agreements, and transfer them to a permanent owner with documentation of how they were produced.

The labeled set in particular is worth protecting. A set of several hundred records reviewed and adjudicated by qualified people is expensive to produce and does not expire when the pilot does.

The evaluation method

If the pilot built a way to measure performance on this class of task, including the held-out set, the adjudication procedure, the agreement statistics, and the reviewer instructions, that method is reusable and should outlive the pilot. Building an evaluation approach that quality assurance accepts is a meaningful piece of work in a regulated setting, and it is a shame to rebuild it every time. How to evaluate model output rigorously is a discipline in its own right, and the harvest step here is simply to make sure the artifact does not disappear with the environment.

The governance artifacts

The risk assessment, the intended use statement, the validation approach, the change control design, and the data provenance documentation are all partly reusable. They were written for this use case, but the structure, the questions asked, and the positions quality assurance took are transferable. Organizations that keep these accumulate a working template. Organizations that do not rewrite them from scratch for every proposal and get different answers each time.

The lessons about the process, not the tool

This is the category most often written down badly, because the natural instinct is to record findings about the technology. “The model struggled with the lower-volume severity bands” is a fact about this pilot and will be stale within a year. The more durable lessons are about how the organization runs this kind of work: that the data extract took nine weeks rather than two and why, that quality assurance needed a different framing of intended use than the team initially offered, that the reviewers who were supposed to test it had no protected time, that the vendor’s model was updated mid-pilot and invalidated three weeks of results. Those findings apply to every pilot that follows.

A practical rule for the closure record. Divide the lessons into two lists: what we learned about this use case, and what we learned about how we run pilots. The second list is the one the next sponsor needs, and it is the one that almost never gets written. Make it a required section rather than an optional one.

Keeping the Option Open: The Named Condition

Stopping a pilot should not permanently close a use case. Most stops in a regulated environment are about timing and prerequisites rather than about the idea being wrong. The data was not ready, the validation path was not settled, the operating owner did not exist, the vendor’s product was immature. Those are conditions, and conditions change.

The way to keep the option open without keeping the pilot open is to write a revisit condition into the closure record. It has three parts, and all three are required for it to function.

THE CONDITION

A specific, observable change

“Historical severity coding remediated across all three sites” or “an operating owner with allocated capacity exists in manufacturing quality” or “the vendor’s model is available under a contract that gives us notice of updates.” Not “when the technology matures.”

THE OWNER

A named person who watches for it

Usually whoever owns the domain rather than whoever ran the pilot, since the pilot team will have moved on. Their responsibility is small: notice when the condition is met and raise it.

THE DATE

A scheduled review even if nothing changes

Twelve or eighteen months out. A calendar entry, not an intention. Without it, the revisit condition is a sentence in a document nobody opens again.

THE STATUS

Visible in the portfolio view

Stopped use cases with live revisit conditions belong on the portfolio register in their own category, so leadership can see them. They are not active work and they are not dead ideas.

There is a governance benefit to keeping this register that goes beyond the individual use cases. A portfolio that shows what was tried, what was decided, and on what evidence is a much stronger artifact in front of a board, or an inspector, than a portfolio that shows only what is currently running. It demonstrates that the organization makes decisions about AI investment rather than accumulating them. That is a maturity signal, and it is the kind of thing that is easy to build if you start and impossible to reconstruct later.

It also has a practical effect on how proposals are written. When sponsors know that stopped pilots are recorded with their reasons, they write sharper proposals, because the closure record of the last attempt at a similar idea is sitting in the same repository and someone will read it. That feedback loop is worth more than most governance mechanisms an organization can install.

What a good revisit looks like in practice

Eighteen months after a classification pilot was stopped for inconsistent historical coding, the remediation program completes. The domain owner raises it. The original closure record is retrieved, and it contains the labeled set, the evaluation method, the quality-accepted validation approach, and a clear statement of what performance was reached and why. The new proposal starts from that position rather than from nothing, and it can be scoped as a four-week confirmation rather than a twelve-week pilot, because most of the uncertainty was already reduced the first time.

That is the return on closing a pilot well. It is not only the reputational damage avoided. It is that the work done the first time remains usable, which is the difference between an organization that learns and one that repeats.

Conclusion

The published evidence on AI project completion is noisier than the confident percentages in circulation suggest, and leaders should treat any single figure with suspicion until they can see the sample and the method. What does hold up across the credible sources is consistent and specific: the dominant reason AI projects fail is not that the models underperform, but that the organization never established what the project was supposed to establish. RAND’s practitioners named it as the leading root cause, the academic interview studies find the same pattern, and the survey data on abandonment shows the consequence at scale.138 The fix is not more governance. It is one page written before the work starts, saying what question this pilot answers, what result means scale, what result means stop, when the decision happens, and who makes it.

The organizations that get the most out of AI in pharma and biotech are not the ones with the highest pilot success rate. They are the ones that can stop things cleanly, harvest what the attempt produced, and move the budget to the next question without anyone’s standing being damaged in the process. That capability compounds. Every well-closed pilot makes the next proposal sharper, the next evaluation faster, and the next stop easier. Sakara Digital works with pharma and biotech organizations building this kind of decision discipline into their AI portfolios, from writing exit criteria that hold up under pressure to structuring the closure records that make the next attempt cheaper. If you are running pilots that have no defined end, or you are about to start one and want the criteria written properly the first time, we are happy to have that conversation.

For Further Reading