In This Article
- Executive Summary
- Why ROI Arrives Too Late to Steer By
- The Adoption Ladder: Five Tiers That Map to a Real Progression
- Substitution: The Tier That Decides Whether Anything Changed
- The Vanity Metrics, Named
- Leading Indicators of Abandonment
- Measuring Adoption in a Regulated Environment
- Small Populations: When Percentages Are Noise
- When Adoption Stalls: Three Causes, Three Different Fixes
- A 90-Day Instrumentation Plan
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Almost every AI program in pharma and biotech is asked the same question at the same moment: what is the return? The question is fair. The timing is wrong. Return on investment is a lagging indicator. It can only be computed after benefits have accumulated, after a baseline has held still long enough to compare against, and after enough time has passed that attribution is arguable. By the time an honest ROI number exists, the decisions that determined it were made months earlier, by people who had no evidence to work from.
There is a second set of numbers that can be measured now and that actually predicts whether a deployment will reach scale. They are adoption metrics, and most organizations measure them badly. They count licenses, logins, and total queries, all of which can rise while a deployment quietly fails. The metric that matters is substitution: whether the old method has actually stopped. A tool used alongside the old process has added work rather than removed it, and no amount of enthusiasm in a usage dashboard changes that.
This article sets out a five-tier adoption ladder (access, activation, habitual use, depth, substitution), names the vanity metrics and explains precisely why each misleads, identifies the leading indicators that predict abandonment before usage falls, and gives an honest account of what can and cannot be measured on a validated GxP system. It closes with the three real causes of a stalled deployment and a practical answer for populations too small for percentages to mean anything.
Why ROI Arrives Too Late to Steer By
Return on investment is a settlement, not a signal. To calculate it honestly you need three things that are rarely available early: a baseline that was measured before the deployment started, an attribution window long enough for the benefit to appear in a financial system, and a counterfactual that tells you what would have happened anyway. Most AI programs in life sciences have none of the three at the point where the steering decisions get made.
That is not an argument against measuring ROI. Sakara Digital has written at length about how to do it properly in Measuring ROI from AI Investments in Life Sciences, and the discipline of building a defensible benefit case is worth the effort. The argument here is narrower and more practical: ROI cannot be your steering instrument, because it reports on a period that has already closed. If the answer arrives in month eighteen, the twelve months of decisions that produced it were made blind.
The industry evidence on what happens in those blind months is now substantial. MIT’s Project NANDA reviewed more than 300 publicly disclosed AI initiatives alongside structured interviews and survey responses from senior leaders, and reported that despite tens of billions of dollars in enterprise investment, roughly 95 percent of generative AI pilots produced no measurable impact on the profit and loss statement.1 McKinsey’s global survey found that the large majority of organizations now use AI in at least one function, while around two thirds have not begun scaling it across the enterprise and only a minority can point to any measurable effect on earnings.2 Gartner has forecast that more than 40 percent of agentic AI projects will be canceled by the end of 2027, citing unclear business value and inadequate risk controls among the reasons.3
Deloitte’s quarterly enterprise survey work has made a related point more gently: organizational change only happens so fast, and most organizations end up setting their own pace toward a return rather than hitting a schedule set by the technology.17 That is a reasonable position for a board to take. It is an impossible position for a program manager, who has to make decisions every week and cannot wait for the pace to reveal itself.
The picture inside pharma and biotech specifically is consistent with that. A 2025 survey of 115 United States technology executives at pharmaceutical and biotechnology companies found that only 40 percent of AI pilots reach scaled deployment, and that the proportion able to demonstrate measurable value varies sharply by function: 49 percent in enterprise technology and data operations, 47 percent in commercial, and 17 percent in research and discovery.4 Those are not failures of arithmetic. They are failures that happened long before anyone tried to do the arithmetic.
Sector research is also starting to separate signal from noise on the clinical side. The IQVIA Institute’s 2026 review of global research and development trends reported a credible signal on AI-enabled programs for the first time, rather than the general enthusiasm that characterized earlier years.18 That is encouraging and it also illustrates the timing problem exactly: the signal is credible now because the programs generating it started several years ago, under conditions where nobody could yet see it.
What a leading indicator has to do
A leading indicator earns its place by satisfying three conditions. It has to move early, before the outcome you care about. It has to be causally connected to that outcome rather than merely correlated with enthusiasm. And it has to be actionable, meaning that when it moves in the wrong direction there is a specific thing you can do about it.
Adoption satisfies all three, but only if it is defined carefully. Adoption defined as “people are using it” satisfies none of them, because that sentence covers everything from a single curious login to a fully substituted workflow. The rest of this article is about making the word precise enough to be useful.
The Adoption Ladder: Five Tiers That Map to a Real Progression
Adoption is not one number. It is a sequence of states that a person passes through, and each transition has a different failure mode and a different fix. Measuring the sequence tells you where the deployment is losing people. Measuring the aggregate tells you nothing except that a number went up or down.
The five tiers below describe that sequence. The important structural rule is that each tier should be reported as a conversion from the tier above it, not as a share of total headcount. A deployment where 90 percent of eligible people have access but only 20 percent of those with access ever activate has a very different problem from one where only 30 percent have access but 85 percent of those activate. Reported as a share of headcount, both look like roughly the same number.
Two pieces of evidence explain why the aggregate number is so unreliable. The 2025 DORA research on AI-assisted software development, drawing on nearly 5,000 responses from technology professionals, found that 90 percent report using AI at work and more than 80 percent believe it has increased their productivity, while around 30 percent report little or no trust in what it produces.14 A single headline usage figure cannot contain that much disagreement and still mean anything. The same research concluded that AI acts as an amplifier of an organization’s existing strengths and weaknesses, and that the largest returns come from the surrounding system rather than the tool.15 If that is right, then the measurements worth taking are the ones that describe the surrounding system, which is what the ladder does.
The second piece of evidence is older and more fundamental. Decades of research into technology acceptance, consolidated in the unified theory of acceptance and use of technology, has consistently found that models predict intention to use a system far better than they predict actual use, with the gap between the two substantial and persistent.16 That is a warning about a specific and common practice: asking people whether they intend to use a tool, or how useful they find it, and treating the answer as an adoption measure. Intention is a real construct and it is not behavior.
Tier 1: Access granted
Access means the account exists, the permission is mapped to the person’s actual role, single sign-on works, and the person can reach the tool from the device and network they normally work on. It is the cheapest tier to achieve and the one most often reported as though it were adoption. It is not adoption. It is a procurement and provisioning outcome.
Access still deserves measurement, because the failures here are real and invisible. Role mappings that were built from an outdated organization chart. Permissions granted to a job title that no longer performs the task. Validated environments where the tool is reachable from a laboratory workstation but not from the office machine where the work is actually written up. Measure access as a conversion from the eligible population, and expect to spend real effort defining that population, which is the single most commonly skipped step in the entire exercise.
Tier 2: Activation
Activation means the person completed one meaningful unit of work with the tool. Not logged in. Not opened the interface and closed it. Completed the smallest unit of work that the tool exists to produce: one draft, one query returned and read, one record classified, one document summarized and used.
The distinction between a login and a completed unit of work is not pedantic. Login-based activation counts are the reason so many deployment dashboards look healthy in month one and collapse in month four. A login records curiosity. A completed unit of work records that the person got far enough to see whether the output was any good.
Tier 3: Habitual use
Habitual use means the tool was used during a normal working week by someone whose role includes the task, in the ordinary course of the work rather than during a training session or a demonstration. The practical definition most organizations can support is something like “used in at least three of the last four weeks by a member of the eligible population.”
Habit takes longer than most deployment plans allow. The best-known field study of habit formation followed 96 volunteers performing a chosen daily behavior over twelve weeks and found that the median time to reach 95 percent of maximum automaticity was 66 days, with a range from 18 to 254 days across individuals.5 That study concerned simple personal behaviors, not multi-step professional tasks inside a controlled process, so the honest reading is that 66 days is a floor rather than a target. A twelve-week pilot that reports low habitual use is not necessarily reporting failure. A twelve-week pilot that reports high habitual use in week three usually is: it is measuring something other than habit, most often mandated participation.
Tier 4: Depth
Depth means the tool is being used for the task it was built for, rather than for a trivial subset of that task. This is the tier that separates a deployment that will scale from one that will look fine on a dashboard and deliver nothing.
The pattern is easy to recognize once you look for it. A system built to draft a full deviation investigation narrative gets used to rephrase a paragraph. A system built to screen an entire literature corpus gets used on the five papers someone already found. A system built to reconcile a data set gets used to check a single field. In each case, usage is genuine, activation is high, habit may even be forming, and the intended benefit is absent because the intended task is not the task being performed.
Measuring depth requires defining, before deployment, what the primary intended use is and what a complete instance of it looks like. That definition should be a short written statement produced with the process owner, not with the vendor. Depth is then the share of habitual users whose usage includes the primary intended use at least once in the measurement period, and, more usefully, the share of eligible task instances that were performed end to end with the tool.
Tier 5: Substitution
Substitution means the old method has stopped. Not declined. Stopped, for a defined scope of work, with the legacy path formally retired for that scope. This is the tier that determines whether the deployment removed work or added it, and it is covered in full in the next section.
| Tier | What it measures | Denominator | What a drop here tells you |
|---|---|---|---|
| 1. Access | Account exists, permission matches the real role, tool is reachable from the working environment | Eligible population (people whose job includes the task in the period) | Provisioning, role mapping, or environment problem. Rarely a people problem. |
| 2. Activation | Completed one meaningful unit of work, not a login | People with access | Onboarding, first-run experience, or the tool failed on the first honest attempt. |
| 3. Habitual use | Used in a normal working week (for example, 3 of the last 4 weeks) | Activated users | The tool did not survive contact with real workload, or there is no trigger in the workflow. |
| 4. Depth | Used for the primary intended task, end to end, not a trivial subset | Habitual users, and separately, eligible task instances | Output quality is not good enough for the full task, or scope was never communicated. |
| 5. Substitution | Legacy path formally retired for a defined scope of work | Legacy transaction volume at baseline | The deployment is additive. It has increased total work, whatever the usage numbers say. |
Substitution: The Tier That Decides Whether Anything Changed
Substitution is the tier that matters most and the one almost nobody measures, for a straightforward reason: it is the only tier whose data does not live in the new tool. Every other tier can be read from the AI system’s own telemetry. Substitution has to be read from the legacy system, the old queue, the shared mailbox, the template document, the spreadsheet, the paper form. Nobody instruments those, because they are the things everyone assumes are going away.
They usually do not go away. The health informatics literature is the richest available body of evidence on what actually happens when a new system is introduced into a documented, controlled clinical process, and its consistent finding is that the old method persists alongside the new one. A study across 11 primary care clinics at three benchmark institutions, observing 120 staff and providers, documented ten categories of paper-based workarounds and five categories of computer-based workarounds, with efficiency, memory, and awareness as the most consistent reasons given across all three institutions.6 A separate cross-case study of eleven practices found double documentation and duplicate data entry, scanning of paper documents, reliance on recall for information the system made inaccessible, and freestanding tracking systems maintained in parallel.7 Reviews of this literature have argued that studying workflow and workarounds systematically is a prerequisite for improving health system performance, precisely because the workarounds are where the real process lives.8
None of those studies are about AI. That is the point. The persistence of the old method is not a property of AI tools. It is a property of introducing anything into a process that has not itself been changed. AI deployments inherit the pattern in full, and they inherit it in a sharper form, because AI output usually needs review before it can be used, and review is exactly where the old method reattaches itself.
The one-line substitution test. What did you turn off? If the honest answer is nothing, then whatever the usage dashboard says, the deployment has added a step to the process rather than replaced one. Total work has gone up. The benefit case, whenever it is eventually calculated, will be calculated against a process that now runs two ways at once.
How to measure substitution concretely
Substitution is measurable, but it requires deliberately instrumenting things you are trying to retire. Four measures cover most situations.
- Legacy transaction volume. Count how many times the old path was used per period: records created in the legacy system, messages into the shared mailbox, copies of the old template opened, requests submitted on the old form. You need a baseline number from before the deployment started, which means this has to be set up early. A substitution rate is meaningless without a before.
- Share of final records produced by the new path. For any process with a defined output, the question is which system produced the record that became official. Not which system was consulted. Which one produced the thing that goes into the file.
- Retirement events. These are binary and they are the most honest measure available. The step was removed from the procedure. The old form was withdrawn from the document management system. The scheduled legacy report stopped being generated. The legacy queue was closed to new items. Each of these has a date and an owner, and each is verifiable by someone other than the deployment team.
- Reconciliation effort. Where both paths still run, someone is comparing them. Count that work explicitly. It is the clearest evidence of an additive deployment, and it is usually invisible in every report because it is performed by people who were not part of the pilot.
The honest exception: deployments that are additive by design
Some AI deployments genuinely do not replace anything, and pretending otherwise produces a false failure signal. Screening a corpus that was previously never screened, monitoring a data stream nobody had capacity to watch, generating a second independent read on a decision that previously had one: these create new capability rather than displacing old work. For these, substitution is the wrong tier-five metric and applying it will make a good deployment look bad.
The replacement metric for additive deployments is coverage: the share of the eligible population of items, cases, records, or events that the new capability actually reaches, measured against the total that exists rather than against the total that was previously handled. A monitoring capability that reaches 8 percent of the relevant data stream has not scaled, even if everyone who uses it loves it.
What is not acceptable is deciding after the fact that a deployment was additive by design because substitution turned out to be zero. Whether the deployment is substitutive or additive should be written down at the start, in the same document that defines the eligible population and the primary intended use. If it is substitutive, it needs a named legacy path and a target retirement date. If it is additive, it needs a coverage denominator. Deployments that cannot answer which one they are have not been scoped.
The Vanity Metrics, Named
Three numbers appear on nearly every AI deployment dashboard in the industry, and all three can rise steadily while the deployment fails. They are worth naming individually, because each misleads for a different reason and each has a specific replacement.
Total queries
Total queries, prompts, messages, or interactions is an unbounded numerator with no denominator. It rises with headcount, with time, and with retries. That last one matters most: a user who asks the same question four different ways because the first three answers were wrong contributes four times as much to the total as a user who got a usable answer immediately. A rising query count is therefore consistent with a tool that is getting worse.
Total queries is also almost always concentrated. In most deployments a small group of enthusiastic users generates the majority of the volume, which means the headline number can grow while the median user does nothing at all. Reporting a total conceals exactly the distribution you need to see.
The replacement: report the distribution, not the sum. The median number of interactions per eligible user, the share of eligible users at or above a defined usage threshold, and the share of total volume generated by the top decile of users. If the top decile generates 80 percent of the volume, you have a pilot with extra steps, not a deployment.
Licenses assigned
Assigning a license is a procurement event. It records a decision made by a budget holder, not a behavior performed by a user. Software asset management data makes the size of the gap plain: analysis of a large managed portfolio of enterprise software found license utilization at 54 percent in 2025, improved from 47 percent the year before, with average annual license waste of around 19.8 million dollars per organization studied.9 Roughly half of assigned seats, across a very large sample, were not being used.
Life sciences organizations have a particular version of this problem, because AI capability often arrives bundled inside a productivity suite or a platform the company already licenses. In that situation “licenses assigned” can jump to 100 percent of the workforce on a single day, purely as a contract event, with no relationship to anything happening in the work.
The replacement: the assigned-to-active conversion rate, measured over a defined period, against the eligible population rather than total headcount. And a standing question at every deployment review: how many assigned seats have never reached tier two?
Users who logged in once
Cumulative unique users is a ratchet. It can only go up. Once a person has logged in, they are permanently included, whether they returned the next day or never opened the tool again. Any metric that cannot decline cannot warn you about anything.
The fix is well established in ordinary product analytics and is worth borrowing directly: use period-bounded active counts, where an active user is defined as a unique user who interacted at least once during the reporting period, and report both daily and monthly figures.10 A period-bounded number can fall, which is the entire reason to use it.
The replacement: period-bounded actives at tier two or above, plus the ratio of period actives to cumulative uniques. That ratio is one of the fastest ways to see a deployment decaying: when it drops steadily while onboarding continues, new users are arriving faster than existing users are leaving, and the headline count keeps rising anyway.
Two more that deserve the same treatment.
- Pilot satisfaction scores. Pilot cohorts are usually volunteers, and volunteers are the most favorably disposed population you will ever measure. A high satisfaction score from a self-selected group predicts almost nothing about the second wave. Report satisfaction separately for volunteers and for people assigned to the tool without being asked.
- Self-reported time saved. Asking people how much time a tool saved them produces an estimate influenced by the framing of the question, by what people believe leadership wants to hear, and by the difficulty of recalling how long the old task actually took. If time saved matters to the benefit case, it needs a measured before and after on a defined unit of work, not a survey.
Leading Indicators of Abandonment
By the time aggregate usage falls, abandonment has already happened and been happening for a while. Three signals move earlier, and all three are measurable with data most deployments already have.
Falling repeat-use rate among early adopters
Early adopters are the most forgiving population in the organization. They chose to be there, they tolerate rough edges, and they will work around problems that would stop anyone else. When their repeat-use rate starts to fall, the tool has failed a group that wanted it to succeed. That is a much stronger signal than the same decline in a general population.
This signal is invisible in aggregate reporting, because early-adopter decline is usually masked by continued onboarding. The total number of active users can rise for months while every individual cohort decays. The measurement that exposes it is cohort retention: group users by the week they first activated, then track each cohort’s return rate over subsequent weeks separately. If the week-four return rate is lower for the March cohort than for the January cohort, the deployment is getting worse at retaining people, regardless of what the total says.
A widening gap between the pilot team and everyone else
The pilot team almost always had advantages that were never written down: physical or organizational proximity to the people who built the tool, an informal channel for getting problems fixed within hours, an exception to a control that slowed everyone else down, a colleague who knew the prompts that worked. None of those advantages travel to the second wave.
The measurement is a direct comparison of the same tier at the same elapsed time. Take the pilot team’s habitual-use rate at week eight after their own start date, and compare it with the second wave’s habitual-use rate at week eight after theirs. If the second wave is materially below the pilot at equivalent elapsed time, and the gap is not closing by week twelve, the deployment depends on something the pilot team had and nobody else does. That is a scaling defect, and it will not be fixed by more training. It will be fixed by finding the undocumented advantage and either building it into the product or building it into the support model.
Requests to export output into the old system
This is the clearest structural signal available, and it arrives early because users raise it themselves. Requests for a comma-separated export. Requests for a print view. Requests for an integration that pushes the output back into the legacy application. Requests for the output in the format of the old template. Users asking how to copy the result into the system they were supposed to have stopped using.
Every one of these requests is the same message stated as a feature request: the tool has become a step in the old process rather than a replacement for it. It is worth treating export requests as a tracked category in the help desk taxonomy, separate from general enhancement requests, precisely because they carry this diagnostic weight. A deployment where export requests are the top enhancement category has a substitution problem that no amount of usage growth will resolve.
| Signal | What to measure | What it means | First action |
|---|---|---|---|
| Falling repeat use among early adopters | Cohort retention by week of first activation, tracked separately per cohort | The tool failed the most forgiving users available | Interview the lapsed early adopters individually, before changing anything |
| Widening pilot-to-wave gap | Habitual-use rate at equal elapsed weeks, pilot versus each later wave | The pilot depended on an undocumented advantage | Find the advantage and productize it, or reproduce it in the support model |
| Export and copy-back requests | Help desk requests tagged as export, print view, or legacy integration | The tool is a step in the old process, not a replacement | Re-open the workflow design. This is a process problem, not a product one. |
| Concentration of use | Share of total volume from the top decile of users | A pilot is being reported as a deployment | Report the median user, not the total, at every review |
| Drifting depth | Share of usage on the primary intended task versus trivial subsets | Output quality is insufficient for the full task | Sample the failed full-task attempts and read them |
| Ticket language shift | Ratio of “how do I do X” to “how do I get around X” tickets | Users have moved from learning to circumventing | Treat the workaround as the requirement it is revealing |
Measuring Adoption in a Regulated Environment
Everything above assumes you can instrument the tool and report on who used it. In a GxP environment that assumption needs qualification in three separate directions, and being honest about the limits is more useful than pretending they do not exist.
Adding telemetry to a validated system is itself a change
If the AI capability sits inside a validated computerized system, adding usage logging is a configuration change. It requires an impact assessment, it goes through change control, it may require regression testing depending on how the logging is implemented, and it updates the validated state of the system. That is not an argument against doing it. It is an argument for doing it at the right time.
The practical consequence is that the measurement architecture has to be specified before validation, not after deployment. Telemetry requirements belong in the user requirements specification alongside functional requirements. Teams that skip this find themselves three months into a deployment, unable to answer basic questions about usage, facing a change control cycle to add the instrumentation that would have answered them, and by the time the change is through the questions have been overtaken by events.
The audit trail is not a usage analytics data set
It is tempting to solve the telemetry problem by reading the audit trail. The audit trail already records who did what to which record and when, it is already validated, and it is already retained. Using it for adoption reporting is technically possible and legally more complicated than it looks.
The audit trail exists for a specific regulatory purpose: reconstructing the history of a GxP record. Repurposing attributable audit trail entries to report on individual staff behavior is a new processing purpose applied to personal data, and it engages a different set of obligations. Regulators supervising workplace monitoring have been explicit that monitoring workers requires a documented lawful basis, transparency to the people being monitored, and a data protection impact assessment where the monitoring is systematic.11 In several European jurisdictions there is a further layer: works councils hold co-determination rights over the introduction of technical systems capable of monitoring employee performance, which means the decision is not the sponsor’s alone to make.
The workable answer is to keep the two data sets separate by design. Attributable audit trail data stays inside its GxP purpose. Adoption reporting runs on aggregated, role-level, non-attributable counts, with a stated minimum group size below which no figure is published. This is not only the defensible position, it is usually the more useful one, because individual-level adoption reporting invites the exact behavior that makes the metrics worthless: people using the tool to be seen using the tool.
The sixth tier that only regulated deployments need
Where AI output feeds a decision that is itself controlled, adoption metrics need one tier the commercial frameworks do not include: conformance of the controlled step. Was the required human review actually performed, by someone qualified, and is the record of that review complete and contemporaneous?
This matters because usage and conformance can move in opposite directions. A team that has fully adopted a tool and quietly stopped performing the review step has high adoption and a compliance problem. FDA’s draft guidance on AI supporting regulatory decision-making frames credibility around the specific context of use, and the credibility of a model whose human oversight step has silently lapsed is not what the plan said it was.12 Notably, that framework applies where AI output supports safety, effectiveness, or quality determinations. It expressly does not cover uses aimed purely at operational efficiency, which is where a great deal of pharma AI actually sits. Knowing which side of that line a given deployment falls on determines how much of this tier you need.
What you can measure without touching the validated system
Several of the most valuable measures require no telemetry at all, which makes them available immediately and unaffected by change control:
- Legacy path volume. Count the old queue, the old form, the old report. This is substitution data and it usually lives outside the validated system.
- Retirement events. Procedure revisions, form withdrawals, decommissioned reports. These are document control records, already maintained, already dated, already attributable to an owner.
- Help desk taxonomy. Ticket categories and volumes, particularly the export-request category described earlier.
- Task instance counts. How many deviations, batch reviews, literature screens, or data reconciliations occurred, and how many of those used the new path. This is process data, not system telemetry.
- Structured user conversations. Underrated, and in small populations superior to anything a dashboard produces, as the next section argues.
Small Populations: When Percentages Are Noise
A great deal of regulated work is performed by small groups. Seven qualified persons. Eleven validation engineers. Fourteen people who perform a specific batch review. Nine medical writers with a particular therapeutic area. In populations this size, percentages are not merely imprecise. They are actively misleading, because a single person changes the headline figure by ten to fifteen points and the reader has no way to know that.
Reporting “adoption fell from 71 percent to 57 percent this month” when the underlying change is one person out of seven taking parental leave is worse than reporting nothing, because it triggers action against a signal that does not exist. Six practical rules make small-population measurement honest.
Six rules for measuring adoption in small populations
- Report counts, not percentages, below a stated denominator. Set a floor (30 is a reasonable default) and below it write “5 of 7” rather than “71 percent”. The count carries its own uncertainty; the percentage hides it.
- If you must use a proportion, use an interval method that behaves at small N. The normal approximation is unreliable at small sample sizes and near the extremes. The Wilson score interval is the standard alternative and gives coverage much closer to nominal in exactly these conditions.
- Do not read zero as evidence of absence. The rule of three gives the upper bound: observing zero events in n trials leaves a 95 percent upper limit of roughly 3/n.13 With seven users and zero abandonments, the true abandonment rate could still be as high as about 43 percent. A clean small pilot is not evidence of a clean deployment.
- Switch from sampling to census. With twelve eligible users you do not survey a sample, you talk to all twelve. Twelve structured fifteen-minute conversations produce more usable information than a survey with five responses, and take about the same amount of calendar time.
- Count task instances, not people. A team of eight that performs 400 batch reviews a year gives you 400 observations rather than eight. Wherever the unit of analysis can be moved from the person to the work item, the statistics improve immediately and the privacy position improves with them.
- Use run charts rather than cross-sectional percentages. Plot the count over time for the small group and look for shifts (a run of consecutive points on one side of the median) and trends. These patterns are interpretable at small N in a way that a month-over-month percentage change is not.
One further discipline is worth adopting: pre-register the threshold. Before the pilot starts, write down what count would constitute success, what count would trigger a redesign, and what count would end the deployment. In a small population, the temptation to interpret the result after seeing it is overwhelming, and the numbers are small enough to support almost any interpretation. Writing the threshold down in advance is a cheap protection against that, and it makes the eventual decision defensible to people who were not in the room.
None of this is an argument that small populations cannot be measured. It is an argument for measuring them with methods appropriate to their size. The most common failure in regulated AI deployments is not the absence of measurement. It is applying consumer product analytics, designed for populations of hundreds of thousands, to a group of eleven people, and then acting on the result.
When Adoption Stalls: Three Causes, Three Different Fixes
Adoption stalls have three real causes. Distinguishing between them matters enormously, because the fixes are unrelated to each other and applying the wrong one wastes a full deployment cycle. The default organizational response, more training, addresses none of the three in most cases.
The tool is not good enough
Users try it, read the output, and redo the work themselves. Depth is low, rework is high, and complaints are specific and reproducible rather than general. This is a product problem and it is fixed by product work: better retrieval, better grounding, a narrower scope where quality is acceptable. Running more training against a quality problem teaches people to use something that does not work.
The workflow was never changed
Users like the tool. Depth may even be fine. Substitution is zero, because the procedure still names the old system, the approval routing still points at the legacy form, and the definition of done still requires the old artifact. This is a documented-change problem. It is fixed by revising procedures, forms, routing, and acceptance criteria, which is slower and less enjoyable than product work and is the actual bottleneck in most regulated deployments.
The incentive rewards the old method
Usage is high in groups where nothing personal depends on the old method and low where individual performance is measured on something the old method optimizes. The reviewer measured on findings per hour has no reason to want pre-screening. The analyst whose standing rests on producing the legacy report has no reason to retire it. This is fixed by changing what is measured and rewarded, or by accepting the ceiling.
“People need more training”
Training fixes one specific and comparatively rare failure: capability. The test is simple. Sit with a user and watch them attempt the task. If they can perform it correctly when observed, capability is not the problem and training will not change the outcome. In most stalled deployments users know exactly how to use the tool and have concluded, correctly, that doing so does not help them.
Diagnosing in the right order
Test the three causes in order of how cheaply they can be ruled out, which happens to also be the order in which the fixes get progressively harder.
Read twenty real outputs
Not demonstration outputs. Twenty actual attempts at the primary intended task, pulled from real use, read by someone who knows what good looks like in that process. If a meaningful share are not usable without substantial rework, stop here. You have a quality problem and nothing downstream will fix it.
Read the procedure
Open the current effective version of the SOP, work instruction, or form that governs the task. Does it require the old system by name? Does the approval route still terminate in the legacy application? Does the definition of a complete record still reference the old artifact? If yes, adoption is capped by the document, not by the users, and the fix is a change control.
Ask what each person is measured on
For the group with the lowest habitual use, write down what their individual performance is actually assessed against, including informal expectations. Then ask whether using the tool helps or hurts that number. This conversation is uncomfortable and it usually produces the answer in under an hour.
Only then consider enablement
If output quality holds up, the procedure has been changed, and the incentives are neutral or supportive, and adoption is still low, then capability and confidence are genuinely worth addressing. At that point enablement works, because there is nothing else standing in the way.
The order matters because each step is progressively more expensive to run and progressively harder to reverse. It also matters politically. Organizations reach for training first because training is the intervention that blames nobody and requires no one to change a document or a target. That is precisely why it so rarely works.
There is a fourth possibility worth naming, which is that the deployment was scoped against a task that does not need doing. This is rarer than the three causes above but it does happen, particularly where a tool was selected before a problem was defined. The tell is that adoption is low, the tool works, the procedure permits it, the incentives are neutral, and users describe the task itself as low value. In that case the correct response is to stop the deployment and redirect the capability, which is a good outcome reached early rather than a failure.
A 90-Day Instrumentation Plan
The measurement work described here is not large, but almost all of it has to happen before or alongside the deployment rather than after it. The following sequence fits inside a normal deployment timeline and produces a usable adoption read by day 90.
Days 1 to 15: define the eligible population and the unit of work
Write down who is expected to perform which task, how often, and what a complete instance of that task looks like. This is the denominator for every tier and the definition of depth. It is produced with the process owner, not the vendor, and it is the step teams skip most often.
Days 1 to 15: settle the measurement architecture
Decide what telemetry is needed, where it will live, at what level of aggregation it will be reported, how long it will be retained, and on what lawful basis. If the system is validated or heading for validation, these become requirements now rather than a change control later. Agree the minimum group size below which no figure is published.
Days 16 to 30: baseline the legacy path
Count current volume on the old method and record it. Name the legacy path explicitly, name its owner, and state whether the deployment is substitutive or additive by design. If substitutive, put a target retirement date on the plan. Without this baseline, substitution cannot be measured at any point later.
Days 31 to 60: instrument the ladder and publish conversions
Report the five tiers as conversion rates, weekly to the deployment team and monthly to the sponsor. Report counts alongside percentages, and the median user alongside the total. Add the export-request category to the help desk taxonomy on day 31, before anyone needs it.
Days 61 to 90: first cohort read and first stall diagnosis
Run cohort retention by week of first activation. Compare the second wave against the pilot at equal elapsed time. If any tier conversion is below the pre-registered threshold, run the four-step diagnosis in order and record which cause was found. Do not run training before step three is complete.
Ongoing: fold adoption into periodic review
For validated systems, adoption reporting belongs inside the existing periodic review cycle rather than running as a parallel exercise. That places it in front of the people who can act on it, keeps it in a controlled document, and makes the retirement of the legacy path a reviewable commitment rather than an aspiration.
What this produces by day 90
At the end of the first quarter you will not have an ROI number, and you should not pretend to. What you will have is a defensible answer to five questions that ROI cannot answer for another year: how many eligible people can actually reach the tool, how many of those got far enough to complete real work with it, how many use it in an ordinary week, how many use it for the task it exists for, and whether anything has been turned off. Those five answers are enough to decide whether to expand, redesign, or stop.
They are also enough to have a credible conversation with a finance function that is asking for the return. “We cannot tell you the return yet, and here is the evidence that we are on a path to one” is a substantially stronger position than a benefit estimate assembled from self-reported time savings. Organizations that build the adoption read early tend to find that the ROI conversation gets easier, not harder, because by the time the question is formally asked they have twelve months of leading evidence behind the answer.
Conclusion
The reason adoption measurement is done badly is not that the metrics are difficult. It is that the easy metrics are available immediately from the vendor dashboard and the useful ones require deciding, in advance, who is supposed to be doing what and what is supposed to stop. That work is uncomfortable because it exposes scope decisions that were never really made. It is also the work that determines whether a deployment reaches scale, which is why the discomfort is worth having early rather than in month eighteen when the return question arrives and the honest answer is that nobody knows.
Our consistent observation across pharma and biotech deployments is that the single question separating programs that scale from programs that stall is the substitution question, and that it is almost never on the dashboard. Usage can look excellent in a deployment that has quietly doubled the work, because the old path is running in parallel and nobody is counting it. Ask what was turned off. Ask for the date and the owner. If there is no answer, the adoption numbers are describing activity rather than change, and the return will eventually reflect that.
Sakara Digital works with pharma and biotech organizations designing the measurement layer around AI deployments in validated and non-validated environments, including the parts that have to be settled before a system reaches validation. If you are deploying AI capability and want an independent view on what to instrument, what you are allowed to instrument, and what the early numbers are really telling you, we are happy to have that conversation.
For Further Reading
For Further Reading
- Measuring ROI from AI Investments in Life Sciences
- Moving from AI Pilot to Enterprise Scale in Regulated Environments
- From AI Pilots to Production: Why 95% of Pharma AI Projects Fail
- AI Maturity Self-Assessment: A 30-Minute Diagnostic for Pharma IT Leaders
- Change Management for Digital Transformation in Pharma
- Upskilling the Life Sciences Workforce for AI: A Practical Approach
References & Sources
- Challapally, A., et al. “The GenAI Divide: State of AI in Business 2025.” MIT Project NANDA, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- QuantumBlack, AI by McKinsey. “The State of AI: Global Survey.” McKinsey & Company, 2025. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Gartner. “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027.” Press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- ZS. “Scaling AI in Pharma and Biotech: 2026 ZS CDIO Research.” Survey of 115 US technology executives conducted July 2025. https://www.zs.com/insights/scaling-ai-in-pharma-cdio-2026
- Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., and Wardle, J. “How are habits formed: Modelling habit formation in the real world.” European Journal of Social Psychology, 40(6), 2010, 998-1009. https://onlinelibrary.wiley.com/doi/10.1002/ejsp.674
- Flanagan, M. E., Saleem, J. J., Millitello, L. G., Russ, A. L., and Doebbeling, B. N. “Paper- and computer-based workarounds to electronic health record use at three benchmark institutions.” Journal of the American Medical Informatics Association, 20(e1), 2013, e59-66. https://pubmed.ncbi.nlm.nih.gov/23492593/
- Cifuentes, M., et al. “Electronic Health Record Challenges, Workarounds, and Solutions Observed in Practices Integrating Behavioral Health and Primary Care.” PMC7304941. https://pmc.ncbi.nlm.nih.gov/articles/PMC7304941/
- Zheng, K., Ratwani, R. M., and Adler-Milstein, J. “Studying Workflow and Workarounds in Electronic Health Record-Supported Work to Improve Health System Performance.” Annals of Internal Medicine, 2020. https://pubmed.ncbi.nlm.nih.gov/32479181/
- Zylo. “SaaS Statistics: license utilization and waste, 2026 SaaS Management Index.” https://zylo.com/blog/saas-statistics
- Microsoft. “Analyze active user metrics.” Microsoft Copilot Studio documentation, Microsoft Learn. https://learn.microsoft.com/en-us/microsoft-copilot-studio/analytics-active-users
- Information Commissioner’s Office. “Monitoring workers: guidance for employers.” ICO, UK GDPR guidance and resources. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/employment/monitoring-workers/
- US Food and Drug Administration. “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products.” Draft guidance, January 2025. https://www.fda.gov/media/184830/download
- Hanley, J. A., and Lippman-Hand, A. “If nothing goes wrong, is everything all right? Interpreting zero numerators.” JAMA, 249(13), 1983, 1743-1745. https://pubmed.ncbi.nlm.nih.gov/6827763/
- DORA. “State of AI-assisted Software Development 2025.” Research report based on nearly 5,000 survey responses. https://dora.dev/dora-report-2025/
- Google Cloud. “Announcing the 2025 DORA Report.” Google Cloud Blog, September 2025. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
- Venkatesh, V., Thong, J. Y. L., and Xu, X. “Unified Theory of Acceptance and Use of Technology: A Synthesis and the Road Ahead.” Journal of the Association for Information Systems, 17(5), 2016. https://aisel.aisnet.org/jais/vol17/iss5/1/
- Deloitte. “State of Generative AI in the Enterprise.” Deloitte AI Institute quarterly survey series. https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-generative-ai-in-enterprise.html
- IQVIA Institute. “Global R&D Trends 2026 Report Finds Credible Signal on AI-Enabled Programs.” IQVIA, May 2026. https://www.iqvia.com/blogs/2026/05/iqvia-institutes-global-r-and-d-trends-2026-report-finds-credible-signal-on-ai-enabled-programs








Your perspective matters—join the conversation.