In This Article
- Executive Summary
- What Your Current Metrics Actually Measure
- The Statistics You Should Stop Repeating
- Six Behavioral Measures, Defined Precisely
- Where the Data Comes From
- How Each Measure Fails When It Gets Gamed
- A Measurement Plan for One System Rollout
- Reporting a Number That Got Worse
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Most digital transformation programs in pharma and biotech report the same four numbers: licenses deployed, active users, training completions, and go-live dates hit. All four are real, all four are easy to collect, and all four have a property that should worry you. They only move in one direction. None of them can tell you whether the work changed how a decision gets made, and none of them can go down when the program is failing.
This article proposes six behavioral measures that can go down: decision completion rate, time to first value for a new user, task effort score, rework rate, workaround rate, and residual work in the system the new one was supposed to replace, with shadow spreadsheets treated as a separate signal. Each is defined precisely enough to implement, sourced to data you already generate, and paired with the specific way it fails once someone has an incentive to make it look good.
It also deals with the sourcing problem. This topic attracts recycled statistics, and the most famous one, that 70 percent of transformations fail, does not survive contact with its own citation trail. We trace it, explain what the underlying research actually measured, and then set out a measurement plan for a single system rollout: the baseline to capture before go-live, checkpoints at 30, 90, and 180 days, who reports each number, and how to present a measure that got worse without the program being cancelled.
What Your Current Metrics Actually Measure
Pick up the last steering committee pack for any large system program in a life sciences organization. A laboratory information management system replacement, an electronic quality management system rollout, a clinical data platform, a manufacturing execution system extension. The status slide will carry some version of the same four numbers.
Licenses deployed. Unique users who logged in during the reporting period. Training modules completed, usually with a percentage against the assigned population. Milestones hit against the plan, with go-live at the end.
Take each one seriously for a moment, because they are not worthless. They are just measuring something other than what the program was funded to change.
Licenses deployed measures procurement, not use
A license count is a record of a contractual commitment. It tells you how much capacity the organization has bought and how much it is paying. It is a perfectly good number for a vendor management conversation and it belongs in the budget review. It carries no information at all about whether anyone opened the software. In programs where licenses were bought in a block at contract signature, the number is fixed on day one and never moves again, which makes it useless as a progress measure and yet it keeps appearing on progress slides.
Logins measure presence, not work
Active user counts are the most seductive of the four, because they feel like behavior. A person logged in. Something must have happened.
What actually happened is unknown. In regulated environments a large share of logins are compliance-driven: someone opens the system to acknowledge a document, apply an electronic signature to something decided elsewhere, or check a status a colleague asked about in a meeting. A single sign-on arrangement can inflate the count further, because the user never experiences a login at all and the count reflects session creation rather than intent. If your monthly active user figure includes people who opened the system once to approve a record whose content was agreed in a meeting three days earlier, the number is measuring meeting follow-up, not adoption.
Training completions measure attendance, not capability
A completion record confirms that a person was assigned a module and marked it complete. In GxP settings this record has genuine value, because training records are part of the controlled documentation and inspectors look at them. It is evidence that a qualification requirement was met.
It is not evidence that the person can now do the task. Completion is scored against exposure, not performance, and the assessment attached to most system training is a short knowledge check written by the same team that built the module. Training effectiveness beyond a read-and-understand acknowledgment is a large subject in its own right, and one worth separating from adoption measurement entirely.
Go-live dates measure project management, not outcomes
Hitting a go-live date proves that the project team delivered a build into a production environment on a schedule. That is real work and it deserves credit. But go-live is a project event, not a business event. It says nothing about whether the process moved. Every experienced program leader has been in an organization where the system went live on time and the work carried on in the old place for another eight months.
The shared property. Licenses, logins, training completions, and milestones can all be driven up by program activity alone, without a single decision changing. They are cumulative or near-cumulative, so they trend upward by construction. That is exactly why they are attractive. A metric that only goes up produces a report that is never uncomfortable, and a report that is never uncomfortable is never read carefully.
Why programs choose them anyway
It would be unfair to treat these metrics as a failure of intelligence. They are chosen for reasons that make sense from inside a program.
They are cheap. Every one of them comes out of a system that already reports it, with no design work, no definition negotiation, and no data collection burden on the business. They are unambiguous, so nobody argues about the definition in a steering meeting. They are available on day one, which matters when a sponsor wants a status number six weeks in. And they are defensible, because when a program is challenged, a rising line is easier to present than an explanation of why a behavioral measure moved sideways for a quarter.
There is a final reason, less comfortable. Activity metrics protect the people reporting them. If the only numbers on the slide are ones that cannot fall, the reporting relationship never becomes adversarial. That is a real organizational benefit and it is worth naming honestly, because any replacement set of measures gives that protection up. The measures proposed later in this article can all get worse. Adopting them is a decision to accept a harder conversation in exchange for information you can act on.
The question the metrics do not answer
The purpose of most digital transformation work in pharma and biotech is not to install software. It is to change where and how decisions get made: which data a batch disposition decision rests on, how a deviation gets triaged, whether a stability trend is spotted by a person scanning a report or by a rule that fires, how long a study team waits for a query to resolve.
None of the four standard metrics can detect a change in any of those. A behavioral measure is one that moves when the decision path moves and stays flat when it does not. That is the whole idea, and everything below is an attempt to make it operational.
The Statistics You Should Stop Repeating
Before proposing new numbers, it is worth cleaning up the old ones. Digital transformation is unusually prone to statistics that circulate for years without anyone checking where they came from. If you are going to ask a board to accept measures that can go down, you cannot open the argument with a figure that falls apart under inspection.
The 70 percent failure claim
The most repeated statistic in this field is that roughly 70 percent of transformations fail. It appears in vendor decks, conference keynotes, and consulting proposals, usually attributed to McKinsey, to Kotter, or to no one at all.
Trace it properly and the trail is short. McKinsey’s own December 2021 transformation report footnotes the 70 percent figure not to a study but to two of John Kotter’s books, Leading Change (1996) and A Sense of Urgency (2008), where it appears as an authorial assertion rather than a research finding.3 Kotter’s original 1995 Harvard Business Review article, the piece most often credited with the number, does not contain it. That article draws on his observation of more than 100 companies over a decade and describes eight errors that derail change efforts. Its characterization of outcomes is qualitative: a few efforts were very successful, a few were outright failures, and most fell somewhere in between with a tilt toward the lower end. The only percentage in the piece is a separate observation that well over 50 percent of the companies he watched failed at the first stage, establishing urgency.2
Mark Hughes examined this directly in the Journal of Change Management in 2011, reviewing five separate published instances of the 70 percent claim and tracing each back to its cited origin. His conclusion was that there is no valid and reliable empirical evidence supporting the figure, and that the narrative persists through citation of citation rather than through data.1 The trail generally leads back to informal estimates made in early 1990s business process reengineering literature about a specific method, not to a measured failure rate for organizational change in general.
Practical consequence. If someone opens a business case with the 70 percent figure, the correct response is not to argue about whether transformations succeed. It is to ask what the number measured, over what population, using what definition of failure. In this case there is no answer, because there was never a measurement.
What the McKinsey survey did measure
The 2021 McKinsey report is worth reading on its own terms, provided you describe it accurately. It was an online survey in the field from May 18 to June 29, 2021, with 1,034 respondents drawn from a range of regions, industries, company sizes, and functions. Every respondent had been part of a transformation in the previous five years at their current or a previous employer. Responses were weighted by each respondent’s nation’s contribution to global GDP.3
Fewer than one third of those respondents said their organization’s transformation had been successful at both improving performance and sustaining the improvement over time. Respondents at organizations reporting success estimated that they had realized on average 67 percent of the maximum financial benefit available; respondents elsewhere estimated 37 percent. The report also allocates where value was lost: roughly a quarter during target setting, 55 percent during and after implementation, and 20 percent after implementation was complete.3
Every one of those numbers is a self-reported opinion, collected from memory, about an event the respondent participated in. The success measure combines two judgments in a single question. The benefit figures are respondent estimates of a counterfactual maximum that nobody observed. This is useful directional evidence about how practitioners perceive transformation outcomes. It is not a failure rate, and it should never be presented as one.
The other statistic with the same problem
A parallel case is instructive because the debunking is more technical. The Standish Group’s CHAOS report, source of the widely quoted claim that only around 16 percent of software projects succeed, was examined by Eveleens and Verhoef in IEEE Software in 2010. They identified four problems: the success definition rests entirely on estimation accuracy for cost, time, and functionality rather than on whether the software delivered value; the accuracy measure is one-sided, so it penalizes overruns but treats padded estimates as success; steering an organization on that definition rewards deliberately inflated estimates; and the aggregate figures average numbers whose bias is unknown.4
The lesson generalizes. A headline percentage is only as good as its denominator and its definition of the event being counted. Before you use one, you need three things: the population, the operational definition of success or failure, and the collection method. If you cannot state all three, do not use the number. Writing the sentence without it is almost always fine.
Why this matters for your own metrics
The discipline you apply to somebody else’s statistic is the same discipline you owe your own. Every measure proposed in the next section is specified with a numerator, a denominator, a time window, and a data source, because a metric without those four things becomes a number people argue about instead of a number people act on.
Six Behavioral Measures, Defined Precisely
These six are chosen because each can move in either direction, each detects a different kind of change, and each can be built from data a validated system already produces. They are not a scorecard to adopt wholesale. Most programs should pick three.
1. Decision completion rate
Definition. Of the decisions of a named type that the system was built to support, the share that were made inside the system, on the system’s own data, within the defined service window, without a parallel approval recorded outside it.
How to specify it. Name the decision type first and write it down before go-live. Not “quality decisions” but, for example, “disposition of a released batch”, “triage classification of a new deviation”, or “acceptance of an out-of-trend stability result”. The denominator is every decision of that type in the period, counted from the process record rather than from the new system, so decisions made elsewhere still appear in the denominator. The numerator is the subset that meets all three conditions: recorded in the system, supported by data drawn from the system rather than pasted in from another source, and closed within the window the process defines.
What it detects. Whether the decision path actually moved. This is the single measure that most directly answers the question the standard four cannot answer. It also exposes the common pattern where a system becomes a place decisions are recorded after being made somewhere else.
Reading it. A rate below 60 percent at 90 days usually means the process design and the system design disagree, not that users are resistant. Look at the excluded cases before you look at the people.
2. Time to first value for a new user
Definition. The elapsed time from account activation to the user’s first independently completed real task of the type the system exists to support. Report the median and the interquartile range, never the mean.
How to specify it. Two definitions have to be fixed in advance and never changed mid-program. First, when the clock starts: account activation, meaning the moment credentials become usable, not the day training was assigned and not the day the user first logged in. Second, what counts as a first task: a real transaction of a defined type, on real data, completed without an assist recorded in the support log. A training sandbox exercise does not count. Neither does a task where a superuser sat alongside the user, which is why the support log matters.
What it detects. How much of the system’s benefit is available to someone who was not part of the project. Programs consistently over-estimate this, because the people designing the rollout have been using the system for months. A long tail here is the clearest early warning that the system will end up operated by a small group of specialists rather than by the process owners.
Reading it. Report by cohort, grouped by the month the user was activated. If the median for the cohort activated in month six is no better than the cohort activated in month one, the support model is not learning.
3. Task effort score
Definition. A single rating question asked immediately after a specific real task, on a seven-point scale, capturing how difficult the user found that task.
How to specify it. The instrument matters less than the timing. Sauro and Dumas compared three single-question post-task measures with 26 participants performing five tasks across two applications, and found that a simple Likert-type item and a subjective mental effort question both distinguished between the applications and were more sensitive at small sample sizes than a magnitude estimation approach.10 The seven-point single-item version has since been used widely as a post-task measure.11 If you need to know which kind of effort is high rather than just how high, the NASA Task Load Index provides six workload dimensions developed from a multi-year program covering 16 experiments.12
The related insight from customer research is that effort predicts behavior better than satisfaction does. Dixon, Freeman, and Toman, working from a study of 75,000 customer interactions, found that the effort a person had to expend was a better predictor of loyalty than satisfaction or willingness to recommend.13 The internal analogue is direct: a user who finds the task hard will route around it, whatever they said in the satisfaction survey.
What it detects. The friction that produces workarounds, before the workarounds show up in the transaction data. Effort scores are a leading indicator for the two measures below.
Reading it. Ask after the task, not at the end of the month, and never at the end of a training session. Report the response rate alongside the score, and report by role and site, because a good average with a 12 percent response rate concentrated in the superuser group tells you nothing.
4. Rework rate
Definition. The share of records of a given type that were edited, corrected, reopened, rejected, or re-approved after first submission, within a defined window after submission.
How to specify it. Fix the window, commonly 30 days, and fix which events count. Corrections before submission are a different thing and should be counted separately if at all. Rejections that send a record back a stage count. Post-approval edits count and are the most informative subset. Exclude changes driven by an upstream master data correction, and record that exclusion rule in the metric definition so nobody has to reconstruct it later.
What it detects. Whether the system is producing right-first-time work or moving the error later in the process. A new system frequently reduces the number of visible errors at entry while increasing the number of corrections after approval, which is worse, not better, in a regulated setting.
Reading it. Rework almost always rises in the first 60 days. That is expected and should be stated in the charter so that nobody treats the first data point as a failure.
5. Workaround rate
Definition. The share of transactions of a given type completed by a route other than the designed one.
How to specify it. List the alternative routes explicitly before go-live: manual entry where an interface exists, bulk upload where transaction entry was designed, use of an override or exception reason code, entry by a delegate rather than the performer, or completion outside the intended sequence. Count each separately, then report the total as a share of transactions.
The health informatics literature is the best available guide to what this looks like in practice, because the observation work has been done there. Koppel, Wetterneck, Telles, and Karsh studied barcode medication administration and identified 15 distinct types of workaround, grouped as omitted steps, steps performed out of sequence, and unauthorized steps, together with 31 probable causes spanning technology, organization, task, patient, and environment. Their override log covered 307,698 administrations, with overrides recorded for 10.3 percent of medications charted.15 Ash, Berg, and Coiera made the broader argument several years earlier: clinical information systems can generate new categories of error precisely where the software’s linear design meets the non-linear reality of the work.16
What it detects. The gap between the process as designed and the process as performed. A high workaround rate is design feedback, not a discipline problem, and treating it as the latter is the fastest way to make the measure disappear.
6. Residual work in the system it was meant to replace
Definition. The share of transactions of the target type still occurring in the predecessor system, plus a separate count of transactions occurring in spreadsheets.
How to specify it. Two components, reported separately. The first is straightforward: the predecessor system’s own transaction log, filtered to the transaction types in scope, expressed as a share of the combined total across old and new. The second, the shadow spreadsheet component, needs proxies because there is no log. Use export volume from the new system, counting distinct users who export to spreadsheet formats and the frequency with which they do it; file creation and modification activity in the shared locations the team uses; and, where available, template file names that persist from the old process.
Why shadow spreadsheets deserve their own signal. They are the clearest evidence that a system did not take over the analytical work, and they carry real risk of their own. Panko’s pooled review of six inspection studies covering 85 operational spreadsheets found errors in 94 percent of them, with individual studies reporting between 86 and 100 percent. Cell error rates observed in development studies fall in the same 1 to 5 percent range as base error rates for other non-trivial cognitive tasks, which is why large spreadsheets almost always carry at least one incorrect result.14 A rising export count after go-live is not a neutral observation. It means the number that informs a decision is being produced somewhere with no audit trail.
Time to first value, task effort score
Move within weeks. Tell you whether the system is usable by people who were not on the project team, before the transaction data shows anything.
Rework rate, workaround rate
Move over one to three months. Tell you whether the designed process and the performed process have converged.
Decision completion rate
Moves over three to six months. The only one that answers whether the decision path itself changed. Do not judge it early.
Residual legacy work and shadow spreadsheets
Moves last and often not at all without deliberate action. The measure most likely to contradict a green status report.
Where the Data Comes From
The objection to behavioral measures is always the same: we do not have the data. In regulated life sciences organizations this is usually untrue. You have more event-level data than almost any other industry, because keeping it is a requirement rather than a choice.
The audit trail is the primary source
Validated GxP systems keep an audit trail recording who did what to which record and when. That record exists for compliance reasons, and quality organizations review it for compliance reasons. The same data supports four of the six measures directly.
Decision completion rate comes from workflow state transitions with their timestamps and actor identifiers. Rework rate comes from post-submission modification and state-reversal events. Workaround rate comes from the events attached to alternative routes: bulk load records, override reason codes, delegated actions. Time to first value comes from joining the provisioning record to the first qualifying transaction by the same user.
Two cautions apply. First, using the audit trail for program measurement is a new purpose for existing data, and the extraction and any reporting built on it need to be handled under your normal controls rather than as a side activity. Second, an audit trail records what the system saw. A decision made in a meeting and entered afterward looks identical to a decision made in the system, which is why the decision completion rate definition includes a timestamp condition and why a periodic sample review is part of the method rather than an optional extra.
Process mining turns the log into a process picture
Event logs become substantially more useful when read as process data rather than as record data. Process mining, set out as a discipline in the Process Mining Manifesto produced by the IEEE Task Force on Process Mining in 2011, provides two capabilities that matter here: discovery, which reconstructs the process actually followed from the log, and conformance checking, which compares the observed behavior against the intended model and quantifies deviation.17
Conformance checking is close to a direct implementation of the workaround rate. Instead of listing alternative routes by hand, you specify the intended path and the analysis reports every case that departed from it and where. Where a process mining capability already exists in the organization, the measurement work is configuration rather than construction.
Survey data has to be collected, but not much of it
Task effort score is the only measure of the six that requires asking a person. Keep the burden proportionate. One question, presented in the system at the completion of a defined task type, sampled rather than universal, with the sampling rate set so that each site and role produces enough responses per month to be readable. Twenty responses per site per month is more useful than 400 responses concentrated among people who volunteer opinions.
The comparison problem, and the method that fixes it
Every measure above is meaningless without something to compare it to, and the comparison most programs make is the wrong one. A single pre-period average against a single post-period average will attribute to the program any trend that was already running, and any seasonal effect that happens to fall on the wrong side of go-live.
The established method for this situation is an interrupted time series read. Wagner, Soumerai, Zhang, and Ross-Degnan set out the approach for medication use research, using segmented regression to separate the level change at the point of intervention from the change in slope afterward, and describing it as the strongest quasi-experimental design available for evaluating the longitudinal effect of an intervention when randomization is not possible.18 The practical requirement it imposes is straightforward and it is the reason the baseline section below matters: you need enough pre-period points to establish the existing trend. Monthly points for eight to twelve months before go-live are usually sufficient, and a single pre-period point is never sufficient.
What this means in practice. If you can extract twelve months of history for a measure from the predecessor system, you can establish a baseline trend retrospectively, even if nobody thought about measurement before the program started. Do this first, before spending anything on new instrumentation. In most organizations two of the six measures turn out to be reconstructable from data already retained.
How Each Measure Fails When It Gets Gamed
Any measure attached to an outcome someone is accountable for will be managed. This is not cynicism, it is the oldest finding in performance measurement. Donald Campbell put it in 1979: the more a quantitative social indicator is used for social decision-making, the more subject it becomes to corruption pressures and the more likely it is to distort the process it was meant to monitor.6 Marilyn Strathern’s 1997 formulation, describing the growth of audit in British universities, is the version most often quoted: when a measure becomes a target, it ceases to be a good measure.5
The correct response is not to abandon measurement. It is to name the failure mode for each measure in advance, in the same document that defines the measure, and to specify the detection check that goes with it. A measure whose gaming pattern is written down is much harder to game.
| Measure | How it gets gamed | Detection check |
|---|---|---|
| Decision completion rate | The decision type is narrowed until only decisions already made in the system qualify. Records are created after the fact so the log shows a decision the system did not support. | Version-control the decision type definition and require change approval. Report the gap between the decision date field and the record creation timestamp; a growing gap is the tell. |
| Time to first value | A trivial task is designated as the qualifying first task. Accounts are provisioned late so the clock starts after the user is already trained and shadowed. | Fix the qualifying task list before go-live and treat changes as a definition amendment. Report the interval between training completion and provisioning as a companion figure. |
| Task effort score | The question is asked of superusers, or at the end of a training session while the trainer is present. Low-scoring responses are excluded as unrepresentative. | Publish response rate, role mix, and site mix with every score. Set a minimum coverage threshold below which the score is reported as not available rather than as a number. |
| Rework rate | Corrections are moved into a draft state that the metric does not count, or reclassified as data maintenance. The window is shortened until most rework falls outside it. | Track total edit events per record alongside the rate. If edits per record rise while the rework rate falls, the classification changed, not the behavior. |
| Workaround rate | The generic reason code is removed so alternative routes become unclassifiable. Blanket exceptions are granted so a route stops being an exception. | Count granted exceptions as a separate reported figure. Review any reason code whose volume drops sharply without a process change to explain it. |
| Residual legacy work | The predecessor system is switched off before the work moves, which pushes the work into spreadsheets where it cannot be counted. | Never report legacy residual without the spreadsheet export signal beside it. A legacy figure falling to zero while exports rise is a transfer, not a migration. |
The structural protection
Detection checks help, but the strongest protection is structural: do not attach individual incentives to any of these six measures. Use them to steer the program, not to evaluate the people using the system. Campbell’s argument was specifically about indicators used for decision-making about people and institutions, and the pressure scales with the consequence attached.6
The second protection is to keep the definitions owned by someone who does not report to the program. The measure definition document should have a named owner in the process organization or quality function, and changes to a definition mid-program should require the same approval as a change to a controlled procedure. Most metric corruption in practice is not falsification. It is a series of small, individually reasonable definition changes, each one made to fix a legitimate problem, whose combined effect is that the number at month nine is not measuring what the number at month one measured.
A Measurement Plan for One System Rollout
What follows is a plan for a single system going live in a defined population, not a portfolio dashboard. Portfolio-level reconciliation of executive KPIs is a separate discipline with different problems.
Before go-live: the baseline nobody captures
The baseline is the part that gets dropped, and it is worth being honest about why. Three reasons, in order of how often they are the real one.
The first is funding. Programs are funded to build and deploy. The measurement work has no deliverable of its own and it competes for the same analyst time as user acceptance testing in the weeks when both are needed. When something has to give, it is the thing without a milestone attached.
The second is timing. The window when a baseline must be captured is the eight to twelve weeks before go-live, which is the busiest period of the entire program. Asking a process owner to instrument the current process while they are validating the replacement is asking for the least available time in the schedule.
The third reason is the one people do not say aloud. A baseline is evidence. If the current process turns out to be performing better than the business case assumed, the business case is weakened. If it is performing far worse, someone has to explain why nobody noticed. Either way, a good baseline creates an obligation, and there is a quiet organizational preference for not creating one.
The consequence of skipping it. Without a baseline you cannot distinguish a system that made things worse from a system that inherited a process that was already deteriorating. Every post-go-live number becomes arguable, and arguable numbers get resolved by whoever holds more authority in the room rather than by evidence. This is the single highest-return item in the plan and it is almost always the first thing cut.
What to capture, minimally. Eight to twelve monthly points for whichever of the six measures can be reconstructed from the predecessor system, so that the pre-period trend is visible rather than a single average. Then, in the final four to six weeks before go-live, one round of task effort scores against the current process, a manual count of the decision type over four weeks, and a documented inventory of the spreadsheets currently in use for the in-scope work, including who maintains each one. That inventory takes about two days and is the most useful artifact in the whole plan, because it becomes the denominator for the shadow spreadsheet signal later.
The checkpoint schedule
Day 30: leading indicators only
Report time to first value for the first cohort, task effort scores, and workaround rate. Report rework rate as an observation with an explicit note that an increase is expected. Do not report decision completion rate at all, and say in the pack that you are not reporting it and why. Reporting an outcome measure before it can move trains the audience to ignore it later.
Day 90: the first real read
All six measures, with the baseline trend shown on the same chart rather than as a separate figure. This is the checkpoint where decision completion rate becomes meaningful and where residual legacy work becomes actionable. It is also the point to review whether any measure definition needs amendment, done once, formally, and documented, rather than continuously.
Day 180: the decision point
All six against baseline, read as a trend rather than as a level. The question at 180 days is not whether the numbers are good. It is whether they are moving in the intended direction at a rate that reaches the target within the planned horizon. A measure that is worse than baseline but improving steadily is a different situation from one that improved for 60 days and then flattened.
Beyond 180 days: hand over or stop
Either the measures move into the standing operational reporting owned by the process, or they stop. A program-owned metric that outlives the program becomes an orphan report that nobody reads and nobody can cancel. Decide this explicitly at the 180-day review.
Who reports what
The separation of reporting responsibility matters more than the reporting format, and it is a small structural decision that prevents a large problem.
| Role | Reports | Does not report |
|---|---|---|
| Program manager | Delivery status, scope, schedule, budget, risks, and dependencies | Any of the six behavioral measures. The person accountable for delivery should not also be the source for the numbers that judge it. |
| Process owner | Decision completion rate, rework rate, workaround rate, residual legacy work | Delivery status. The process owner is reporting on their own process, which is the point. |
| Data or analytics owner | Extraction method, data quality caveats, and any change to how a measure was calculated | Interpretation. Separating calculation from interpretation keeps definition disputes away from outcome disputes. |
| Quality or compliance | Validation of the measure definitions and approval of definition changes | Routine reporting. Involvement is at definition and amendment, not monthly. |
| Steering sponsor | Decisions taken in response to the measures | The measures themselves. A sponsor who presents the numbers cannot also be the person the numbers are presented to. |
In a smaller organization one person may hold two of these roles. That is workable. The one combination to avoid in every case is program manager and process owner, because it removes the only independent check in the structure.
Reporting a Number That Got Worse
This is where behavioral measurement usually fails, and it fails for a reason that has nothing to do with the measures. The first time a number goes down, the program has to survive the meeting. If it does not, the next report contains only measures that cannot go down, and the organization is back where it started within two quarters.
Agree the expected shape at charter stage
Some of these measures are supposed to get worse before they get better, and saying so in advance costs nothing while saying so afterward sounds like an excuse.
Task effort scores usually worsen in the first 30 to 60 days, because a competent user of the old process is a novice in the new one. Rework rate usually rises for a similar period as users learn what a complete record requires. Workaround rate often spikes at go-live and then falls sharply as gaps get closed, and the shape of that fall is more informative than the peak. Residual legacy work typically does not move at all until someone makes an explicit decision to stop accepting work through the old route.
Write these expectations into the charter, with rough magnitudes and durations, and have the sponsor sign them. A number moving as predicted is a sign the measurement is working. That framing is only available if the prediction was on the record first.
Set thresholds and attach an action to each
Define, before go-live, three bands for each measure and what each triggers. Within expected range: no action, note it. Outside expected range but improving: named investigation with a return date, no change to plan. Outside expected range and not improving after two consecutive periods: a defined intervention, which might be a design change, a process change, additional support, or a scope reduction.
The purpose of the bands is to make the response to a bad number automatic rather than negotiated. When there is no pre-agreed response, a bad number becomes a debate about whether it is really bad, and that debate is won by whoever has the most to lose.
Use a fixed three-line format
Every measure outside its expected range gets the same three lines, in the same order, every time.
- What moved. The measure, the direction, the magnitude, and the comparison basis. “Rework rate on batch records is 18 percent at day 90, against a pre-go-live baseline of 11 percent and an expected day-90 range of 12 to 16 percent.”
- What we think is causing it. The current best explanation, with its evidence and its confidence. “Two thirds of the corrections are on three fields that the previous form pre-populated and the new one does not. Confidence is moderate; based on a sample of 40 corrected records, not on the full population.”
- What we are changing and when we will know. One action, one owner, one date by which the measure should show a response. “Default values for those three fields are in the next configuration release, planned for day 110. We expect the rate to return within range by day 150 and will report at day 120 whether it is trending.”
Three lines, no narrative, no adjectives. The format is deliberately dull, and the dullness is the point. It moves the conversation from whether the news is bad to whether the diagnosis is right and the action is sufficient, which is the only conversation worth having.
The rule that protects the whole system: a measure that gets worse and is reported accurately, with a diagnosis and an action, should never on its own trigger cancellation. Cancellation decisions belong to a separate set of criteria agreed in advance. If reporting a bad number can end a program, nobody will report one, and the measurement system becomes decorative within a quarter.
Two supporting practices
First, report the same measures in the same order every period, including the ones that are fine. A pack whose contents change with the news teaches the audience to read the table of contents for the story, and it makes any omission look like concealment.
Second, borrow the framing from software delivery measurement. The DORA measures pair speed with stability precisely so that neither can be improved by damaging the other, and the practice guidance is to use them for team improvement rather than for comparison between teams.19 The same balance applies here. Decision completion rate paired with rework rate is a matched set: pushing decisions into the system while breaking quality shows up immediately in the second number. Report them together, always, and never report one without the other.
A note on what the academic models contribute
The information systems literature has spent forty years on the question of what predicts and constitutes system success, and two strands are directly useful when you are defending this approach to a skeptical sponsor.
Davis established in 1989 that perceived usefulness and perceived ease of use are the two constructs that most strongly relate to whether people use a system, validating scales across two studies with a total of 152 users and four applications.7 Venkatesh, Morris, Davis, and Davis later combined eight competing acceptance models into a unified account built on performance expectancy, effort expectancy, social influence, and facilitating conditions, reporting that the combined model accounted for around 70 percent of variance in intention to use, against 17 to 41 percent for the individual models it replaced.8 Effort is not a soft measure. It is one of the two or four constructs that decades of empirical work keep returning to.
DeLone and McLean’s updated model of information systems success makes the complementary point: system success is multidimensional, spanning system quality, information quality, service quality, use, user satisfaction, and net benefits, and measuring any one of them alone gives an incomplete and often misleading picture.9 Which is a formal way of saying that a login count on its own was never going to be enough.
Conclusion
The reason activity metrics persist is not that anyone believes them. It is that they are cheap, unambiguous, available immediately, and incapable of embarrassing the person reporting them. Replacing them means accepting a harder conversation, and that trade is only worth making if the organization intends to act on what it learns. A program that will not change course regardless of the evidence is better off with the cheap numbers, and it is worth being clear about which kind of program you are running before investing in measurement.
If you do intend to act, the requirements are modest. Pick three measures, not six. Define each one with a numerator, a denominator, a window, and a source, and write down how it will be gamed. Reconstruct eight to twelve months of baseline from the data you already retain, because that is what makes every later number readable. Separate the person who reports delivery from the person who reports outcomes. Agree in advance which measures are expected to get worse and by how much, and agree that reporting a bad number honestly will not end the program. None of that requires new technology, and most of it requires about two weeks of work by people who are already in the room.
Sakara Digital works with pharma and biotech organizations on measurement that survives contact with a steering committee, including baselining before a system rollout and defining metrics that can move in both directions. If you are planning a rollout and want an independent view on what to measure before go-live rather than after, we are happy to have that conversation.
For Further Reading
For Further Reading
- Measuring AI Adoption, Not ROI: The Metrics That Predict Scale
- Change Management for Digital Transformation in Pharma
- Quality Metrics That Actually Drive Improvement in Pharma
- Building a Digital Transformation Office in Life Sciences
- The Manufacturing Data Quality Scorecard: KPIs Beyond Regulatory Submissions
References & Sources
- Hughes, Mark. “Do 70 Per Cent of All Organizational Change Initiatives Really Fail?” Journal of Change Management, vol. 11, no. 4, 2011, pp. 451-464. https://research.brighton.ac.uk/en/publications/do-70-per-cent-of-all-organizational-change-initiatives-really-fa/
- Kotter, John P. “Leading Change: Why Transformation Efforts Fail.” Harvard Business Review, May-June 1995. https://hbr.org/1995/05/leading-change-why-transformation-efforts-fail-2
- McKinsey & Company, People & Organizational Performance and Transformation Practices. “Losing from Day One: Why Even Successful Transformations Fall Short.” December 2021. Losing from Day One (PDF)
- Eveleens, J. Laurenz, and Chris Verhoef. “The Rise and Fall of the Chaos Report Figures.” IEEE Software, vol. 27, no. 1, 2010, pp. 30-36. https://www.cs.vu.nl/~x/the_rise_and_fall_of_the_chaos_report_figures.pdf
- Strathern, Marilyn. “‘Improving ratings’: audit in the British University system.” European Review, vol. 5, no. 3, 1997, pp. 305-321. Cambridge Core record
- Campbell, Donald T. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning, vol. 2, no. 1, 1979, pp. 67-90. https://ideas.repec.org/a/eee/epplan/v2y1979i1p67-90.html
- Davis, Fred D. “Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology.” MIS Quarterly, vol. 13, no. 3, 1989, pp. 319-340. https://www.jstor.org/stable/249008
- Venkatesh, Viswanath, Michael G. Morris, Gordon B. Davis, and Fred D. Davis. “User Acceptance of Information Technology: Toward a Unified View.” MIS Quarterly, vol. 27, no. 3, 2003, pp. 425-478. https://aisel.aisnet.org/misq/vol27/iss3/5/
- DeLone, William H., and Ephraim R. McLean. “The DeLone and McLean Model of Information Systems Success: A Ten-Year Update.” Journal of Management Information Systems, vol. 19, no. 4, 2003, pp. 9-30. https://www.tandfonline.com/doi/abs/10.1080/07421222.2003.11045748
- Sauro, Jeff, and Joseph S. Dumas. “Comparison of Three One-Question, Post-Task Usability Questionnaires.” Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI 2009), pp. 1599-1608. https://measuringu.com/papers/Sauro_Dumas_CHI2009.pdf
- Sauro, Jeff. “The Evolution of the Single Ease Question (SEQ).” MeasuringU. https://measuringu.com/evolution-of-seq/
- Hart, Sandra G., and Lowell E. Staveland. “Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research.” NASA Ames Research Center, 1988. https://archive.org/details/nasa_techdoc_20000004342
- Dixon, Matthew, Karen Freeman, and Nicholas Toman. “Stop Trying to Delight Your Customers.” Harvard Business Review, July-August 2010. https://hbr.org/2010/07/stop-trying-to-delight-your-customers
- Panko, Raymond R. “What We Don’t Know About Spreadsheet Errors Today: The Facts, Why We Don’t Believe Them, and What We Need to Do.” Proceedings of the 16th EuSpRIG Conference, 2015. https://arxiv.org/abs/1602.02601
- Koppel, Ross, Tosha Wetterneck, Joel Leon Telles, and Ben-Tzion Karsh. “Workarounds to Barcode Medication Administration Systems: Their Occurrences, Causes, and Threats to Patient Safety.” Journal of the American Medical Informatics Association, vol. 15, no. 4, 2008, pp. 408-423. https://pmc.ncbi.nlm.nih.gov/articles/PMC2442264/
- Ash, Joan S., Marc Berg, and Enrico Coiera. “Some Unintended Consequences of Information Technology in Health Care: The Nature of Patient Care Information System-related Errors.” Journal of the American Medical Informatics Association, vol. 11, no. 2, 2004, pp. 104-112. https://pmc.ncbi.nlm.nih.gov/articles/PMC353015/
- van der Aalst, Wil, et al. (IEEE Task Force on Process Mining). “Process Mining Manifesto.” Business Process Management Workshops (BPM 2011), Lecture Notes in Business Information Processing, vol. 99, 2011. https://research.tue.nl/en/publications/process-mining-manifesto/
- Wagner, A. K., S. B. Soumerai, F. Zhang, and D. Ross-Degnan. “Segmented Regression Analysis of Interrupted Time Series Studies in Medication Use Research.” Journal of Clinical Pharmacy and Therapeutics, vol. 27, no. 4, 2002, pp. 299-309. https://pubmed.ncbi.nlm.nih.gov/12174032/
- Google Cloud. “Use Four Keys Metrics Like Change Failure Rate to Measure Your DevOps Performance.” Google Cloud Blog. https://cloud.google.com/blog/products/devops-sre/using-the-four-keys-to-measure-your-devops-performance








Your perspective matters—join the conversation.