Why Automating a Broken Process Makes It Fail Faster

There is a pattern we see often in life sciences AI projects. A team picks a process that everyone agrees is slow. A vendor demonstrates a tool that handles a clean sample of that work well. The pilot starts, and within weeks the project team is spending most of its time on exceptions: records that do not fit the template, reviewers who disagree with the output, cases where nobody can say what the right answer should have been. The tool is blamed. In most cases the tool is doing exactly what it was built to do. The process underneath it was never stable enough to automate.

This is the core reason to fix processes before AI automation. It is not a philosophical point about technology. It is a practical point about what automation does to a process that has problems in it.

Automation Repeats Whatever It Is Given

A manual process with problems has a built-in brake. People notice when something looks wrong. An experienced reviewer knows that a certain form is often filled in incorrectly on night shift and checks it more closely. A quality lead remembers that two sites classify the same event differently and adjusts. These fixes are informal, undocumented, and unevenly applied, but they catch a share of the errors.

When you automate that process, the informal fixes usually disappear. The AI system is trained or configured on the documented process, or on historical records that contain the same inconsistencies. It then applies that pattern at scale, with no reviewer pausing to say “that looks odd.” Three things happen at once:

  • Errors move faster. A misclassified deviation that once waited in a queue for a day is now routed in seconds to the wrong owner.
  • Errors become more uniform. A manual process makes scattered mistakes. An automated process makes the same mistake every time the same conditions appear.
  • Errors become harder to see. Output from a system looks authoritative. Reviewers who once did the work start checking it less carefully, particularly when volume goes up.

That combination is what “fail faster” means in practice. The process does not fail in a new way. It fails in the old way, more often, with less human attention on it.

The Evidence From AI Programs

Industry research on AI value points in the same direction. In an October 2024 survey of 1,000 executives across 59 countries, Boston Consulting Group found that 74% of companies had yet to show tangible value from their use of AI. BCG reported that the leaders in its sample follow what it calls a 10-20-70 rule: 10% of resources into algorithms, 20% into technology and data, and 70% into people and processes.1

Gartner made a related prediction in July 2024: that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs, or unclear business value.2 Several of those reasons trace back to process problems. Poor data quality is often the record of an inconsistent process. Unclear business value is often the result of automating a process nobody had measured, so there was no baseline to show improvement against.

74%of companies had yet to show tangible value from AI in BCG’s October 2024 survey of 1,000 executives1
70%of resources go to people and processes among the AI leaders BCG identified (the 10-20-70 rule)1
30%of generative AI projects predicted by Gartner to be abandoned after proof of concept by end of 20252

An Old Lesson With Higher Stakes

None of this is new. In 1990, Michael Hammer published “Reengineering Work: Don’t Automate, Obliterate” in the Harvard Business Review. His argument was that companies were using computers to speed up processes designed around old constraints, when they should have been asking whether those processes needed to exist in their current form at all.3 The technology then was mainframes and early networks. The lesson carries over directly to AI.

What is different in pharma and biotech is the consequence of getting it wrong. In most industries, automating a broken process wastes money and annoys customers. In GxP work, it can produce records that do not support release decisions, investigations that do not find root causes, and review steps that no longer do what the regulations expect them to do. Those outcomes show up during inspections.

What a Broken Process Looks Like in GxP Work

“Broken” does not mean noncompliant. Many of the processes that most need fixing before automation are fully compliant on paper. They have approved SOPs, trained staff, and records that pass audit. They are broken in a narrower, operational sense: they produce different results for the same input, they depend on rework to reach an acceptable outcome, or they only work because a few experienced people fill the gaps.

Five Signs a Process Is Not Ready for AI

The table below lists the signs we look for, what they look like day to day, and what an AI system tends to do with them.

SignWhat You See Day to DayWhat Automation Does With It
Undefined termsStaff disagree on what counts as “major,” “complete,” or “reviewed.” Answers depend on who you ask.The system learns an average of inconsistent judgments, or applies one reading that half the organization disagrees with.
Heavy reworkRecords routinely go back for correction. A document takes three or four review rounds before approval.Rework loops get faster but not fewer. The system may add a new loop when reviewers reject its output.
Loose handoffsWork waits between teams. Nobody is sure who owns a step. Items are chased by email.Automated routing sends work to the wrong queue at machine speed, or stalls at the same unclear handoff.
Hidden workaroundsSpreadsheets, side lists, and “the way we do it here” that differ from the SOP.The system follows the SOP and breaks the workaround that was keeping the process running.
High variationThe same type of task takes one day or three weeks depending on site, shift, or person.No stable pattern to learn or test against, so accuracy claims from a pilot do not hold in production.

If three or more of these signs apply to the process you want to automate, the process work will likely deliver more value in the first six months than the AI system will.

Why Regulators Care About These Particular Processes

The processes most often chosen for AI in quality operations are review and investigation processes. Those are also the processes FDA cites most often. FDA’s Report on the State of Pharmaceutical Quality for FY2025, published in June 2026, looked at nearly 10,000 original applications and supplements received in FY2024 and FY2025 that required a facility assessment. As of May 2026, 28% had received a Complete Response Letter, and 43% of those (12% of all the submissions) were due to facility withholds.4

The report named three common inspection citations that resulted in facility withholds. The first was failure to thoroughly review discrepancies or batch failures within production record review, which FDA summarized as inadequate investigations. The second was quality control unit responsibilities or procedures not written or followed, with examples that include inadequate oversight of deviations.4 Batch record review, deviation investigation, and quality unit oversight are exactly the areas where companies are now proposing AI support. If the underlying process has gaps, automating it puts a new layer on top of a known inspection risk.

~10,000original applications and supplements requiring a facility assessment, received FY2024-2025, analyzed by FDA4
28%of those submissions had received a Complete Response Letter as of May 20264
43%of those Complete Response Letters were due to facility withholds (12% of all submissions)4

The Human Error Trap

One sign of a broken process deserves its own mention because it hides the others. When investigations regularly conclude “human error,” the process looks fine and the people look like the problem. That conclusion closes the record, but it rarely fixes anything.

EU GMP Chapter 1, which covers the pharmaceutical quality system, addresses this directly. It states: “Where human error is suspected or identified as the cause, this should be justified having taken care to ensure that process, procedural or system-based errors or problems have not been overlooked, if present.”5 In other words, the regulator expects the investigation to rule out process causes before it settles on the person. (Chapter 1 has carried that expectation since it came into operation in January 2013.)

Susan Schniepp made a similar point in Pharmaceutical Technology, arguing that “human error” is overused as a root cause and that the more useful question is what caused the employee to make the error. Her recommendations focus on clearer documentation, better training, and monitoring that looks for system problems rather than individual blame.6

This matters for AI in a specific way. If your deviation history is full of “human error” conclusions, that history is a poor training set and a poor benchmark. An AI system trained to predict root cause from it will learn to predict “human error,” because that is what the records say. It will be accurate against the history and wrong about the plant.

Watch for this: A high share of “human error” or “undetermined” conclusions in a process’s records is not a reason to automate the investigation step. It is a sign that the investigation process itself needs work first, and that the historical records should not be used as ground truth for any AI system.

Step One: Map the Process as People Run It Today

The first step to fix processes before AI automation is to find out what the process is. That sounds obvious. In practice, most automation projects start from the SOP, and the SOP describes the process as it was designed, approved, and last revised. It rarely describes the process as it runs today.

The SOP Is Where You Start, Not Where You Stop

Read the SOP first so you know the intended design. Then go and watch the work. Sit with the people who do it. Follow a few real records from start to finish. Ask where they wait, what they check that the procedure does not mention, and what they do when something does not fit.

Lean practitioners call this a current state map. The Lean Enterprise Institute describes value-stream mapping as diagramming every step in the material and information flows needed to deliver a product, starting with a current state map of how the flow works today and then a future state map of how it should work.7 The same method applies to a flow of documents, records, and decisions. The discipline is to map what happens, including the steps nobody is proud of.

1

Pick the Boundaries

Define where the process starts and ends. For a deviation, is the start the event on the floor, or the record being opened in the QMS? Is the end closure, or CAPA effectiveness check? Different boundaries give different maps.

2

Walk Real Records

Choose 10 to 20 recent records, including some that went smoothly and some that did not. Trace each one through every person, system, and queue it touched.

3

Capture Steps, Waits, and Decisions

Record each step, who performs it, what they need to start, what they produce, how long the work takes, and how long the item waits before the next step. Mark every decision point and note who makes it and on what basis.

4

Find the Workarounds

List every spreadsheet, email chain, side conversation, and informal check that keeps the process moving. These are often the most important parts of the map.

5

Compare to the SOP

Mark every place the real process differs from the written one. Each difference is either a procedure that needs updating or a practice that needs correcting.

Use the Data You Already Have

Most GxP processes already produce an event log. A QMS records when a deviation was opened, assigned, investigated, reviewed, and closed. A document management system records each review cycle and approval. An electronic batch record system records each entry and each review. That data can show you the real process at a scale no set of interviews can match.

This is the idea behind process mining. The Process Mining Manifesto, written by Wil van der Aalst and a large group of co-authors and published in 2012, describes the goal as discovering, monitoring, and improving real processes, not assumed ones, by extracting knowledge from event logs in existing information systems. It also describes conformance checking: comparing a process model against the event log to find where actual behavior departs from the model.8

You do not need a process mining platform to use this idea. An export of timestamps and status changes from your QMS, analyzed in a spreadsheet, will often show you how many records loop back, where they wait longest, and which paths through the process are common and which are rare. Pair that with a few days of observation and you have a map you can trust.

Where Workarounds Hide

Workarounds are easy to miss because the people using them no longer think of them as separate from the process. Some places to look:

  • Personal trackers. Many quality staff keep their own list of open items because the system view does not show what they need. Ask to see it.
  • Pre-review checks. Authors often send a draft to a trusted colleague before formal review, to avoid rejection. That is an unofficial review step.
  • Standing meetings. A weekly meeting where open deviations are discussed and assigned may be doing the triage work the SOP assigns to a single role.
  • Known exceptions. “We always hold those for the site head” or “that product line goes to a different reviewer” are rules that exist nowhere in writing.

Questions to ask in every walkthrough:

  • What do you need before you can start this step, and how often is it missing?
  • When do you send something back, and what is the most common reason?
  • What do you check that the procedure does not ask you to check?
  • If a new colleague did this step exactly as written, what would go wrong?
  • Where does work wait, and who is it waiting for?

Step Two: Measure Variation and Rework

A map tells you how the process flows. Measurement tells you how well. Before you change anything, you need numbers that show where the process is unstable, how much rework it carries, and what normal performance looks like. Those numbers do two jobs. They tell you what to fix in Step Three, and they give you the baseline you will need later to show that automation helped.

Borrow the Logic of Process Validation

Manufacturing teams already know how to think about variation. FDA’s 2011 guidance on process validation states that manufacturers should understand the sources of variation, detect the presence and degree of variation, understand the impact of variation on the process and ultimately on product attributes, and control the variation in a manner commensurate with the risk it represents to the process and product.9 The same guidance recommends using quantitative, statistical methods whenever appropriate and feasible.

That guidance is written for manufacturing processes, not for document review or deviation handling. But the logic transfers well. A quality review process also has inputs, steps, and outputs, and its output also varies. Before you automate it, you should be able to say where that variation comes from, how large it is, and how much it matters.

What to Measure

The following measures are practical to collect from existing systems and a short sampling exercise. None requires new technology.

MeasureHow to Collect ItWhat It Tells You
Cycle time distributionTimestamps from the QMS, DMS, or EBR system, by record type, site, and monthThe spread, not just the average. A wide spread points to variation you need to explain.
Wait time shareTime between steps compared to time spent in stepsWhere handoffs are slow. In many review processes, waiting is most of the elapsed time.
Rework rateCount of records returned for correction, and the number of loops per recordHow often the process fails the first time, and at which step.
First-pass acceptanceShare of records approved without any returnThe simplest single measure of process health.
Reviewer agreementTwo qualified people classify or review the same sample independentlyWhether the organization has a shared standard. Low agreement means the definitions need work.
Reopen and recurrence rateRecords reopened after closure; repeat events with the same causeWhether the process fixes problems or only closes records.

Averages hide the problem. An average deviation closure time of 28 days can mean that nearly every record closes in about four weeks, or that half close in a week and the rest take two months. Those are very different processes, and only the second one has a variation problem. Always look at the distribution and the outliers.

Measure Agreement Between Experts

Of all the measures above, reviewer agreement is the one most often skipped and the one that matters most for AI. The test is simple. Take 30 to 50 recent records. Ask two or three qualified people to classify or review them independently, without seeing each other’s answers or the original conclusion. Then compare.

If your experts agree most of the time, you have a shared standard, and an AI system has something stable to learn and be tested against. If they disagree often, you do not have a standard yet. You have several personal standards. No AI system can be more consistent than the definition it is built on, and you cannot validate a system against a “right answer” your own experts do not agree on.

Even carefully curated data carries this problem. A 2021 study by Curtis Northcutt, Anish Athalye, and Jonas Mueller found an average of at least 3.3% label errors across ten widely used machine learning benchmark test sets, with at least 6% in the ImageNet validation set. The authors showed that these errors were enough to change which model looked best.10 Those were benchmark datasets that researchers had checked with care. There is no reason to assume a company’s deviation classification history is cleaner, and good reason to test it before trusting it.

Set the Baseline You Will Need Later

Write down the numbers from this step and keep them. When the automation project is proposed, the business case will depend on them. When the system is validated, acceptance criteria should refer to them. When leadership asks a year later whether the AI investment paid off, these are the numbers you will compare against. Projects that skip this step often cannot answer that question, which is one reason so many AI programs struggle to show value.

Step Three: Fix Handoffs and Definitions

Steps One and Two show you where the process breaks. Step Three is where you fix it. In our experience, most of the fixes fall into four groups: definitions, decision rules, handoffs, and inputs. Very few require new software.

Definitions

Say What the Words Mean

Write down what “minor,” “major,” and “critical” mean for this process, with examples of each. Define “complete,” “reviewed,” and “approved” in terms of what must be true, not who signed.

Decision Rules

Turn Judgment Into Criteria

Where experienced staff make a call, ask them how they decide. Capture the criteria they use. Keep human judgment for the cases the criteria do not cover, and say so explicitly.

Handoffs

Give Every Step an Owner

Each step needs one named role, clear entry criteria (what must be ready before it starts), and clear exit criteria (what must be true before it moves on). Remove steps that exist only to check the previous step.

Inputs

Fix Problems at the Source

If a large share of rework comes from incomplete or incorrect inputs, fix the form, template, or upstream process that creates them. Downstream review should not be the main quality control for upstream errors.

Definitions Come First

Most disagreement in quality processes comes from terms that everyone uses and nobody has defined in operational terms. “Major deviation” may be defined in the SOP as a deviation with “potential impact on product quality,” but that phrase does not tell a reviewer whether a two-minute excursion on a hold-time parameter qualifies. Two reviewers reading the same definition will reach different answers.

The fix is to add criteria and worked examples. For each category, list the conditions that place an event in it, give three or four real examples from your own history, and give one or two borderline cases with the reasoning for where they fall. Then repeat the reviewer agreement test from Step Two. If agreement improves, the definition is working. If it does not, keep refining.

This work is what an AI system will depend on later. A classification model, a set of rules, or a prompt given to a language model all need a clear definition of each category. If you write that definition now, for people, you have written most of the specification for the system.

Handoffs Are Where Time Goes

In most review and investigation processes, the time spent doing the work is small compared to the time spent waiting between steps. A deviation may need four hours of investigation effort and still take five weeks to close, because it waits for assignment, waits for a subject matter expert, waits for quality review, and waits for approval.

Fixing handoffs usually means a few simple changes. Give each queue a named owner who is accountable for how long items wait in it. Define what “ready for the next step” means so that items are not passed forward incomplete and then sent back. Where two steps exist only because one checks the other, ask whether the first step can be done right so the second is not needed. Hammer’s argument from 1990 applies here: some steps exist because of constraints that no longer apply, and the right move is to remove them, not to speed them up.3

Change the Procedure Through the Proper Route

In a GxP setting, fixing the process means changing the procedure, and that goes through change control. Under 21 CFR 211.100(a), written procedures for production and process control, including any changes, must be drafted, reviewed, and approved by the appropriate organizational units and reviewed and approved by the quality control unit. Under 211.100(b), those procedures must be followed, and any deviation from them must be recorded and justified.11 Similar expectations apply to quality system procedures under EU GMP.

That requirement is a help, not a hurdle. When the workarounds found in Step One are either written into the procedure or removed, the gap between the documented process and the real process closes. That gap is itself an inspection risk, whether or not AI is ever added.

Run the Fixed Process Manually First

Once the procedure is updated and staff are trained, run the fixed process by hand for a period long enough to measure it again. For high-volume processes such as document review, a few weeks may be enough. For lower-volume processes, it may take a quarter. Then repeat the measurements from Step Two.

Two things usually happen. First, performance improves, sometimes a great deal, before any automation is added. Second, the remaining problems become clearer, because the noise from undefined terms and loose handoffs is gone. Both outcomes are useful. The first gives the business an early return. The second tells you precisely which parts of the process AI should target.

A good result, not a failed project: Sometimes the process fix delivers most of the improvement the AI project was supposed to deliver. When that happens, leaders sometimes see it as a reason the AI project “was not needed.” A better reading is that the company got the gain sooner and for less money, and can now decide whether automation adds enough on top to justify the investment.

Step Four: Automate, With the Process as the Specification

When the process is mapped, measured, fixed, and stable, automation becomes much simpler. The work from the first three steps turns directly into the inputs an AI project needs.

Output of Steps One to ThreeWhat It Becomes in the AI Project
Current and future state process mapThe scope: which steps the system performs, which stay with people, and where the handoff between them happens
Written definitions and decision rulesThe functional requirements, and the instructions or labels the system is built on
Worked examples and borderline casesTest cases for validation, including the hard cases where the system is most likely to fail
Reviewer agreement resultsA realistic accuracy target. The system should be measured against the level of agreement your experts achieve.
Baseline cycle time, rework, and first-pass acceptanceAcceptance criteria and the business case measures to track after go-live

Write Down the System’s Role Before You Choose a Tool

Before any tool is selected, the project team should be able to state in a few sentences what question the system answers, which steps it performs, what it produces, and how people will use that output. For example: “The system reads each new deviation record and proposes a classification of minor, major, or critical using the criteria in the procedure. A quality reviewer confirms or changes the classification before the record is assigned. The proposed classification is not used for any other decision.” A statement like that sets the scope, the risk, and the validation effort.

That statement is hard to write for a process that has not been mapped or whose decision rules exist only in people’s heads. It is straightforward to write after Steps One to Three, because every term in it has already been defined. It also tells you how much the system’s output matters. A system whose output a person always confirms before anything happens carries less risk than one whose output feeds a release decision directly, and the testing and oversight should reflect that difference.

Automate the Stable, Rule-Based Parts First

Not every step in a fixed process should be automated. The best early candidates share a few traits:

  • High volume. The step happens often enough that saving time on each instance adds up.
  • Clear criteria. The decision rules from Step Three cover most cases, and your experts agree on them.
  • Checkable output. A person can quickly confirm whether the system got it right.
  • Contained impact. An error is caught by a later step before it can affect product or patient.

Steps that depend on judgment the criteria do not capture, or where an error would reach a release decision without another check, should keep a person making the decision, with AI in a supporting role at most. The process map shows you where those boundaries are.

Keep Measuring After Go-Live

Automation changes the process, so the measurements from Step Two need to continue. Track the same cycle time, rework, and first-pass acceptance measures after go-live, and compare them to the baseline. Add measures for the system itself: how often reviewers override it, which categories it gets wrong, and whether its error rate changes over time.

Watch in particular for a drop in human attention. If reviewers stop overriding the system, that can mean it is working well, or it can mean they have stopped checking. Periodic blind sampling, where a reviewer checks a random set of records without seeing the system’s output first, is a simple way to tell the difference.

What good looks like at this stage: A written process that matches how work is done. Definitions your experts apply consistently. A measured baseline. A system scope that names which steps are automated and which stay human. Test cases drawn from your own records, including the hard ones. And a monitoring plan that uses the same measures you collected before you started.

Three Life Sciences Examples: Document Review, Deviations, and Batch Records

The four steps are general. How they apply depends on the process. The three examples below are the processes we are asked about most often when companies plan AI in quality and operations.

Document Review

The goal companies bring: Use AI to check SOPs, validation documents, and reports for completeness, consistency, and format before or during human review, so review cycles get shorter.

What usually breaks first: Document review processes often have no shared definition of what a reviewer is responsible for. One reviewer checks technical content. Another rewrites sentences for style. A third comments on formatting. Comments arrive at different times, conflict with each other, and trigger another round. Authors learn to expect four or five rounds and stop trying to get it right the first time. The rework is built into the process.

If you add an AI checker to that process, it becomes one more reviewer with its own standard. Its comments add to the pile. Reviewers argue with its suggestions. Cycle time may not improve at all.

Fix before you automate:

  • Define what each reviewer role owns. Technical reviewers own technical accuracy. Quality reviewers own compliance with procedure and regulation. Style and formatting belong to a template and a style guide, not to every reviewer.
  • Write the checklist each role reviews against. Under 21 CFR 211.100(a), procedures must be reviewed and approved by the appropriate organizational units and by the quality control unit,11 but the regulation does not tell you what each reviewer should look at. Your procedure should.
  • Measure first-pass acceptance and the number of review rounds per document type. Find the most common reasons for return.
  • Fix the templates so the most common return reasons cannot occur, for example by making required sections mandatory and pre-filling standard text.

What AI can then do well: Once each role has a written checklist, an AI system can run the checks that are rule-based (required sections present, references valid, terms used consistently, numbering correct) before the document reaches a human. Reviewers then spend their time on the technical and compliance questions only they can answer. The checklist is the specification. The record of past return reasons is the test set.

Deviations

The goal companies bring: Use AI to triage and classify new deviations, suggest likely root causes from past events, draft parts of the investigation report, and flag trends.

What usually breaks first: Deviation classification varies by site and by person, as described in Step Two. Root cause categories are too broad, so “human error” and “procedure not followed” absorb a large share of events. Investigations are closed on time because the metric is closure, not effectiveness. And the historical record, which any AI system would learn from, reflects all of this.

Regulators are watching this process closely. In an August 27, 2026 warning letter to Jabil Inc., a contract manufacturer of sterile injectables, FDA cited 21 CFR 211.192 and stated that the firm’s root cause finding of “undetermined” for a sterility test failure was inadequate. FDA wrote that “investigations must be thorough, well-documented, scientifically sound, and timely,” and that “Procedural updates and training alone do not address the systemic failures that allowed deficient investigations to persist undetected by quality unit (QU) oversight.”12 FDA asked the firm for a comprehensive, independent assessment of its overall system for investigating deviations, discrepancies, complaints, out-of-specification results, and failures, with an action plan covering investigation competencies, scope determination, root cause evaluation, CAPA effectiveness, and quality unit oversight.12

That list is, almost item for item, the process work this article recommends before automation. FDA is asking firms to fix the investigation process itself, and it is explicit that retraining people and updating procedures are not enough on their own.

Fix before you automate:

  • Rewrite the classification criteria with examples and borderline cases, then test reviewer agreement across sites until it is acceptable.
  • Break broad root cause categories into specific, actionable ones. Require that any “human error” conclusion show that process, procedural, and system causes were considered, as EU GMP Chapter 1 expects.5
  • Define what a complete investigation contains, so quality review checks against a standard rather than a reviewer’s preference.
  • Add recurrence and CAPA effectiveness to the measures, not only on-time closure.
  • Clean or relabel a sample of historical records under the new criteria before using any of them to train or test an AI system.

What AI can then do well: With consistent classification criteria and a relabeled sample, AI can suggest an initial classification for a person to confirm, find similar past events across sites that a single investigator would not know about, and flag records whose investigation is missing required elements. The investigator still owns the conclusion. AI makes the evidence easier to find and the gaps easier to see.

Batch Record Review

The goal companies bring: Move to review by exception, where reviewers look only at entries that fall outside defined limits or conditions, with AI or rules-based logic identifying those exceptions.

What usually breaks first: Review by exception only works if the organization has already decided which entries are critical, what the acceptable range is for each, and what counts as an exception. As described in a BioPharm International article by Johan Zebib, ISPE defines review by exception as “an approach in which manufacturing and quality data are screened to present or report only critical process exceptions as required by approvers for review and disposition of intermediates and products.”13 Every word in that definition assumes prior process work: someone has to have decided what is critical and what the approvers require.

The regulatory requirement is clear. Under 21 CFR 211.192, production and control records must be reviewed and approved by the quality control unit to determine compliance with established, approved written procedures before a batch is released or distributed, and any unexplained discrepancy must be thoroughly investigated.14 Review by exception does not remove that responsibility. It changes how the review is performed, which means the logic that selects exceptions has to be right.

If the master batch record is poorly designed, with ambiguous instructions, fields that are routinely left blank or corrected, or limits that do not match the validated process, exception logic will either flag too much (and reviewers go back to reading everything) or flag too little (and real discrepancies pass through). Good documentation practice errors that are corrected by hand on every batch are a sign that the record, not the operator, needs fixing.

Fix before you automate:

  • Measure which fields generate the most corrections, comments, and review queries. Those point to record design problems.
  • Redesign the master batch record so that instructions are unambiguous and each entry has a clear expected value or range.
  • Agree, with quality and manufacturing together, which entries are critical and what the exception criteria are for each. Tie them to the process knowledge from validation.
  • Define what the reviewer must do with each type of exception, and what evidence they must record.
  • Run the new exception criteria in parallel with full manual review for a defined number of batches, and compare what each finds.

What AI can then do well: With agreed critical entries and exception criteria, rules-based logic handles most exception detection. AI can add value at the edges: reading free-text comments to spot entries that look routine but describe a problem, finding patterns across batches that suggest drift before any single batch is out of limits, and helping reviewers assemble the evidence for a disposition. The exception criteria remain the core, and they come from the process work, not from the tool.

The Three Examples Side by Side

ProcessCommon BreakFix Before AIWhere AI Then Helps
Document reviewNo defined reviewer roles; many review rounds; comments conflictRole-based checklists, better templates, measured return reasonsRule-based pre-checks before human review
DeviationsInconsistent classification; “human error” and “undetermined” conclusions; closure measured over effectivenessClassification criteria with examples, specific root cause categories, recurrence measures, relabeled sampleSuggested classification, similar-event search, completeness checks
Batch record reviewNo agreed critical entries; record design drives repeated correctionsRecord redesign, agreed exception criteria, parallel run against full reviewFree-text review, cross-batch trend detection, evidence assembly

A Readiness Check Before You Automate

Leaders are often asked to approve an AI project for a process they do not work in day to day. The following check is designed for that situation. Each question has a yes or no answer, and each “no” points to work that should happen first.

#QuestionIf the Answer Is No
1Is there a current state map of the process as it runs today, checked against real records?Map the process (Step One) before scoping the system.
2Do we know the distribution of cycle time and where work waits?Pull timestamps from the existing system and analyze them.
3Do we know the rework rate and the top three reasons for rework?Sample recent records and categorize the returns.
4Have two or more qualified people independently reviewed the same sample, and do they agree most of the time?Run the agreement test. If agreement is low, fix definitions first.
5Are the key terms and decision rules written down, with examples?Write them (Step Three). This becomes the system specification.
6Does every step have one owner with clear entry and exit criteria?Fix the handoffs before automating any routing.
7Does the written procedure match the real process, with workarounds either documented or removed?Update the procedure through change control.
8Has the fixed process run manually long enough to measure its performance again?Run it and re-measure before building.
9Can we state which steps the system will perform, which stay with people, and how outputs will be used?Write down the system's role and how its outputs will be used before selecting a tool.
10Is the historical data we plan to train or test on consistent with the current definitions?Relabel a sample under current definitions, or do not use it as ground truth.

A process that scores “yes” on eight or more of these questions is usually ready to automate. A process that scores five or fewer will likely spend most of the automation budget on problems the process work would have solved more cheaply.

How to Present This to Leadership

Process work before automation can look like delay. The way to frame it is as the first phase of the automation project, with its own deliverables and its own measured return. A typical first phase for one process can be time-boxed to six to ten weeks: two to three weeks to map and measure, three to four weeks to fix definitions and handoffs and update the procedure, and the balance to run the fixed process and re-measure.

At the end of that phase, leadership receives three things: a measured improvement from the process fix, a clear decision on whether automation adds enough to justify the next investment, and, if it does, a specification and test set that make the automation project faster and easier to validate.

There is also a broader benefit. FDA’s Quality Management Maturity program, which is voluntary, aims to encourage drug manufacturers to adopt quality management practices that go beyond CGMP requirements. Its stated goals include fostering a strong quality culture mindset and acknowledging establishments that strive to continually improve quality management practices.15 Measured, documented process improvement is the kind of work that program is designed to encourage, whether or not a site takes part in an assessment.

When You Can Move Faster

Not every process needs the full treatment. If a process already has written definitions, high reviewer agreement, low rework, and a procedure that matches practice, you can move to automation quickly. The readiness check will show that. The point is not to delay every AI project. It is to avoid automating the ones that are not ready, and to know which is which before the money is spent.

It is also reasonable to run process work and early technical work in parallel, as long as the technical work does not lock in design decisions that depend on the process fix. Exploring what a tool can do with a sample of records is fine. Configuring it for production against definitions that are about to change is not.

Conclusion

The processes companies most want to automate with AI in pharma and biotech are review and investigation processes. They are high-volume, labor-intensive, and slow. They are also the processes FDA cites most often, and the ones most likely to carry undefined terms, loose handoffs, heavy rework, and historical records that reflect those problems. Automating them as they are produces the same failures faster and makes them harder to see. The sound approach is to fix processes before AI automation: map the process as it runs today, measure variation and rework, fix handoffs and definitions, and only then automate, using the fixed process as the specification. That sequence often delivers much of the improvement on its own, and it makes whatever automation follows easier to scope, validate, and defend.

Sakara Digital works with pharma and biotech organizations planning AI in quality and operations, and we often start with the process underneath the tool. If you are weighing an AI project for document review, deviations, batch record review, or another GxP process and want an independent view of whether the process is ready, we are happy to have that conversation.

For Further Reading