What a Backlog Is Actually Telling You

A deviation backlog is a capacity statement before it is a compliance statement. It says that over some period, the number of events entering the system was larger than the number of records the organization could close to its own standard. Everything else follows from that arithmetic. The compliance consequences are real, but they are consequences, not the disease.

This matters for how you respond. A team that treats a backlog as a discipline problem will add pressure, and the queue will shrink for about six weeks while the quality of the closed records falls. A team that treats it as an arithmetic problem will ask three questions instead: how many events arrive per week, how many can we close per week, and which of those events genuinely required an individual investigation in the first place.

In a single-site company the arithmetic is uncomfortable but visible. In a shared quality function covering several sites it is usually invisible, because nobody owns the total. Each site sees its own open records and believes its own list is small. The central team sees a queue with no natural owner and a set of priorities that shift every time a site escalates by telephone. The number that would tell you the truth, arrival rate against clearance rate across the whole network, is very often not being calculated by anyone.

84 Open and past due deviation investigations documented at one API facility as of June 21, 2024, per the FDA warning letter issued to Sanofi in January 2025
187 days The longest single past due example cited in that same letter, measured against the firm’s own procedural closure period2
~20% Share of bioreactor runs attempted between January 2022 and July 2024 that were rejected for contamination or other quality failures at that site2

The four causes, and only one of them is headcount

When we look at a queue that will not clear, the contributing causes almost always fall into four groups, and they need different responses.

Arrival volume. The process is generating more events than a stable process should. A rejection rate near one run in five, as in the example above, is not an investigation capacity problem. It is a process control problem that presents as an investigation capacity problem. No amount of investigator hiring fixes it, and clearing the queue without fixing the process only resets the clock.

Routing. Every event is being given the same treatment. A logbook entry made in the wrong column and a confirmed sterility failure both open a record with the same template, the same review chain and the same approval signatures. This is the most common cause we see, and it is the cheapest to fix, because the fix is a written rule rather than a person.

Authorship. The people who know what happened are not the people writing it up. Investigations queue behind a small number of trained writers in the central team who were not in the room, have to reconstruct the event from records and interviews, and produce a document that the site then disputes. The Sanofi response to the FDA identified exactly this pattern in its own root causes: attrition of trained investigators and process knowledge gaps among newer investigators.2

Approval. The record is written but waits for a signature. In shared functions this is usually a queue of one or two people who approve everything for every site, and it is the easiest bottleneck to miss because the aging report shows the record as open without saying which of the eight steps it is stuck on.

Diagnose before you plan. Split your open records by the step they are currently sitting in, not by their age. If most of the queue is waiting for approval, hiring investigators will make the queue longer. We have seen both failure modes, and the aging report alone cannot tell them apart.

What the Rules Require, and What They Leave to You

Before designing a triage rule it is worth being precise about what the regulations actually say, because a great deal of unnecessary investigation volume comes from internal procedures that are stricter than the regulation and were never revisited.

The United States: 21 CFR 211.192

The regulation requires that all drug product production and control records be reviewed by the quality control unit, and that any unexplained discrepancy or failure of a batch or any of its components to meet any of its specifications be thoroughly investigated, whether or not the batch has already been distributed. It requires a written record of the investigation including the conclusions and follow-up, and it requires the investigation to extend to other batches of the same drug product and other drug products that may have been associated with the specific failure or discrepancy.1

Two things are worth noticing. First, the trigger is an unexplained discrepancy or a specification failure, not every recorded event. Second, the regulation sets no closure period at all. It says thorough. The FDA’s own quality systems guidance frames the same point from the management side: quality risk management helps determine the extent of discrepancy investigations and corrective actions, and management should assign priorities to activities based on an assessment of risk including both the probability of occurrence of harm and the severity of that harm.10

The European Union: Chapter 1 and Chapter 8

EudraLex Volume 4, Chapter 1 states as a basic requirement of GMP that any significant deviations are fully recorded, investigated with the objective of determining the root cause, and appropriate corrective and preventive action implemented.6 The qualifier “significant” is doing real work in that sentence. Chapter 1 also requires that an appropriate level of root cause analysis be applied during the investigation of deviations, suspected product defects and other problems, and says this can be determined using quality risk management principles.

Chapter 1 contains one more sentence that every shared quality function should have on the wall. It says that while some aspects of the pharmaceutical quality system can be company-wide and others site-specific, the effectiveness of the system is normally demonstrated at the site level.6 A centralized model does not move the demonstration of effectiveness to the center. The inspector will still open a site and expect that site to show a working system.

Chapter 8, which covers complaints, quality defects and product recalls, is the closest thing in the regulations to a direct statement about the model this article describes. Clause 8.4 addresses handling that is managed centrally and requires the roles and responsibilities of the parties to be documented. It then adds a plain limit: central management should not, however, result in delays in the investigation and management of the issue.7 Clause 8.2 is equally direct about resourcing, requiring that sufficient trained personnel and resources be made available for the handling, assessment, investigation and review of complaints and quality defects.7

The sentence that defines the obligation

Centralizing is permitted. Centralizing and being slower than a site-based model would have been is not. That single clause converts backlog from an operational embarrassment into a documented deviation from the GMP guide, and it applies to the design of the model, not only to its worst week.

ICH Q10 and the basis for triage

ICH Q10 provides the sentence that makes a triage matrix defensible. Describing the corrective action and preventive action system, it states that a structured approach to the investigation process should be used with the objective of determining the root cause, and that the level of effort, formality and documentation of the investigation should be commensurate with the level of risk.9 Q10 also makes resourcing an explicit management responsibility, requiring management to provide adequate resources and to ensure that resources are appropriately applied to a specific product, process or site, and it requires communication processes that ensure appropriate and timely escalation of product quality and quality system issues.9

Q10 section 3.2.4 then closes the loop on reporting. Management review should include a timely and effective communication and escalation process to raise appropriate quality issues to senior levels of management, and among the actions it should identify is the provision, training or realignment of resources.9 A backlog reported into management review with a resourcing decision attached is Q10 working as designed. A backlog that never reaches management review is a Q10 gap in addition to whatever else it is.

The release consequence people forget

In the EU, Annex 16 makes an open investigation a supply chain event rather than a records event. A Qualified Person may certify a batch where an unexpected deviation has occurred provided registered specifications are met, but the deviation should be thoroughly investigated and the root cause corrected, and the impact assessed through a quality risk management process. Annex 16 also notes that responsibilities may be shared between more than one Qualified Person involved in the manufacture and control of a batch, and that the certifying QP should be aware of and take into consideration any deviations with the potential to affect compliance with GMP or the marketing authorization.8

For a multi-site network this is the practical reason the backlog gets political. An unclosed record at site A can hold certification of a batch that site B needs, and the person under pressure is a QP who has no authority over the investigator’s workload.

Triage: Three Gates Before Anything Gets a Number

Triage is not a way of investigating less. It is a way of deciding, in writing and before the fact, which events receive an individual investigation with its own root cause and CAPA, which are grouped into a periodic investigation of a known and previously characterized pattern, and which are recorded and trended with defined thresholds that promote them if the pattern changes.

The design principle we use is that triage should be a sequence of gates rather than a scoring model. Scoring models invite negotiation. Gates produce a decision that can be explained in one sentence.

Gate 1: Is product impact confirmed, plausible, or excluded?

This is a question about evidence, not about opinion. Confirmed means a specification failure, a confirmed out-of-specification result, a sterility or container closure integrity failure, a confirmed contamination event, or a complaint alleging harm. Plausible means the event touched a control that protects product and the evidence to exclude impact does not yet exist. Excluded means there is contemporaneous evidence, not an assumption, showing the product was not affected.

The rule that keeps this gate honest is simple: an event cannot be routed as excluded on the basis of a judgment that has not been written down and approved by someone independent of the operation. This is where the Genzyme Ireland warning letter of June 2026 is instructive. The FDA cited the firm under 21 CFR 211.192 for cancelling numerous deviations without investigating the root cause or assessing product impact, and framed it as a pattern of the quality unit failing to exercise proper oversight.5 Cancelling records is what triage looks like when there is no written gate behind it.

Gate 2: Is the batch pending release, or already distributed?

This gate sets urgency rather than depth. A record attached to a batch pending certification is on the critical path for supply and for the QP. A record attached to a distributed batch is on the critical path for patients, and 211.192 explicitly applies whether or not the batch has been distributed.1 A record attached to neither can be scheduled.

Separating urgency from depth is the single change that does the most to unstick a shared function. Teams that conflate the two end up giving deep treatment to whatever is loudest and shallow treatment to whatever is quiet, which is the opposite of a risk-based system.

Gate 3: Is the cause already characterized and under an open CAPA?

If an event is the twelfth instance of a pattern whose root cause was determined three months ago, and a CAPA is open and on schedule to address it, writing a twelfth individual root cause analysis adds no knowledge. It adds a record. Those events belong in a grouped investigation with a defined review cycle, an explicit link to the parent record and the open CAPA, and a threshold rule that promotes the group back to individual treatment if the rate changes, the severity changes, or the CAPA due date passes without effect.

The condition is that the grouping rule must be written and approved before the events arrive, not applied afterwards to a pile that has already grown. A rule written after the fact to explain a backlog is not a triage rule. It is a rationalization, and an inspector will read it as one.

The worked triage matrix

The matrix below is the shape we use with shared quality functions. The tiers are deliberately few. The last column, which defines what moves an event up a tier, is the part that most internal procedures omit and the part that makes the whole thing defensible.

Tier What belongs here Treatment Author / approver What promotes it
Tier 1
Critical
Confirmed OOS or batch failure; sterility failure; container closure integrity failure; Grade A or B environmental excursion with product exposure; confirmed cross-contamination; complaint alleging patient harm; any event triggering a field alert or regulatory notification Individual investigation. Opened on the day of awareness. Immediate containment decision documented. Scope expansion to other batches and products assessed and recorded regardless of outcome Site technical author with central quality investigator paired from day one. Approved by central quality head and, in the EU, reviewed with the certifying QP Nothing. This is the ceiling
Tier 2
Major
Deviation affecting a validated state or a GMP control with product impact not yet excluded; equipment fault during a batch; procedural nonconformance at a critical step; out-of-trend result; excursion between alert and action limits in a classified area Individual investigation on the standard path. Product impact assessment mandatory and independently reviewed. CAPA required if a root cause is determined Site technical author. Technical review by area management. Approved by central quality Product impact confirmed; second occurrence within the defined window; batch pending release becomes blocked
Tier 3
Grouped
Recurring event with a determined root cause and an open, on-schedule CAPA; repeat documentation practice findings of the same type; repeat minor equipment faults already characterized Recorded individually, investigated as a group on a defined cycle (we normally use monthly). Group record carries the parent root cause, the CAPA link, and the rate Site authors the entries. Central quality authors and approves the group record Rate exceeds the pre-approved threshold; severity changes; parent CAPA passes its due date; effectiveness check fails
Tier 4
Trended
Events that carry meaning only in aggregate: single alert-level counts in lower-grade areas within limits, minor logbook corrections captured at review, transient utility alarms with no product contact Recorded, coded and trended. No individual investigation. Reviewed at a defined frequency with written escalation thresholds Site records. Central quality owns the trend review and its output Any threshold breach; any coincidence with a Tier 1 or Tier 2 event on the same equipment, product or shift

Two rules make this matrix survivable. First, tier assignment is recorded with a one-line written rationale at the time of assignment, by name. Second, down-tiering after the fact requires the same approval as the original assignment plus a second reviewer. Without those two rules, a triage matrix becomes an instrument for making a backlog disappear on paper, which is a substantially worse finding than the backlog.

Risk Ranking That Survives a Site Challenge

In a shared function, ranking is not a private decision. Every ranking tells one site that its record will wait while another site’s record moves. That site will challenge it, sometimes with a good argument, and the design question is whether your ranking method can absorb the challenge without collapsing into whoever escalates hardest.

What inspectors look for in a risk decision

The PIC/S aide-memoire used by inspectors to assess quality risk management implementation is unusually specific about this, and it is worth reading as a design brief rather than an audit tool. It sets out that evidence should support the decisions made, that where specific risk elements are discounted this should be supported by appropriate rationale, and that inappropriate conclusions may be drawn where assessments rest on unjustified assumptions or incomplete identification.11 It also expects the scope, planning and scheduling of risk management activities to be organized, and risk management activities to be monitored, evaluated and reviewed for effectiveness.11

Read that as three obligations on your ranking method. It has to be based on stated evidence. It has to record what you decided not to worry about and why. It has to be reviewable, meaning someone can come back later and see whether the ranking turned out to be right.

Rank on four inputs, and publish them

We use four inputs, in this order, and the order is part of the method because it prevents the last one from dominating.

INPUT 1

Patient exposure

Is affected product on the market, in distribution, or held? Distributed product outranks everything, and 211.192 requires the investigation regardless of distribution status, so this is a ranking input rather than a scope question.

INPUT 2

Severity of the potential harm

Sterility, identity, strength and cross-contamination outrank appearance and packaging. Severity is assessed on the potential harm if the worst plausible reading of the evidence is true, not on the expected outcome.

INPUT 3

Evidence decay

How quickly does the ability to find the root cause disappear? Retained solutions, environmental plates, operator recollection and equipment state all degrade. An event with fast-decaying evidence moves up even at moderate severity, because waiting converts a solvable investigation into an unresolvable one.

INPUT 4

Supply consequence

Whether a batch is blocked for certification. This is real and it belongs in the ranking, but it is fourth, and it is recorded separately so that nobody can later claim the ranking was driven by supply pressure.

The challenge path, written down in advance

A site that disagrees with its ranking needs a route that is faster than an escalation to a vice president and slower than nothing. Ours has four properties.

  • A named challenger and a named decider. The site quality representative or site head may challenge. The decider is the central quality head or a delegate, and it is the same person every time so that outcomes are consistent.
  • A fixed window. Challenges are answered within two working days. A ranking dispute that takes two weeks has already imposed the delay it was trying to prevent.
  • New evidence, not new emphasis. The challenge must identify a fact the ranking did not have: a distribution record, a second occurrence, a test result, a shared component with another product. Restating the importance of the site’s schedule is not a challenge.
  • A written outcome that goes into the record. The decision, its basis, and who made it. This is what turns a disagreement into evidence of a functioning system rather than evidence of an argument.

The FDA’s quality systems guidance supports the underlying point, noting that it is important to engage appropriate parties in assessing risk, including manufacturing personnel and other stakeholders.10 A ranking produced by a central team with no site input is technically a ranking and practically a source of appeals.

A useful test. Take the three oldest open records in your queue and ask the site that owns each one to state, in a sentence, why it is not higher in the ranking. If they cannot, the ranking was never communicated. If they disagree, you have found a real dispute that has been resolving itself as delay rather than as a decision.

Who Writes the Investigation When the Site Has No Quality Headcount

This is the question that decides whether a shared model works. The central team cannot write every investigation for every site and remain current. The site has no quality headcount to write them. The answer that works is that authorship moves to the site and stays with people who are not in quality, while assessment and approval stay in the center.

That sentence tends to make quality leaders uneasy, so it is worth being clear about what it does and does not mean. It does not mean production investigates itself and approves the outcome. It means the person who was present writes the account of what happened and gathers the evidence, and an independent quality reviewer assesses the adequacy of the investigation, the product impact conclusion and the CAPA, and approves or rejects it. The independence that the regulations require is in the assessment and the approval. Nothing requires the narrative to be typed by a quality employee, and there is a strong practical argument that it should not be, because the person who was in the room writes a better account than someone reconstructing it two weeks later.

The roles, drawn explicitly

Activity Site operations Site technical author Central quality investigator Central quality approver / QP
Record the event and the time of awareness Does it Verifies completeness Monitors the daily intake list Not involved
Immediate containment (hold, quarantine, stop) Executes Documents Confirms sufficiency within the day Informed for Tier 1
Assign the tier Provides facts Proposes Assigns and records the rationale Arbitrates challenges
Write the investigation narrative and gather evidence Provides statements Owns and writes Coaches; writes only for Tier 1 Not involved
Root cause determination Contributes Proposes Challenges and confirms Reviews for Tier 1
Product impact and scope expansion Provides batch and distribution data Drafts Owns the assessment Approves
CAPA definition and owner Accepts actions Proposes Reviews adequacy Approves
Closure decision Not involved Not involved Recommends Approves
Trend, group review, periodic review Supplies data Attends Owns Receives output
Escalation to management review Not involved Not involved Prepares Presents

Making site authorship real rather than nominal

Naming a site technical author in a procedure does not create one. Four things do.

Named individuals with time protected in the shift plan. Two to four people per site, by name, with investigation time in their schedule. If investigation writing is what happens after the shift, it will be the first thing dropped in a busy week, which is exactly the week that generates the events.

Training that includes writing, not only the procedure. The failure mode is not that people do not know the SOP. It is that they write a chronology instead of an investigation, omit the evidence that excludes alternative causes, and stop at the first plausible explanation. Paired writing with a central investigator for the first three or four records fixes most of it.

A content template with required fields. Time of occurrence and time of awareness as separate entries. The alternative causes considered and the evidence that excluded them. The batches and products checked for the same signal, including those checked and cleared. Human error, if concluded, with the justification the guidance requires and evidence that process and system causes were examined first.

Rejection with a reason, early. A central reviewer who returns weak records within two days with a specific reason teaches faster than any training course. A central reviewer who rewrites the record instead trains the site to submit weak records forever, and moves the entire workload back to the center without anyone deciding to. That second pattern is the most common way shared quality models fail.

The independence question, answered plainly

Both 21 CFR 211.192 and the EU guide place the review and approval obligation on the quality unit, and the FDA restated the expectation directly in 2025: the quality unit has the responsibility to ensure that investigations are scientifically sound, consider all available data, expand to include all potentially affected drug products, and result in timely and appropriate actions to protect public health.4 None of that is compromised by a production supervisor writing the narrative. All of it is compromised if the same person writes and approves.

The Handoff and the Escalation Path

Most of the delay in a shared model is in the gaps between people, not in the work itself. A record moves from site to center and back, and each transfer has a queue in front of it. Reducing the number of transfers is usually worth more than speeding any one of them.

The handoff design

1

Same-day intake, once a day, in one place

Every site posts new events to a single network list by a fixed time each day. The central team reviews the list once, assigns tiers, and publishes the assignments the same day. One review event rather than a stream of individual notifications is what makes this affordable for a small central team.

2

Tier 1 gets a person, not a ticket

A named central investigator is assigned to a Tier 1 event within the day and works alongside the site author from the start. Evidence decay is fastest in the first 48 hours, and this is where central expertise earns its keep.

3

One consolidated review, not serial commentary

The central reviewer returns all comments once, with a decision: accept, accept with changes, or reject with reasons. Serial rounds of single comments are the largest hidden delay in most shared functions and the easiest to remove.

4

Approval slots, scheduled

Approvers hold fixed daily or twice-weekly slots for closure decisions. Approval that happens whenever the approver has a gap is the step that produces the most aging with the least visibility.

5

Extensions requested before the due date, never after

If a record will not close on time, the extension request is raised before the date, with a reason and a new date, and approved by the named approver. This is a procedural obligation in most quality systems and one of the most reliably cited failures when it is skipped.

The escalation path

Escalation should be triggered by defined conditions rather than by frustration. ICH Q10 requires communication processes that ensure appropriate and timely escalation of product quality and quality system issues, and describes management review as potentially a series of reviews at various levels of management with an escalation process attached.9 The conditions we normally define are these.

  • Any Tier 1 event, immediately, to the central quality head and the affected QP or release authority.
  • Any record that passes its due date without an approved extension, weekly, by name and by site.
  • Any record blocking release of a batch for more than five working days, to the supply and quality leads together.
  • Any Tier 3 group whose promotion threshold is breached, at the next trend review and not later.
  • Any site whose arrival rate has exceeded its clearance rate for four consecutive weeks, to management review with a resourcing recommendation attached.

That last one is the trigger most organizations lack, and it is the one that converts a capacity problem into a management decision before it becomes an inspection finding.

A 90-Day Plan to Work Down an Existing Backlog

This plan assumes the common starting position: several sites, one shared quality function, a queue in the low hundreds with a long tail, no reliable arrival and clearance figures, and a management team that has been told twice that it is under control.

It also assumes something that most plans skip. Reallocating people to clear a backlog is itself a change that affects other work, and regulators treat it as one. In the Sanofi case the firm committed to reassigning existing personnel and bringing in subject matter experts to address the investigation backlog. The FDA called the response inadequate in part because the firm did not explain how this would affect other operations, how these personnel would be trained to conduct specific aspects of the investigations, and how their performance would be monitored and assessed.2 Plan the reallocation with the same rigor you would plan a process change, and document it.

Week zero: count before you plan

Before day one, produce one table with a row for every open record: site, date of awareness, date opened, current step, tier if assigned, whether an extension exists and whether it was approved before the due date, whether a batch is blocked, and whether the same root cause appears elsewhere in the list. This usually takes three to five days and it is the only part of the plan that cannot be skipped. Every organization we have done this with found records nobody was working on and records that had been closed in the system but not in fact.

1

Days 1 to 30: stop the queue from growing

Approve and issue the triage matrix, including the grouping rules and the promotion thresholds, so that new arrivals route correctly from day one. Restore extension discipline: every open record past its date gets a documented, approved extension with a real new date this month, or it gets closed. Stand up the daily intake list and the scheduled approval slots. Publish the arrival and clearance rate weekly by site from week one, because the crossover point is the number the plan is actually managing. Do not re-triage the existing backlog yet.

2

Days 31 to 60: clear in priority order, not age order

Now re-triage the backlog against the matrix, with the rationale recorded for each. Work three streams in parallel: release-blocking records first, then Tier 1 and Tier 2 records with the fastest evidence decay, then the grouped records, which usually collapse a substantial share of the queue into a small number of parent investigations. Run paired writing sessions at each site for the first records so that site authorship is established while the queue is being cleared rather than afterwards. Expect the count to fall slowly in this window and the aging profile to improve first.

3

Days 61 to 90: close out and prove it held

Complete the remaining individual records, convert group findings into CAPAs with owners and dates, and run an independent review of a sample of the records closed during the effort, checking content quality rather than closure speed. Present the whole thing to management review with the arrival and clearance crossover, the aging profile, the sample review result, and an explicit resourcing recommendation. If the crossover has not happened, say so and quantify the gap in people rather than in effort.

The trap in the middle third. Around day 45 the count stops falling fast and pressure appears to close records with thinner content. This is the moment that determines whether the exercise leaves you better or worse off. Hold the sample review, keep the rejection rate visible, and be explicit with management that a slower curve with sound records is the intended outcome. A queue cleared with weak investigations reappears as a repeat finding, and the second time it is a network-level conclusion rather than a site-level one.

What good looks like at day 90

  • Clearance rate exceeds arrival rate for at least four consecutive weeks, at every site rather than in aggregate.
  • No open record without either a current due date or an extension that was approved before the original date passed.
  • The oldest open record is younger than the oldest record on day one by more than 90 days, which shows the tail was worked and not just the new arrivals.
  • Every site has at least two trained authors who have written and had accepted at least two records.
  • A management review record exists containing the numbers, the resourcing decision, and the named owner of the remaining gap.

Reporting That Keeps the Backlog Visible Without Punishing the People Clearing It

Reporting on a backlog is where good intentions do the most damage. A single overdue count published by site produces a league table, and league tables in quality produce down-tiering, early closure and quiet reclassification. The reporting design has to make the queue visible to management while making it unprofitable to game.

Report pairs, never single numbers

Every measure of speed is published next to a measure of quality, and neither is published alone. That pairing is the whole method.

Speed or volume measure Paired quality measure What the pair prevents
Records closed per week Share of a random sample of closed records judged adequate on independent review Closing thin records to make the count move
Open record count by site Tier distribution of new arrivals, month over month Down-tiering new events to keep the count low
Percent closed within the procedural period Percent of extensions approved before the due date Retroactive extensions used to erase lateness
Average age of open records Age of the oldest open record, and the count over 90 days An improving average that hides an untouched tail
CAPA closure rate Share of CAPAs with a completed effectiveness check that passed Closing actions without evidence they worked
Root causes determined Repeat rate: same root cause recurring within a defined window Assigning a convenient cause to close the record

Report the arrival rate as prominently as the queue

The queue is a stock. The arrival and clearance rates are the flows that produce it. A report showing only the stock invites management to ask the quality team why the number is not falling, which is the wrong question when the answer is that the process is generating more events than the system was designed to handle. Showing arrivals by site and by category alongside clearance moves the conversation to where it belongs, which is process control and resourcing rather than individual performance.

Route it into management review, with a decision attached

ICH Q10 identifies the provision, training or realignment of resources as an expected output of management review, and the FDA’s quality systems guidance lists realignment of resources among typical review outcomes.910 Use that. A backlog presented to management review without a resourcing recommendation is information. A backlog presented with a quantified gap, two options and a recommendation is a decision, and the record of that decision is a substantial part of what protects the organization later.

Say the quiet part in the report. If the honest position is that the network needs three more investigators and cannot have them this year, write that in the management review record together with the interim risk-based approach being used instead. A documented, risk-based decision to operate with known constraints reads very differently from an undocumented drift into the same position.

The Regulatory Exposure of the Backlog Itself

Leaders often assume that an open investigation is neutral until it is closed badly. It is not. The queue is itself evidence, and inspectors read it as evidence about the quality unit rather than about the individual events.

What an inspector concludes from a large queue

First, that the quality unit is not resourced or not empowered. The clearest illustration is the January 2025 warning letter in which the FDA documented approximately 84 open and past due deviation investigations at an API facility. The observation was cited as a failure of the quality unit to exercise its responsibility to ensure the API manufactured at the facility complied with CGMP, and the agency asked for a comprehensive assessment and remediation plan to ensure the quality unit is given the authority and resources, including adequate training, to function effectively.2 The backlog was not treated as a set of late documents. It was treated as proof about the unit.

Second, that your own procedure is the standard you will be measured against. In that same letter the FDA noted that the governing procedure required non-significant and significant deviations to be closed within defined periods and required a documented and approved extension request where a record could not be closed in time, and that multiple past due investigations lacked those extensions.2 There is no argument available about whether the period was reasonable. The firm wrote it, and the records did not meet it. This is why extension discipline is worth more than a faster closure target.

Third, that delay and inadequacy are the same finding. The FDA has repeatedly linked the two. In the Apotex letter of October 2025 the agency found that the firm did not properly investigate critical equipment failures and container closure integrity issues to evaluate root causes or assess impact on other potentially affected batches in a timely manner, and stated that the quality unit is responsible for ensuring investigations are scientifically sound, consider all available data, expand to include all potentially affected drug products, and result in timely and appropriate actions to protect public health.4 Timeliness appears in the same sentence as scope and scientific soundness.

Fourth, that a late investigation and a missing one converge. In the Turbare Manufacturing letter of September 2025 the FDA observed that the firm’s investigation failed to determine why an excursion had not been investigated at the time it occurred, and concluded that there was consequently no assurance of the firm’s ability to initiate investigations in a timely manner.3 A queue that grows because events are recorded late produces the same conclusion as one that grows because records are closed late.

Fifth, that a backlog at more than one site is a management finding, not a site finding. This is the exposure specific to the model this article is about. The Glenmark letter of July 2025 contains a section headed “Repeat Violations at Multiple Sites” listing prior warning letters at three facilities in the same network, two of them citing 21 CFR 211.192, and concludes that these repeated failures at multiple sites demonstrate that management oversight and control over the manufacture of drugs is inadequate.12 The same construction appears in the Medline letter of May 2026, which cites similar observations including inadequate investigations into OOS results at a second site in the company’s network and draws the same conclusion about management oversight.13

Sixth, that your resourcing answer will be tested for evidence. When Glenmark responded to a stability testing backlog with plans to increase laboratory capacity, purchase and qualify equipment and hire personnel, the FDA asked for the stability load assessment protocol with pre-defined criteria to evaluate the adequacy of resources.12 In other words, show the model. A capacity commitment without a capacity calculation is not a corrective action.

The reactive response problem. In the Apotex matter the FDA described the firm’s post-inspection commitments as a reactive approach that did not sufficiently explain how the quality system would be improved, and separately noted that meaningful corrective actions were not proposed until a regulatory inspection identified violations that had occurred over an extended period.4 A backlog that is identified, quantified and managed by you, with a documented plan and a management review record, is a different conversation from one identified by an investigator.

What the record should show

Assume an inspector pulls ten open records and five closed ones from your queue. What should be there.

  • Date of occurrence and date of awareness, recorded separately, with a short gap to the date opened. A long gap between awareness and opening is the finding that produced the Turbare conclusion above.
  • The tier and a written rationale, dated and attributable. Not a score with no explanation. One or two sentences saying why this event received this treatment.
  • The containment decision and its timing, including the decision not to contain where that was the conclusion, with its basis.
  • Scope expansion considered and recorded, listing other batches and products checked and cleared, not only those found affected. 211.192 requires the extension to other batches and products that may have been associated with the failure, and the record should show the question was asked.1
  • Extensions approved before the original due date, with reasons and revised dates. Retroactive extensions are worse than none.
  • Author and approver, with the approver independent of the operation. Site authorship is defensible. Site self-approval is not.
  • For grouped records, the pre-approved grouping rule and the promotion thresholds, plus evidence that a threshold breach actually triggered action on at least one occasion. A threshold that has never fired invites the question of whether it was set to avoid firing.
  • For trended events, the trend review output and its recipient, showing that aggregation led somewhere rather than into a report nobody read.
  • The management review record containing the backlog, the resourcing decision, and the follow-up from the previous review.

A note on the coming EU revision

PIC/S published a consultation document in 2025 proposing a revised Chapter 1 to reflect the changes introduced in ICH Q9(R1) on quality risk management, including language on a proactive approach to quality risk management and on informed and timely decisions.14 It is a consultation document and not in force. The in-force text remains the version applicable since January 2013.6 We mention it because the direction of travel is toward more explicit expectations on formality, proportionality and timeliness of risk-based decisions, which is precisely the ground a documented triage matrix occupies. Organizations that write those rules down now will have less to retrofit later.

Conclusion

A shared quality function with a backlog has usually been trying to solve the problem with effort, and effort is the one input that does not scale. The change that works is narrower and less dramatic: write down the rule that decides which events receive which treatment, move authorship to the people who were present while keeping assessment and approval independent, make the handoffs fewer and scheduled, and report the flows rather than only the stock. The regulations support every part of that. ICH Q10 asks for effort commensurate with risk, EU GMP Chapter 1 expects the level of root cause analysis to be determined using quality risk management principles, and Chapter 8 permits central management on the condition that it does not become the source of delay. What none of them permit is an undocumented rule, and an undocumented rule is what most backlogs actually are.

The uncomfortable part is that the queue is evidence whether or not you manage it. An inspector who finds a large aged queue will conclude something about your quality unit’s authority and resources, and the conclusion will be network-wide if the same pattern appears at a second site. The organizations that come through this well are not the ones with the smallest queues. They are the ones that counted honestly, decided in writing what would wait and why, took the resourcing question to management review with a number attached, and can show the record of that decision.

Sakara Digital works with pharma and biotech organizations designing quality operating models that hold up across several sites and one team. If you are looking at a deviation and investigation queue that will not clear and want an independent view of whether the problem is capacity, routing or authorship, we are happy to have that conversation.

For Further Reading