In This Article
- Executive Summary
- The Calendar Is Not a Risk Assessment
- The Constraint That Comes First: The Coverage Floor
- Building the Risk Model: Seven Inputs and How to Weight Them
- From Score to Schedule: Allocating Depth, Not Just Dates
- Reading the Data Honestly: Where Your Inputs Mislead
- Auditing the Areas Nobody Audits
- Program Metrics: Proving the Audit Function Works
- Governing the Model: Refresh, Override, and Inspection Defense
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Most internal audit programs in pharma and biotech run on a fixed cycle. Every area gets audited every two or three years because that is the schedule, and the schedule was set years ago by someone who is no longer in the role. The plan is predictable, easy to resource, and easy to explain. It is also almost entirely disconnected from where risk actually sits this year. Two areas can be equally overdue on the calendar while one has absorbed three system changes, a new product introduction, and a 40 percent turnover in supervisors, and the other has run the same process with the same people since the last audit.
The fix is not to abandon the schedule. It is to put a scored risk model in front of it, so that audit effort follows evidence rather than the date of the last visit. This article gives a working model with seven named inputs, scoring anchors, and weights you can argue with, plus the part most articles skip: the coverage floor. A risk-based program cannot leave an area unaudited indefinitely. Regulators expect the whole quality system to be covered over a defined period. The model reallocates depth and frequency within that floor. Get that wrong and a good idea becomes a finding.
What follows covers how to build and weight the model, how to convert scores into a real annual plan, where the input data misleads you (an area with few deviations may be well controlled or may simply be quiet about them), how to schedule the cross-cutting themes that no department owns, and which program metrics actually indicate the audit function is working. On that last point: findings closed on time is a weak measure. Repeat findings across cycles is a strong one, because a repeat finding means the corrective action did not work.
The Calendar Is Not a Risk Assessment
Walk into most quality organizations and ask how next year’s internal audit plan was built. The honest answer, more often than not, is that last year’s plan was copied forward and the dates were moved. Areas rotate on a two-year or three-year cycle. Sterile operations and quality control get audited more often because everyone agrees they should. Warehouse and facilities get audited less often for the same reason. Beyond that rough tiering, the plan is a rotation, not a risk assessment.
This is not laziness. The fixed cycle solves several real problems at once. It is predictable, so the auditee can plan around it and the audit team can be resourced a year in advance. It is easy to defend, because you can show an inspector a documented schedule and evidence that you followed it. It distributes attention evenly, which matters politically in an organization where nobody wants to feel singled out. And it satisfies the plain reading of the regulations, which ask for audits at defined intervals following a pre-arranged program.1
What the fixed cycle gets right
Before dismantling it, give the calendar its due. A rotation guarantees coverage. It removes the argument about who gets audited, which in a political organization is worth something. It produces a stable annual workload, which is how you keep a small audit team from being overwhelmed in one quarter and idle in the next. And it creates a defensible record: an approved schedule, executed as approved, with deviations from the schedule documented and justified. ICH Q7 asks precisely this of API operations, requiring regular internal audits performed in accordance with an approved schedule, with changes to that schedule justified and documented.8
Any risk-based model that cannot produce those same outputs is worse than the calendar it replaces. That is the bar.
The two failure modes
The fixed cycle fails in two directions at once, and both are expensive in audit-day terms.
The first is over-auditing the stable. An area that has run the same validated process, with the same trained staff, on the same equipment, with no changes and no significant deviations, gets the same three-day audit it got two years ago. The findings are minor and procedural. The audit team spends the time anyway because the calendar said so. Those days came out of a fixed annual budget.
The second is under-auditing the volatile. An area that absorbed a technology transfer, replaced its supervisor, brought a new automated system into GMP use, and generated a cluster of investigations is not scheduled until next year, because it was audited eighteen months ago. When the next audit finally arrives, the problems have compounded, and the corrective actions are harder and slower because the practices have set.
The calendar treats both areas identically. That is the whole problem in a sentence.
The regulator already made this move
There is a useful precedent here, and it is worth pointing at when someone in your organization argues that a fixed cycle is what regulators want. Until 2012, FDA was legally required to inspect domestic drug manufacturing establishments every two years, with no equivalent mandate for foreign sites. Section 705 of the Food and Drug Administration Safety and Innovation Act amended section 510(h) of the Federal Food, Drug, and Cosmetic Act to remove that fixed biennial requirement and replace it with a risk-based inspection schedule that considers named risk factors, including compliance history, the record of recalls, the inherent risk of the drug manufactured, and the time since the last inspection.9
Note the last factor. Time since the last inspection did not disappear from the regulator’s model. It became one input among several. That is exactly the move an internal audit program needs to make, and it is a model built and defended by the agency that will inspect you.
The Constraint That Comes First: The Coverage Floor
Here is where risk-based audit programs most often go wrong, and it is worth stating before any of the modeling detail, because everything else depends on it.
A risk-based audit program cannot leave an area unaudited indefinitely. The model is permitted to change how often you go, how deep you go, and how wide you scope. It is not permitted to drop an element of the quality system out of coverage because it scored low three cycles running. If your first risk-scored plan produces a list of areas that will not be visited for six or seven years, you have not built a risk-based program. You have built a finding waiting to be written.
What the regulations actually require
The expectation of full coverage over a defined period is not implied. It is written down in several places, and the wording is consistent.
EU GMP Chapter 9 states that personnel matters, premises, equipment, documentation, production, quality control, distribution of medicinal products, arrangements for dealing with complaints and recalls, and self-inspection itself should be examined at intervals following a pre-arranged program, to verify conformity with quality assurance principles.1 That list is a coverage scope, and self-inspection appears on its own list, which is a detail programs routinely forget: the audit function audits itself.
ICH Q7 requires regular internal audits performed in accordance with an approved schedule to verify compliance with the principles of GMP for active pharmaceutical ingredients.8 EudraLex Chapter 1 places self-inspection inside the pharmaceutical quality system as a mechanism for monitoring effectiveness and applicability, alongside the product quality review.2 ICH Q10 treats internal audit output as an input to management review of the quality system, which only works if the audits cover the system.7
On the US side, 21 CFR 211.22 assigns the quality control unit responsibility for approving or rejecting all components, procedures, and specifications affecting drug product identity, strength, quality, and purity, and requires written procedures for those responsibilities.12 Section 211.180(e) requires that records be maintained so data can be used to evaluate, at least annually, the quality standards of each drug product.11 Neither section says “internal audit” in those words, but together with FDA’s quality systems guidance they establish an expectation of periodic, systematic self-assessment covering the whole quality system rather than selected pieces of it.13
Then there is the practical test. When an inspector reviews your internal audit program, one of the first things asked for is the list of areas and systems in scope and the record of when each was last audited. A model that cannot answer that question cleanly for every element will not survive the conversation, regardless of how sophisticated the scoring is.
Setting the floor
Define the floor before you build the model, not after. Three decisions:
- Enumerate the audit universe. Every element of the quality system that must be covered, expressed as auditable units. Departments and physical areas, yes, but also the quality system processes themselves: deviation management, CAPA, change control, complaints and recalls, training, document control, supplier management, validation, computerized systems, and self-inspection. FDA’s six-system structure is a useful cross-check against a department list, because it forces you to confirm that the quality, facilities and equipment, materials, production, laboratory control, and packaging and labeling systems are all represented.10
- Set a maximum interval per unit. This is the floor. A common structure is a two-year maximum for units that touch product quality directly and a three-year maximum for supporting units, with certain high-consequence operations set at annual regardless of score. The specific numbers should be your own and should be justified in the audit procedure, not borrowed from a slide.
- Write the floor into the procedure as a constraint the model cannot override. If your scheduling procedure allows a risk score to defer an audit past its maximum interval, the procedure is wrong. The floor is a hard boundary; the score operates above it.
What the model is allowed to change
Inside the floor, the model has real freedom, and this is where the value comes from.
It can pull an audit forward. An area at a two-year maximum can be audited at nine months if the score justifies it. It can change depth: the difference between a four-day full system audit with sampling across every subprocess and a one-day focused audit against a defined set of questions is enormous in effort and in what it finds. It can change scope: a high-scoring area might get three targeted audits across a cycle rather than one broad one. It can change who audits: a high-score area may warrant a lead auditor with specific technical background rather than whoever is available. And it can change follow-up intensity, including whether a verification visit is scheduled independently of the next full audit.
All of that reallocation happens without breaching the coverage floor. That is the design.
Building the Risk Model: Seven Inputs and How to Weight Them
The model below is deliberately concrete. It uses seven inputs, each scored on a one-to-five scale with written anchors, each weighted, and summed to a single score on a 100-point scale. It is a formal quality risk management activity in the ICH Q9(R1) sense, and the revised guideline is explicit that the degree of formality should be matched to the importance, uncertainty, and complexity of the decision, and that subjectivity in risk assessment inputs and outputs is a known weakness to be managed rather than ignored.45
That last point is the reason for written scoring anchors. A score of 4 must mean the same thing to two different assessors, or the model is just opinion with arithmetic attached.
The seven inputs
| Input | What you measure | Typical source | Weight |
|---|---|---|---|
| Prior audit findings and severity | Count and severity mix of findings from the last two internal audits, plus any external or regulatory observations attributed to the area | Audit management system, inspection records | 20 |
| Deviation and CAPA rate | Deviations opened per unit of activity, severity mix, overdue investigations, CAPA effectiveness check failures | QMS | 20 |
| Change volume | Approved changes in the period, weighted by classification, including process, equipment, system, and supplier changes | Change control system | 15 |
| Product criticality and patient risk | Dosage form and route, sterility requirement, patient population, availability of alternatives, position in the process | Product register, quality risk assessments | 15 |
| Regulatory exposure and inspection history | Markets supplied, agencies with jurisdiction, time since last regulatory inspection, open commitments | Regulatory affairs, inspection log | 12 |
| Personnel turnover and experience | Turnover in the period for GMP-relevant roles, vacancy duration, share of staff below a defined experience threshold, supervisory change | HR, training records | 10 |
| Time since last audit | Months elapsed as a proportion of the maximum interval for that unit | Audit schedule | 8 |
Two things about that table deserve attention.
First, time since last audit carries the smallest weight in the model. That is the entire argument of this article expressed as a number. It is a real input, because elapsed time genuinely increases uncertainty about what is happening in an area, and because it is the mechanism by which a quiet area eventually rises up the list. But it is not the schedule. When it carries 8 of 100 points, an area that has been quiet, stable, and unchanged for eighteen months will not outrank an area that has generated repeat findings and absorbed a system replacement in six.
Second, the weights sum to 100 and are chosen, not derived. Nobody has a validated empirical weighting for this. What matters is that the weights are written down, approved, applied consistently, and reviewed when experience says they are wrong. If your program consistently finds serious problems in areas the model scored low, the weights are wrong and you should change them, with a documented rationale. That feedback loop is the difference between a model and a spreadsheet.
Anchoring the scores
Each input is scored one to five against written anchors. Here is what good anchoring looks like for the two heaviest inputs.
Prior audit findings and severity. A score of 1 means the last two audits produced no major or critical findings and fewer than a defined number of minor findings, all closed on first effectiveness check. A score of 3 means one major finding, or a pattern of minors clustered in a single subprocess. A score of 5 means a critical finding, a repeat major finding across cycles, or any finding attributed to the area in an external or regulatory inspection. The repeat condition matters more than the count: a repeat finding is evidence that the corrective action did not work, which is a different and more serious signal than a new problem.
Deviation and CAPA rate. This one needs normalization or it punishes busy areas for being busy. Express deviations per unit of activity, and choose the denominator that fits the area: batches produced, tests performed, shipments processed, records reviewed. A score of 1 is a rate in the lowest quartile across comparable units with no critical deviations and no overdue investigations. A score of 5 is a rate in the top quartile, or any critical deviation, or a CAPA that failed its effectiveness check, or a backlog of investigations past their procedural due date. Note that the last two conditions can push an area to 5 even at a low deviation rate. A quiet area with a failed CAPA effectiveness check is not a low-risk area.
The three inputs people underweight
Change volume is the single most predictive input in most programs and the one most often left out, because change control data sits in a different system from deviation data and nobody has joined the two. Weight changes by classification rather than counting them flat. Ten like-for-like component changes are not equivalent to one process change requiring revalidation, or one migration of a GMP-relevant computerized system to a new platform. Annex 11 is direct about this for computerized systems: risk management is applied across the lifecycle, and changes to systems are made in accordance with a defined procedure with appropriate assessment.3 Change volume is a proxy for how much of what you validated is still what is running.
Personnel turnover is uncomfortable to include because it can be read as blaming a department for its attrition. Include it anyway, and be careful about how you express it. What you are measuring is not people, it is the loss of undocumented process knowledge. An area where three of five experienced operators left in a year and were replaced by staff trained on the same SOPs has a genuinely higher probability of practice drift, because the SOPs never contained everything the experienced operators knew. Score supervisory turnover separately and more heavily than operator turnover, because supervisors are the local enforcement mechanism for procedural compliance.
Regulatory exposure is not the same as product criticality, though programs often collapse the two. An area supporting a product supplied to a market whose agency has not inspected the site in six years carries a different kind of exposure from an area supporting a domestic-only product inspected last year. Include open regulatory commitments in this input, since an area with an outstanding commitment from a previous inspection has a specific reason to be looked at before the next one.
A note on formality
ICH Q9(R1) added explicit treatment of formality in quality risk management, making the point that the degree of formality should be customized to the organization’s needs and the risks involved, and that this is a way to use resources sensibly rather than a way to avoid rigor.46
For audit scheduling, that translates cleanly. The annual scoring pass across the whole audit universe is a formal exercise: documented inputs, written anchors, approved weights, recorded output. A mid-year decision to pull one audit forward after a critical deviation does not need the same machinery. It needs a documented rationale and an approval. Confusing the two is how programs end up with a scoring model so heavy that nobody refreshes it.
From Score to Schedule: Allocating Depth, Not Just Dates
A score by itself does nothing. The step that most programs skip is translating the score into three separate decisions rather than one.
Three dials, not one
The instinct is to treat the score as a frequency dial: high score means audit sooner. That is the least useful of the three available adjustments.
- Frequency. When the next audit happens, bounded above by the coverage floor. High-scoring units move earlier in the year and may be visited more than once per cycle.
- Depth. How much evidence you sample and how far you follow a thread. A focused audit tests a defined set of questions against a small sample. A full system audit samples across every subprocess, follows records end to end, and interviews at multiple levels. The difference in audit days is often three to one.
- Scope. What is in the boundary. A high-scoring unit might be split so that a specific subprocess with a repeat finding history is audited separately, on its own schedule, rather than being one line in a broad audit that never gets to it.
Depth is where most of the value sits, because it is where audit days are actually consumed. Moving a low-scoring unit from a four-day full audit to a one-day focused audit releases three days without touching the coverage floor. Those three days go to the units that scored high. The total program size does not change. What changes is where the effort lands.
Banding
| Band | Score | Frequency | Depth | Follow-up |
|---|---|---|---|---|
| Tier 1 | 70 and above | Within 12 months; consider splitting into two targeted audits | Full system audit, lead auditor with subject matter background | Independent verification visit at 90 days on any major finding |
| Tier 2 | 45 to 69 | Within 18 to 24 months | Full audit at standard depth | Effectiveness check at next scheduled audit unless a major finding is raised |
| Tier 3 | Below 45 | At the coverage floor maximum | Focused audit against a defined question set, with escalation criteria to expand scope on site | Documented closure review |
The escalation criteria for Tier 3 are not optional. A focused audit that uncovers something unexpected must have a written trigger allowing the auditor to expand scope on the spot, and the plan must carry contingency days to absorb that. Without it, a light-touch audit becomes a way of not looking, which is precisely the criticism a risk-based program has to answer.
Budgeting audit days
Build the plan against a fixed pool of audit days, because that is the actual constraint. Reserve fifteen to twenty percent of the pool as unallocated contingency for mid-year triggers: a critical deviation, a recall, a failed regulatory inspection at a comparable site, an unplanned system go-live. A plan with no contingency will either break when something happens or will absorb the shock by quietly cancelling a Tier 3 audit, which puts you back at risk of breaching the floor.
Freeze the audit universe
Confirm the list of auditable units, including quality system processes and cross-cutting themes, not just departments. Reconcile it against the six-system structure and against the elements named in EU GMP Chapter 9. Record additions and removals with rationale.
Pull the input data
Extract the seven inputs for a defined lookback period, usually 24 months. Do this from source systems, not from memory. Where a data element is not available at the unit level, record that gap rather than substituting a guess.
Score and calibrate
Two assessors score independently against the written anchors, then reconcile. Differences greater than one point on any input indicate the anchor wording is ambiguous. Fix the anchor, not the score.
Apply the coverage floor
Before looking at the ranked list, place every unit that hits its maximum interval next year. These are fixed. The score then allocates depth and sequence to what remains.
Allocate days and review overrides
Assign depth by band, sum the days, and reconcile against the pool. Every manual change to the model output is logged as an override with a named approver and a written reason. Overrides are a data set, not an embarrassment.
Approve and publish
The plan goes to the quality council or equivalent governance body with the score table attached. Publishing the scores, not just the dates, is what makes the plan defensible and what makes area management take the inputs seriously.
Reading the Data Honestly: Where Your Inputs Mislead
Every input in the model is a measurement of something you cannot observe directly. Deviation counts are not a measurement of how much goes wrong; they are a measurement of how much goes wrong and gets written down. Confusing the two produces a model that systematically directs audit effort away from the areas that most need it.
The deviation paradox
An area with very few deviations is either well controlled or under-reporting. Those two states produce identical numbers in the QMS and opposite conclusions for the audit plan. A naive model scores both as low risk, which means the under-reporting area is the one place in the organization your audit program is structurally least likely to visit. That is the worst possible failure mode for a risk-based schedule, because it is self-reinforcing: the less you look, the fewer deviations get raised, the lower the score, the less you look.
The mirror image is just as damaging. An area with a high deviation count may have a genuine control problem, or it may have an honest reporting culture, an engaged supervisor, and a habit of writing up anything ambiguous. Penalizing that area with a high risk score and a heavy audit teaches the rest of the organization that reporting attracts scrutiny. You will get exactly the behavior you incentivized.
Interpret rate alongside severity mix
The way out is to stop reading deviation rate on its own and always read it against the severity mix and the detection point.
A healthy reporting culture produces a characteristic shape: a relatively high volume of low-severity events, caught early, raised by the people doing the work, with a small tail of significant events. An under-reporting culture produces the opposite shape: a low total volume, but a severity distribution skewed toward major and critical, because only the events too large to absorb get written up. The small stuff never enters the system.
So the diagnostic is not the count. It is the ratio.
The under-reporting signature. Look for an area where all of the following hold at once. Any one alone is noise; three or more together is a reason to schedule an audit regardless of what the composite score says.
- Low total deviation volume with a high proportion classified major or critical. The minor events are missing, not absent.
- A high share of events raised by someone outside the area. QA on record review, the next process step, the customer, or the receiving site. If the area rarely raises its own events, it is not looking or not telling.
- Late detection relative to occurrence. Measure the gap between the event date and the date the record was opened. Under-reporting areas show long, variable gaps because events surface at review rather than at occurrence.
- Downstream error density that does not match upstream deviation density. If batch record review, laboratory data review, or QA release routinely finds entry errors, missing signatures, or documentation gaps in an area that reports almost no documentation deviations, the two data sets are telling different stories and one of them is wrong.
- Complaint or out-of-specification trends pointing at an area with a clean internal record. The events are being detected. They are being detected somewhere else.
That last set of checks is straightforward to build once. It needs deviation data, batch record review findings, complaint data, and out-of-specification records joined at the area level. Most organizations already hold all four and simply have never put them side by side. The check runs annually as part of the scoring pass and takes an analyst a day or two once the extracts exist.
The wider point is one regulators have been making for years in the quality culture conversation, and it applies directly here: the number of deviations is a weaker signal than how the organization detects, investigates, and prevents recurrence.13 Build the model to reflect that.
Other inputs that lie to you
Change volume undercounts unmanaged change. The change control system only contains changes that went through change control. An area that makes undocumented adjustments to a process shows a low change volume, which lowers its risk score. Cross-check change records against equipment maintenance records, master data change logs, and system audit trails. A mismatch between what the audit trail shows and what change control recorded is a finding in itself and a strong signal for the model.
Prior findings reflect prior audits. If an area was last audited by a light-touch focused audit, it has few findings, and few findings lowers its score, and a lower score means another light-touch audit. Correct for this by scoring findings per audit day rather than findings in absolute terms, and by carrying forward a flag for any unit whose last audit was Tier 3 depth. The flag is a reminder that low findings might mean low looking.
Turnover data lags reality. HR turnover reports usually reflect departures, not the period of degraded capability that precedes them: the notice period, the vacancy, the ramp-up. An area that lost a supervisor eleven months ago may still be operating below normal capability. Use a rolling window rather than a point-in-time count, and weight vacancy duration alongside departure count.
Product criticality is stable, which makes it dominant. Because criticality barely moves year to year, a heavy weight on it will effectively hard-code the same units at the top of the list every cycle, which reproduces the fixed schedule you were trying to escape. Keep the weight moderate, and treat criticality as a modifier on the dynamic inputs rather than as a driver in its own right.
Auditing the Areas Nobody Audits
The most consistent blind spot in internal audit programs has nothing to do with scoring. It is structural. Audit plans are built from the organization chart, and several of the highest-risk subjects in a modern pharma or biotech operation do not live anywhere on the organization chart. They cut across it.
Why the organization chart hides risk
Consider data integrity across systems. Manufacturing owns the execution system. Quality control owns the laboratory system. IT owns the infrastructure and the identity management. Quality assurance owns the procedures. Every one of those functions gets audited on its own schedule, and every one of them can pass. The risk sits in the joins: the interface where a result moves from an instrument to the laboratory system, the account that exists in one system and not the other, the audit trail that is enabled in one application and disabled in another, the data that gets transformed between systems with no record of the transformation.
There is external evidence for where this matters. MHRA has published its GMP inspection deficiency data grouped by the GMP chapter and annex references the deficiencies were cited against, which makes visible how much of the finding volume sits in the cross-cutting parts of the quality system rather than in a single production area.18 An audit plan built purely from the organization chart is not structured to look where those deficiencies are being written.
PIC/S guidance on data management and integrity is built around exactly this view, treating data across its full lifecycle and giving specific attention to computerized systems and to outsourced activities where the data crosses an organizational boundary.14 MHRA’s data integrity guidance takes the same position, expecting a data integrity risk assessment that considers the process, the system, and the way data flows through both.15 Neither of those assessments maps to a department. A department-based audit plan cannot deliver them.
Data integrity across system boundaries
Scope by data flow, not by system owner. Follow one critical data element from generation to reported result to archive, across every system and manual step it passes through. Test audit trail configuration, review practice, interface controls, transformation logic, and access at each hop.
Supplier and third-party oversight
Scope by the oversight process, not by the supplier list. Test how suppliers are risk-classified, how audit frequency is set and evidenced, how quality agreements are kept current when scope changes, and how third-party findings feed your own CAPA system. The contract giver retains responsibility for control of outsourced activities regardless of who performs them.
Computerized system change control
Scope by the change process across the application estate. Test whether GMP-relevant changes go through GMP change control rather than the IT ticket queue, whether validation impact is assessed and evidenced, whether periodic review happens, and whether the configuration in production matches the configuration that was qualified.
The audit program itself
EU GMP Chapter 9 lists self-inspection among the things to be examined. Test whether the plan was executed as approved, whether overrides were justified, whether findings were graded consistently across auditors, and whether the risk model inputs were pulled from source rather than assembled from recollection.
How to scope a theme audit
Theme audits fail when they are scoped as “audit data integrity,” which is unbounded and produces a report nobody can act on. Scope them as a traced path with a defined start and end.
For a data integrity theme audit, pick two or three specific data journeys and follow each one completely. One might start at an analytical instrument and end at the certificate of analysis. Another might start at a weighing operation and end in the batch record and the release decision. Every system, every manual transcription, every review step, every place the data is copied, transformed, or aggregated is in scope for that journey. Everything else is out. The audit produces findings about specific controls at specific hops, which is actionable, rather than a general observation that data integrity needs attention, which is not.
The same discipline applies to the other themes. A supplier oversight theme audit follows three suppliers of different risk classifications from qualification through performance monitoring to requalification, and tests the process at each step. A change control theme audit takes a sample of changes made to GMP-relevant systems in the period, including changes that were made outside GMP change control, and traces each one to its assessment, its validation impact, and its production state.
Who owns the findings
The unavoidable problem with theme audits is that findings land on shared ground, and shared ground has a habit of belonging to nobody. A finding that the interface between two systems loses attribution data is not obviously the manufacturing department’s finding or the IT department’s finding.
Handle this at the planning stage, not at the closing meeting. Every theme audit is assigned a single accountable owner before the audit starts, at a level senior enough to direct work in more than one function. That is usually a quality system process owner rather than a department head. The owner receives the findings, and the corrective action plan names the contributing functions. Without that, theme audit findings age quietly and reappear next cycle as repeat findings, which is the one outcome the whole program exists to avoid.
A practical allocation. Reserve a defined share of the annual audit day pool for theme audits before allocating anything to department audits. Programs that allocate departments first and themes with whatever is left over never run theme audits, because there is never anything left over. Ten to fifteen percent of the pool is enough to run two or three theme audits a year at real depth.
Program Metrics: Proving the Audit Function Works
Ask a quality leader how the internal audit program is performing and you will usually get two numbers: percentage of planned audits completed, and percentage of findings closed on time. Both are weak. They measure whether the audit department is administratively organized. Neither measures whether the audit program is finding the things that matter or whether the organization is getting better as a result.
Why on-time closure is a weak measure
On-time closure measures the speed of paperwork. It is fully satisfiable by closing findings with shallow corrective actions inside the due date, which is exactly what a team under pressure on that metric will do. A finding closed in 30 days with a retraining action and no root cause analysis scores identically to a finding closed in 30 days with a process redesign. The metric cannot tell them apart, and the first one will come back.
It also creates a perverse incentive on grading. If closure timelines are tied to severity, and the metric is closure performance, there is quiet pressure to grade findings lower. Watch for a severity mix that drifts toward minor over time with no corresponding change in what the audits actually found. That drift is a program failure, not an improvement.
Keep on-time closure. It is worth knowing. Just stop treating it as the headline.
Repeat findings are the strong measure
The metric that tells you whether the program works is the repeat finding rate: the proportion of findings in the current cycle that address a condition already found in a previous cycle, in the same unit or in a comparable one.
A repeat finding is unambiguous evidence that the corrective action taken last time did not work. Either the root cause was wrong, or the action did not address it, or the action was implemented and did not hold. All three are serious, and none of them show up in a closure-rate metric, because the original finding was closed on time. Regulators reach the same conclusion from the other direction: recurrence is a standard test of whether a corrective and preventive action was effective, and closing a CAPA without a genuine effectiveness check is among the most common weaknesses in the system.16
Two refinements make the metric far more useful.
Count repeats across units, not just within them. The same finding raised in three different areas in the same year is not three local problems. It is one systemic problem with three symptoms, and treating it locally three times guarantees it returns. A cross-unit repeat should trigger a systemic assessment at the quality system level rather than three separate local CAPAs.
Measure the interval to recurrence. A finding that returns after four years is a different signal from one that returns after nine months. Short-interval recurrence points at implementation failure. Long-interval recurrence points at sustainability failure, which usually means the control depended on specific people and those people left. The two need different responses.
| Metric | What it tells you | Strength |
|---|---|---|
| Repeat finding rate | Proportion of findings addressing a previously found condition, within and across units | Strong. Direct evidence of corrective action effectiveness |
| Internal versus external detection ratio | Share of significant findings first raised by internal audit rather than by a regulator, customer, or partner audit | Strong. The clearest test of whether the program is looking in the right places |
| CAPA effectiveness failure rate | Proportion of audit-driven CAPAs that fail their effectiveness check | Strong. A leading indicator of next cycle’s repeat findings |
| Coverage completeness | Proportion of the audit universe within its maximum interval at any point in time | Strong. The floor is either intact or it is not |
| Model predictive accuracy | Correlation between risk score band and severity of findings actually raised | Moderate. Small sample sizes, but it is the only feedback the weights get |
| Time from event to detection | Average age of the practice or condition at the point the audit found it | Moderate. Hard to measure precisely, useful as a trend |
| Findings closed on time | Administrative discipline in CAPA execution | Weak. Necessary hygiene, easily satisfied without improvement |
| Audits completed against plan | Whether the department executed its own schedule | Weak. Measures the audit team, not the quality system |
The detection ratio deserves more attention than it gets
Of the strong metrics, internal versus external detection is the one most likely to change how leadership thinks about the audit function. It asks a simple question: when something significant is found, who found it first?
If a regulatory inspection, a customer audit, or a partner qualification audit consistently raises significant findings that your internal program had not identified, the program is looking in the wrong places or not looking hard enough. That is a direct and uncomfortable measurement of audit effectiveness, and it is exactly the measurement the risk model exists to improve. It also gives you a clean way to evaluate the model itself: if the areas where external parties found problems were scoring in your lowest band, your weights need work.
FDA’s continuing interest in quality metrics, including the 2022 public docket reopening how a reporting program should be structured after two industry pilots, reflects the same underlying question about which measurements actually indicate quality system performance rather than administrative activity.17 The internal audit program is a good place to apply that thinking, because the data is already yours and nobody has to agree on a common definition before you can use it.
Governing the Model: Refresh, Override, and Inspection Defense
A risk model that is built once and never revisited becomes a fixed schedule with extra steps. The governance around the model matters as much as the model.
Refresh cadence
Run the full scoring pass annually, aligned to the planning cycle. Between passes, define named triggers that cause a single unit to be rescored mid-year without redoing the whole model:
- A critical deviation or a confirmed out-of-specification result with product impact
- A regulatory inspection at the site, or at a comparable site in the network, that raises findings relevant to the unit
- A CAPA that fails its effectiveness check
- Go-live of a GMP-relevant computerized system, or a significant change to one
- Departure of a unit’s quality lead or operations supervisor
- A new product introduction or technology transfer into the unit
Each trigger produces a documented rescore and a decision: pull the audit forward, expand the scope of the planned audit, or note the change and hold the plan. Recording the decision to hold matters as much as recording the decision to act, because it is the evidence that the trigger was assessed.
The override log
Every risk model gets overridden. Someone senior will say that an area cannot be audited in the second quarter because of a campaign, or that a unit should be audited despite a low score because of something the model does not capture. That is legitimate. Human judgment is part of a risk-based system and ICH Q9(R1) does not pretend otherwise.
What matters is that overrides are logged as data rather than absorbed into the plan invisibly. Record the unit, the model output, the change made, the reason, the approver, and the date. Then review the log annually. A pattern of overrides in the same direction is telling you something about the model. If the same senior leader consistently pulls one function’s audits later, that is worth a conversation. If assessors consistently override the score upward for a particular input, the anchor for that input is set wrong.
Data quality of the inputs
The model is only as good as the extracts feeding it, and those extracts come from systems that were not designed to feed an audit risk model. Deviation records may not carry a consistent area code. Change records may be attributed to the requester’s department rather than the affected one. Turnover data may be reported at a level of aggregation that does not match your auditable units.
Deal with this openly. Document the mapping from source data to auditable unit, record known gaps, and where a gap is material, record how it was handled. If turnover data is not available at unit level and you used a departmental proxy, write that down. An inspector who sees a documented limitation and a documented workaround sees a controlled process. An inspector who finds an undocumented assumption inside a model that drives audit scheduling sees something else.
What an inspector will ask
Prepare the program to answer five questions cleanly.
The five questions
- What is in scope of your internal audit program, and how do you know it is complete? Answer with the audit universe and the reconciliation against the quality system elements and the six manufacturing systems.
- When was each element last audited? Answer with the coverage report showing every unit against its maximum interval. Any breach must already have a documented justification.
- How did you decide this year’s plan? Answer with the procedure, the scored model, the weights, and the approval record.
- Show me a finding and what happened to it. Answer with the finding, the root cause analysis, the action, the effectiveness check, and the closure. Choose your own example before they choose one for you.
- How do you know your audit program is effective? Answer with the repeat finding rate and the internal versus external detection ratio, trended, with what you changed in response.
The fifth question is the one that separates programs. Most can answer the first four. A program that can answer the fifth with real trended data and a specific example of a change it made because the data said so is demonstrating a functioning quality system, not just a functioning audit department.
Conclusion
Risk-based audit scheduling is not a way to do fewer audits. It is a way to spend the same audit days where they will find more. The model does that by replacing a single implicit input, time since last audit, with seven explicit ones, then using the result to adjust depth and scope rather than only dates. The coverage floor stays fixed underneath all of it, because regulators expect the whole quality system examined over a defined period and no scoring model changes that. Programs that skip the floor and go straight to the scoring build something clever that will not survive its first serious inspection.
The harder part is not the arithmetic. It is being honest about what the inputs mean. A quiet area is either well controlled or not telling you things, and the model as normally built cannot distinguish them, which sends audit effort away from the one place it is most needed. Reading deviation rate against severity mix, detection point, and downstream error density is what makes the difference, and it uses data almost every organization already holds. The same honesty applies to how the program measures itself. Findings closed on time tells you the department is organized. Repeat findings tell you whether anything is actually getting better, because a repeat finding means the corrective action did not work, and no amount of on-time closure changes that.
Sakara Digital works with pharma and biotech organizations building this kind of quality analytics: risk models that hold up in an inspection, data joins that turn separate quality systems into a usable picture, and program metrics that measure improvement rather than activity. If you are rethinking how your internal audit plan gets built and want an independent perspective on where to start, we are happy to have that conversation.
For Further Reading
For Further Reading
- CAPA Effectiveness Reviews: The 90-Day Look-Back I Recommend
- Deviation Trending Analytics: From Excel to Real-Time Dashboards
- The Quality Metrics Dashboard That Actually Drives Investigations to Closure
- The Supplier Quality Risk Heat Map Template
- Quality Culture Pulse Survey: A Quarterly Cadence for Pharma Teams
- Inspection Readiness Is a Mindset
References & Sources
- European Commission. “EudraLex Volume 4, EU Guidelines for Good Manufacturing Practice, Chapter 9: Self Inspection.” https://health.ec.europa.eu/system/files/2016-11/cap9_en_0.pdf
- European Commission. “EudraLex Volume 4, Chapter 1: Pharmaceutical Quality System.” January 2013. https://health.ec.europa.eu/system/files/2016-11/vol4-chap1_2013-01_en_0.pdf
- European Commission. “EudraLex Volume 4, Annex 11: Computerised Systems.” January 2011. https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf
- ICH. “Q9(R1) Quality Risk Management, Step 4 Guideline.” 18 January 2023. https://database.ich.org/sites/default/files/ICH_Q9(R1)_Guideline_Step4_2022_1219.pdf
- U.S. Food and Drug Administration. “Overview of Changes: ICH Q9(R1) Quality Risk Management.” https://www.fda.gov/media/177720/download
- European Medicines Agency. “ICH Q9 Quality Risk Management, Scientific Guideline.” https://www.ema.europa.eu/en/ich-q9-quality-risk-management-scientific-guideline
- U.S. Food and Drug Administration. “Guidance for Industry: Q10 Pharmaceutical Quality System.” April 2009. https://www.fda.gov/media/71553/download
- U.S. Food and Drug Administration. “Q7 Good Manufacturing Practice Guidance for Active Pharmaceutical Ingredients: Guidance for Industry.” https://www.fda.gov/files/drugs/published/Q7-Good-Manufacturing-Practice-Guidance-for-Active-Pharmaceutical-Ingredients-Guidance-for-Industry.pdf
- U.S. Food and Drug Administration. “FDASIA Section 705 Annual Reports on Inspections of Establishments.” https://www.fda.gov/regulatory-information/food-and-drug-administration-safety-and-innovation-act-fdasia/fdasia-section-705-annual-reports
- U.S. Food and Drug Administration. “Compliance Program 7356.002: Drug Manufacturing Inspections.” https://www.fda.gov/media/75167/download
- Electronic Code of Federal Regulations. “21 CFR 211.180: General Requirements (Records and Reports).” https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-J/section-211.180
- Electronic Code of Federal Regulations. “21 CFR 211.22: Responsibilities of the Quality Control Unit.” https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-B/section-211.22
- U.S. Food and Drug Administration. “Guidance for Industry: Quality Systems Approach to Pharmaceutical CGMP Regulations.” September 2006. https://www.fda.gov/media/71023/download
- Pharmaceutical Inspection Co-operation Scheme. “PI 041-1: Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments.” 1 July 2021. https://picscheme.org/docview/4234
- Medicines and Healthcare products Regulatory Agency. “‘GXP’ Data Integrity Guidance and Definitions.” March 2018. https://www.gov.uk/government/publications/guidance-on-gxp-data-integrity
- Electronic Code of Federal Regulations. “21 CFR 211.192: Production Record Review and Investigation of Discrepancies.” https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-J/section-211.192
- Federal Register. “Food and Drug Administration Quality Metrics Reporting Program; Establishment of a Public Docket; Request for Comments.” 9 March 2022. https://www.federalregister.gov/documents/2022/03/09/2022-04972/food-and-drug-administration-quality-metrics-reporting-program-establishment-of-a-public-docket
- Medicines and Healthcare products Regulatory Agency. “MHRA GMP Inspection Deficiency Data Trend.” https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/609030/MHRA_GMP_Inspection_Deficiency_Data_Trend_2016.pdf








Your perspective matters—join the conversation.