In This Article
- Executive Summary
- What the Rules Actually Require, and What They Leave Open
- The Compliance Work That Is Genuinely Automatable
- The Work That Is Not, and Why
- The Hollowing-Out Path Is the Default
- What a Compliance Leader Has to Learn, and Where to Stop
- Moving Upstream: From Reviewing Proposals to Shaping Them
- A Capability Model for the Evolving Role
- Questions to Ask About Any Proposed AI Use Case
- The Organizational Conditions That Decide the Outcome
- When the Role Is One Person, or a Fraction of One
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
A large share of what compliance and quality staff do every day is checking. Checking that a document has every required section. Checking that an audit trail contains no unexplained edits. Checking a batch of regulatory updates for the two that matter. That work is now partly machine-doable, and the honest question is not whether the role changes but in which direction. Two outcomes are available from the same starting point. In one, the compliance leader becomes an advisor whose judgment shapes decisions before they are made. In the other, the function becomes a smaller team approving machine output it no longer has the time or the standing to question.
The second outcome is the default. It happens without anyone choosing it, because removing review effort is easy to measure and upgrading judgment work is not. Research on automation bias is consistent across four decades: people who monitor an automated aid check it less carefully than they check themselves, and training does not fix it89. A rubber-stamp function passes every internal metric right up until an inspection asks a question the tool did not anticipate.
This article separates the compliance work that automates well from the work that does not, and explains what actually determines which of the two outcomes a company gets. It offers a capability model for the evolving role, a set of questions a compliance leader should be able to ask about any proposed AI use case, the organizational conditions that make the advisory role possible or block it, and a version of all of this for the company where quality and compliance is one person, or part of one.
What the Rules Actually Require, and What They Leave Open
Start with what is written down, because a surprising amount of the argument about this role is conducted as if regulators had already settled it. They have not.
In the United States, 21 CFR 211.22 says there shall be a quality control unit with the responsibility and authority to approve or reject components, containers, closures, in-process materials, packaging, labeling, and drug products, and the authority to review production records to assure that no errors have occurred or, if errors have occurred, that they have been fully investigated. It says that unit approves or rejects all procedures and specifications affecting identity, strength, quality, and purity, and that its responsibilities and procedures shall be in writing and followed1.
Read that carefully. It assigns responsibility and authority. It does not say how many people. It does not say what those people spend their day doing, what tools they use, whether the first read of a record is done by a person or a system, or where the function reports. Everything about the operating model is left to the company.
ICH Q10 is broader and equally quiet on operating models. Its second section is Management Responsibility, and it covers management commitment, quality policy, quality planning, resource management, internal communication, management review, management of outsourced activities and purchased materials, and management of change in product ownership2. Those are obligations placed on senior management, not a job description for a compliance officer.
EU GMP Chapter 1 puts the same weight on the top of the house: senior management has the ultimate responsibility for an effective pharmaceutical quality system, which requires the participation and commitment of staff across the company and must be adequately resourced with competent personnel3. Chapter 2, on personnel, adds two clauses that matter more to this discussion than anything in the AI literature. First, the manufacturer should have an adequate number of personnel with the necessary qualifications and practical experience, and senior management should determine and provide adequate resources. Second, and this is the sentence to remember: the responsibilities placed on any one individual should not be so extensive as to present any risk to quality4.
No regulator prescribes an operating model. There is no clause in 21 CFR 211, ICH Q10, or EU GMP Part 1 that defines the compliance officer role, sets a ratio of reviewers to batches, or blesses a particular division of labor between people and software. What the rules fix is responsibility, authority, competence, and sufficiency of resource. How you arrange the work inside those constraints is a management decision, and it is yours to defend.
That is genuinely freeing and genuinely uncomfortable. It means nobody is going to tell you that your redesigned function is compliant in advance. It also means the failure mode is well documented. FDA warning letters routinely cite firms because the quality unit failed to exercise its responsibility to ensure products were manufactured in compliance with current good manufacturing practice and met established specifications6. The citation is not about headcount. It is about a function that held the authority on paper and did not use it.
The Compliance Work That Is Genuinely Automatable
Being specific here is the whole point. Vague claims that AI will handle routine compliance work are what produce bad plans. Below is the work that current systems handle well enough to change how a team is staffed, with an honest note on what each still needs from a person.
Document review for completeness
Checking a protocol, a validation package, a batch record, or a submission section against a required structure is a matching problem. Systems do it faster than people, consistently, at any hour, across every document rather than a sample. This is the single most reliable win available to a compliance function today, because completeness is close to objective. A section is present or it is not. A signature is captured or it is not. A cross-reference resolves or it does not.
What still needs a person: deciding whether a section that is technically present is actually adequate. A rationale paragraph can exist and be empty of reasoning. The tool will pass it.
Audit trail screening
Audit trail review is the clearest example of work that was never done properly by hand because the volume made it impossible. A person reviewing a sample of entries is performing a ritual. A system that screens every entry against defined patterns, such as records modified after approval, deletions without a reason code, activity outside working hours, or repeated changes by one user on one batch, produces a short list a human can actually study.
What still needs a person: everything after the short list. The screen tells you what is unusual. It cannot tell you whether unusual is wrong. That judgment is the job.
Trend detection
Deviations, complaints, environmental monitoring excursions, and out-of-specification results all carry signal that shows up over months and across sites, which is exactly the timescale human attention handles worst. Statistical monitoring and pattern detection surface candidate trends earlier and with fewer misses than quarterly manual review.
What still needs a person: the difference between a statistical signal and a quality problem. Detection is cheap now. Interpretation is not.
Regulatory change monitoring
Watching agency websites, guidance registers, inspection observation summaries, and pharmacopoeial updates across every market a company operates in is a task that scales badly with people and well with software. Automated monitoring can retrieve, classify, and route changes and flag the subset that touches a company’s products, markets, or systems.
What still needs a person: deciding what a change means for this company. The regulators themselves are working through this. The joint HMA and EMA multi-annual AI workplan through 2028 sets out how the European regulatory network intends to use AI for productivity, process automation, and decision support, and it treats capability building and change management for regulators as a workstream in its own right14. If the agencies think the human capability piece needs a dedicated program, industry should not assume it comes free.
First-pass gap assessment
Comparing an existing procedure set against a new guidance, an updated standard, or another company’s quality system during due diligence used to take weeks. A first-pass mapping is now hours. The output is a draft gap list with citations.
What still needs a person: confirming that a mapped gap is real and that an unmapped area is genuinely covered. First-pass output is a starting position for expert review, not a finding.
| Compliance task | What the system does well | What a person must still supply | Evidence you need to keep |
|---|---|---|---|
| Document completeness review | Exhaustive structural checking against a defined template, every document rather than a sample | Judgment on adequacy and scientific soundness of content | Template version, model or rule version, list of documents screened, exceptions raised and dispositioned |
| Audit trail screening | Full-population screening against defined risk patterns | Investigation and disposition of every flagged entry | Screening criteria, date range, hit list, reviewer decision and rationale per hit |
| Trend detection | Cross-site, multi-month pattern detection at a sensitivity people cannot sustain | Decision on whether a signal is a quality problem and what action follows | Method and parameters, signals raised, signals dismissed and why |
| Regulatory change monitoring | Continuous retrieval, classification, and routing across many jurisdictions | Impact assessment against this company’s products, markets, and systems | Sources monitored, coverage gaps acknowledged, impact assessments and owners |
| First-pass gap assessment | Rapid mapping of a procedure set against a new requirement | Confirmation of each gap, and of each area claimed to be covered | Source documents compared, draft output retained, expert review record |
Two points about that table. The first is that the right-hand column is not administrative overhead. It is the evidence that lets you explain the arrangement to an inspector, and it is the reason continuous inspection readiness and AI-supported review belong in the same design conversation rather than separate projects.
The second is that none of these five tasks are eliminated. Each is split into a machine part and a human part, and in every case the human part is the harder half. That is the fact most transition plans get wrong.
The Work That Is Not, and Why
There is a common thread running through the compliance work that resists automation, and naming it is more useful than a list. Every item below requires someone to take a position under genuine uncertainty and then answer for it. Not to compute an answer. To take a position.
Deciding what risk is acceptable
ICH Q9(R1) is unusually direct about this. It devotes a section to managing and minimizing subjectivity, acknowledging that subjectivity can affect every stage of a quality risk management process, that it enters through differences in how risks are perceived and through inadequately defined risk questions, and that while it cannot be completely eliminated it should be managed and minimized5.
Read that as a statement about what risk decisions actually are. They are not calculations with a subjective error term to be scrubbed out. They involve a judgment about how much harm is tolerable in exchange for what benefit, made by people who are accountable for being wrong. A model can supply the inputs to that judgment. It cannot hold the accountability, and a model that produces a confident risk score mostly moves the subjectivity somewhere less visible.
Judging whether an investigation reached root cause
A well-written investigation and a genuinely closed one look identical on the page. Both have a problem statement, an investigation summary, a stated root cause, a corrective action, and an effectiveness check. The difference is whether the stated cause explains the observed facts better than the alternatives someone did not consider.
Detecting that difference means holding the physical process in your head, noticing that the proposed cause would have produced a different pattern of failures than the one observed, and being willing to reopen an investigation the business considers closed. A text system trained on completed investigations learns what an accepted investigation looks like. That is close to exactly the wrong thing to learn, because the population it learned from includes every weak investigation that was accepted.
Interpreting a regulation against a novel situation
Regulations are written before the situations they will be applied to. Somebody has to decide whether a decentralized manufacturing arrangement, a new modality, or a model that updates on new data is covered by an existing clause, covered by analogy, or genuinely unaddressed. That decision requires knowing not only the text but the reasoning behind it, the inspection history around it, and the tolerance of the specific regulators involved for a well-argued position.
A system can retrieve every relevant clause and every published precedent. It cannot judge which argument a particular inspector will accept, and it cannot decide how much regulatory exposure the company should accept in order to move faster.
Saying no to a business sponsor
This is the one that never appears in role descriptions and determines whether the rest of the function is real. Someone senior wants to release, launch, or deploy. The evidence is incomplete. The compliance leader has to say no, or say yes with conditions, in a room where saying no has consequences for them personally.
No tool changes that. What tools can change is the quality of the evidence you bring to the conversation, and that matters more than it sounds. A refusal supported by full-population screening data is a different conversation than a refusal supported by an opinion. But the willingness to hold the position is a human property of a person in an organization that will back them.
The test that separates the two lists. If the work can be judged right or wrong by comparing an output to a defined standard, it is a candidate for automation. If the work involves deciding what the standard should be in a case the standard did not anticipate, or being accountable to a regulator for a decision that could reasonably have gone the other way, it is not. Most compliance activities contain both, which is why the honest unit of redesign is the task, never the role.
The Hollowing-Out Path Is the Default
Here is the part that most writing on this subject avoids. The claim that automating review work frees compliance leaders to be strategic contains a hidden step: someone has to actually build the strategic work, fund it, and protect the time for it. If nobody does, the function does not become strategic. It becomes smaller, and what remains is approval of output it did not produce and cannot easily challenge.
Three separate bodies of evidence say this is the likely path rather than the pessimistic one.
First: people trust automated output more than they should, and know it less than they think
Automation bias is one of the better-established findings in human factors research. A systematic review of automation bias in healthcare documented its frequency, the factors that make it worse, and the interventions that reduce it, concluding that over-reliance on automated aids is common and that mitigations are only partly effective8. A broader integration of the complacency and bias literature found that both appear in novice and expert users alike, that they are driven by how attention is allocated under task load, and that they are not removed by simple practice or by instructions to be careful9.
Translate that into a quality function. A reviewer whose queue is now pre-screened, whose exceptions are pre-highlighted, and whose throughput target has been raised because the tool is helping will check the machine’s work less carefully than they would have checked their own. That is not a discipline failure. It is the predicted result.
Second: the time savings are frequently assumed rather than measured
A randomized controlled trial of experienced open-source developers working on their own repositories found that allowing early-2025 AI tools increased task completion time by 19 percent. The developers had predicted a 24 percent reduction before starting and still believed, after finishing, that the tools had made them roughly 20 percent faster10.
The finding to take from that study is not that AI tools do not help. It is that practitioners’ estimates of their own time savings were wrong by about forty points and confidently held. If a compliance transition plan assumes freed capacity without measuring it, the freed capacity may not exist, while the headcount reduction that was planned against it is entirely real.
Third: the restructuring being recommended looks the same either way
Consultancy guidance on the next generation of compliance leadership, written for banking but broadly applicable, argues for moving from a posture of constraint to one of enablement, for leaner operations with more senior managers rather than large junior-heavy teams, and for the compliance leader taking a leading role in both AI adoption and governance17.
That is a reasonable design. It is also, from the outside, indistinguishable from cutting the junior tier and calling the survivors strategic. The organizational chart after a successful upgrade and after a hollowing-out are the same chart. Only the work is different, and work is harder to see than headcount.
Five signs the function is being hollowed out rather than upgraded
- Exception rates fall and nobody can explain why in terms of the process rather than the tool.
- Reviewers cannot describe what the system would miss. If they cannot name its failure modes, they are not overseeing it.
- The time saved was booked into the budget before it was measured in the workflow.
- No compliance decision in the last year was reversed, escalated, or contested. Real judgment produces friction.
- Compliance is invited to AI conversations at validation, and only at validation.
What determines which way it goes
The difference is not the technology and not the quality of the people. In the companies where this goes well, four things are true and they are all decisions made by senior management, not by the compliance function.
The freed hours are assigned before they are freed
Not “we will use the capacity for higher-value work.” A named piece of work, an owner, and a date. Unassigned capacity is reabsorbed by the queue or removed from the budget within two cycles.
Someone is accountable for the machine part
A defined owner for the screening logic, its performance, its drift, and its failure modes. Where nobody owns the tool, the reviewer inherits accountability for output they cannot interrogate.
The function is measured on something other than throughput
Cycle time and backlog are the metrics automation improves fastest, which makes them the metrics that hide hollowing out. Add a measure of judgment quality, such as investigations reopened on review or decisions revised after challenge.
Compliance is present at use case selection
If the first compliance touchpoint is validation, the role has already been defined as approval. Presence at selection is the structural difference between an advisor and a gate.
What a Compliance Leader Has to Learn, and Where to Stop
The advisory role requires technical understanding, and the most common mistake is misjudging how much. Too little and the advice is generic. Too much and the compliance leader becomes a junior technologist with a compliance title, which removes the independence the role exists to provide.
The target is precise: enough about how these systems behave to ask questions that change a design, and not enough to build one. A compliance leader who can build a model has learned the wrong thing, in the same way a quality director who calibrates the instruments has drifted out of role.
The five things worth learning properly
Where the behavior comes from
That the system’s output is a function of the data it learned from and the instructions it is given at run time, so questions about training data, reference data, and prompt or configuration control are questions about the product, not about the infrastructure. You do not need the math. You need to know that changing any of the three changes the output.
How output varies
That the same input can produce different output on different runs, and that this is a design property to be constrained rather than a defect to be reported. Once you know this, you know to ask how a repeatable result is achieved for a regulated decision, and what is recorded so the decision can be reconstructed later.
What the edges look like
That performance is highest on cases resembling what the system learned from and degrades on cases that do not, without any signal that it has degraded. This is the single most useful piece of technical knowledge a compliance leader can hold, because it converts directly into a question about the operating range and what happens outside it.
How it fails
That the characteristic failure is a fluent, plausible, confident answer that is wrong, rather than an error message. This changes how a review step must be designed, because a human checking for obvious errors will not find these.
What human oversight means in the workflow
The difference between a person who can meaningfully review, one who can only accept or reject, and one whose approval is a keystroke on a queue of two hundred. All three are described as human in the loop. Only the first is oversight.
Where to stop
Do not learn to code models. Do not take ownership of model performance metrics. Do not become the person who defends a technical design in a technical argument, because the moment you do, your independent challenge of it is compromised. Keep a working relationship with people who do own those things, and keep the questions coming from outside the build.
The literacy floor is now partly a legal one
For companies operating in the European Union, part of this is now a legal expectation. Article 4 of the EU AI Act requires providers and deployers of AI systems to take measures to support the development of AI literacy among their staff and others operating systems on their behalf, taking account of those people’s technical knowledge, experience, education, and the context of use. The Commission confirmed that the AI system definition, the AI literacy provisions, and a limited set of prohibited practices became applicable on 2 February 2025, and that it would set up a living repository of AI literacy practices drawn from providers and deployers13.
Treat that as a floor, not a program. A generic AI awareness module satisfies a training record and does nothing for the compliance leader who has to challenge a use case next Tuesday. The industry conversation has caught up here: workforce preparedness and organizational readiness were given equal billing with technical readiness at ISPE’s 2026 AI in Life Sciences programming, with quality teams described as moving toward judgment and context while systems handle exhaustive pattern detection1516. The framing is right. The work of making it true is local.
Moving Upstream: From Reviewing Proposals to Shaping Them
Every compliance leader has had the conversation where a project arrives with a vendor selected, a budget approved, a go-live date announced, and a request for the validation approach. At that point the compliance contribution is limited to making a fixed decision defensible. The advisory role is the same expertise applied four months earlier, when the decision is still open.
The gap between those two positions is not a matter of being invited. It is a matter of being useful at a stage where compliance has historically had nothing to offer except delay.
What early actually means
| Stage | What the business is deciding | What compliance contributes | What it is worth |
|---|---|---|---|
| Problem framing | Whether this problem is worth solving with AI at all | Whether the regulated decision at the end of the process can be supported by this class of evidence | Kills unworkable use cases before anyone spends money |
| Use case selection | Which of six candidates to fund first | A rank ordering by regulatory difficulty alongside the business rank ordering | Changes the sequence, which is the single input with the widest effect |
| Intended use definition | What the system will and will not be used for | Wording that will hold up under inspection, and the boundary of the operating range | Prevents the scope drift that invalidates the original assessment |
| Acceptance criteria | What good enough looks like | The threshold that a regulated decision requires, set before anyone sees a demo | Removes the argument where a demo becomes the standard |
| Vendor evaluation | Which supplier | Which suppliers can produce the evidence you will need, and which cannot | Avoids discovering during validation that the documentation does not exist |
| Validation | Nothing. It has been decided. | Execution of a plan that should already be obvious | Low, which is the point |
The two rows that change outcomes are use case selection and acceptance criteria. Rank ordering candidate use cases by regulatory difficulty is a piece of analysis nobody else in the company can do, it takes about a day, and it routinely changes which project goes first. Setting acceptance criteria before the demo removes the most common failure in AI procurement, where an impressive demonstration on curated examples becomes the implicit performance standard.
How to earn the earlier seat
Nobody hands this over on request. Three things reliably produce the invitation.
Be fast. The reason compliance is engaged late is that engaging compliance early has historically meant a four-week wait for a written position. A compliance leader who gives a specific, provisional, written-down answer within two days becomes someone people call before they are required to. Provisional is fine. Say what would change your view.
Be specific about what is fine. A function that returns a risk to every proposal is a function whose input can be predicted and therefore skipped. Say plainly which parts of a proposal need nothing from you. Credibility is spent by warning about everything.
Bring something they do not have. The regulatory difficulty ranking above. A list of what a given class of use case will require in evidence. The two questions the inspector will ask about this system in 2028. Arrive with material rather than with process.
A practical starting move. Take the AI use case list that already exists somewhere in the company, whether it is a formal portfolio or an informal collection of pilots, and produce a one-page regulatory difficulty ranking of it. Three tiers, one sentence of reasoning each, and a note on what would move an item down a tier. Send it unrequested. It is a day of work and it is the fastest route from reviewer to advisor, because it demonstrates the contribution in a form the business can use rather than describing it.
This runs alongside two adjacent problems worth treating separately: how quality and IT divide responsibility so validation does not become the delay point, and how a company handles staff outside IT who are building their own tools. Both change what the advisory role has to cover, and neither is solved by moving compliance upstream on its own.
A Capability Model for the Evolving Role
Capability models fail when they are aspirational. This one is written so that each level can be evidenced by something a person has actually produced, which also makes it usable for hiring, development planning, and honest self-assessment.
Five domains. Three levels. The levels are not seniority grades. A capable director may be at level 3 in two domains and level 1 in another, and knowing which is more useful than an average.
| Domain | Level 1: Informed | Level 2: Practicing | Level 3: Advisory |
|---|---|---|---|
| Regulatory judgment | Knows the applicable requirements and can locate the clause | Applies requirements to situations they did not anticipate and documents the reasoning | Takes and defends a position where the requirement is genuinely unsettled, and states the exposure being accepted |
| System behavior literacy | Understands what the system does and what data it uses | Can describe how it degrades, where its operating range ends, and what its characteristic failure looks like | Designs the oversight step around the specific failure mode rather than applying a standard review template |
| Evidence design | Knows what records the current process produces | Specifies in advance what evidence a decision will require and confirms the system can produce it | Designs the evidence set for a use case that has no precedent, and can explain to a regulator why it is sufficient |
| Business framing | Understands what the business is trying to achieve | States compliance positions in terms of the business decision at stake, not the clause | Proposes an alternative route to the business goal when the proposed one will not hold |
| Standing | Is consulted when required by procedure | Is consulted before it is required, by people who have a choice | Has held a position against senior pressure, been supported, and is known to have been |
How to evidence each level
For each domain, ask for the artifact rather than the self-assessment. Regulatory judgment at level 3 is evidenced by a written position on an unsettled question, with the reasoning and the accepted exposure stated. System behavior literacy at level 2 is evidenced by a one-page description of a system’s failure modes written by the person, not by the vendor. Evidence design at level 2 is a pre-specified evidence list dated before the build. Business framing at level 3 is a proposal that changed a plan. Standing at level 3 cannot be evidenced by the person at all, which is the point of putting it last: you find out by asking the business.
Reading the model
The domain that most commonly limits an otherwise strong compliance leader is not system behavior literacy. It is business framing. Technical understanding can be built in a quarter with focused effort. The habit of stating a compliance position in terms of the business decision rather than the requirement takes longer, because it requires giving up the safety of citing a clause and instead owning a recommendation.
The domain nobody can develop on their own is standing. It is created by what happens the first time a compliance leader says no to someone more senior. That is a decision made by the executive team, once, and everyone in the building learns the answer.
Questions to Ask About Any Proposed AI Use Case
This is the practical core of the advisory role. A compliance leader who can ask these questions confidently, in a room, without preparation, is already operating as an advisor whatever the title says. They are grouped by what they are actually testing.
Several of these have a parallel in enforcement expectations outside pharma. The US Department of Justice’s Evaluation of Corporate Compliance Programs, updated September 2024, directs prosecutors to ask how a company assesses the impact of new technologies such as AI on its ability to comply with the law, whether management of that risk is integrated into broader enterprise risk management, what controls exist to monitor trustworthiness and reliability, whether controls ensure the technology is used only for its intended purposes, what baseline of human decision-making is used to assess AI, how accountability over its use is monitored and enforced, and how employees are trained on it7. That document is about criminal exposure rather than GMP, but the questions are the right ones and they are on the record.
Group 1: What decision is this actually supporting?
- What decision changes because of this system, and who makes that decision today?
- Is that decision one a regulator will ask us to justify? Which one, and under what clause?
- If the system were removed tomorrow, what would happen to the decision? If the answer is nothing, the system is not doing what the business case claims.
- What is the intended use, in one sentence, that we would be willing to write into a procedure?
Group 2: What does the system do when it is wrong?
- What does a wrong output look like here, and would a reviewer recognize it as wrong?
- Where does performance degrade, and is there any signal when it has?
- What is the defined operating range, and what happens to inputs outside it?
- What is the worst realistic outcome of an undetected error, and who bears it?
Group 3: Is the human oversight real?
- What baseline of human decision-making are we comparing this against, and do we have that baseline measured or assumed?
- How many items will the reviewer handle per hour, and is meaningful review possible at that rate?
- Does the reviewer see the input, or only the output? Review of an output alone is acceptance.
- What has to happen for a reviewer to disagree with the system, and has anyone done it yet?
- Is the reviewer’s throughput target changing because of this system? If so, oversight has been traded for capacity.
Group 4: Can we reconstruct what happened?
- Six months from now, can we reproduce the output for a specific record, including the version of the system and the data that produced it?
- What is recorded about the human review step beyond the fact that it occurred?
- How would we detect that the system’s behavior has changed?
- Who is authorized to change the configuration or the instructions, and how is that controlled?
Group 5: Who owns this?
- Who owns the system’s performance, distinct from who owns the process it supports?
- Who decides it should be turned off, and on what evidence?
- What are we relying on the supplier for, and what happens if they change it?
- If this produced a finding at inspection, whose name is on the response?
The question that ends most weak proposals. “What does a wrong output look like here, and would the reviewer recognize it?” If nobody in the room can answer, the proposal has not been thought through, whatever the demonstration showed. It is not a hostile question and it does not require technical knowledge to ask. It requires only the willingness to ask it in front of people who would rather move on.
The Organizational Conditions That Decide the Outcome
An individual compliance leader can develop every capability above and still be unable to operate as an advisor, because several of the conditions are structural. Being honest about which ones you control matters, because effort spent on conditions you do not control is effort taken from the ones you do.
| Condition | What it looks like when it enables the role | What it looks like when it blocks the role | Who controls it |
|---|---|---|---|
| Reporting line | Compliance reports independently of the function whose work it reviews | Compliance reports to the operational leader whose targets it constrains | Executive team |
| Point of entry | Compliance is on the intake path for new technology proposals | Compliance receives proposals at validation | Executive team, with strong influence from compliance |
| Measurement | Judgment quality is measured alongside cycle time | The only reported metrics are backlog, cycle time, and on-time closure | Compliance leader, mostly |
| Capacity assignment | Freed hours are assigned to named work before the automation goes live | Freed hours are assumed, then removed in the next budget round | Compliance leader, in negotiation with finance |
| Backing on refusal | A compliance no has been upheld against senior pressure, visibly | Refusals are negotiated away or escalated around | Executive team, entirely |
| Ownership of the tooling | A named owner is accountable for screening logic and its performance | The tool is owned by IT as infrastructure and by nobody as a control | Compliance leader and IT jointly |
Two of these deserve comment.
Measurement is more available than it looks. Most compliance functions report cycle time and backlog because those are the numbers the systems produce, not because anyone decided they were the right ones. Adding two judgment measures is usually within the compliance leader’s own authority: the proportion of investigations reopened on second review, and the number of decisions revised after challenge. Both go up when the function is working and down when it is rubber-stamping, which is the opposite direction from throughput and exactly why they are worth reporting.
Backing on refusal is the one you cannot manufacture. If the executive team has never supported a compliance refusal under commercial pressure, no capability model changes the outcome. It is worth knowing this clearly rather than treating it as a personal development problem. A compliance leader in that environment should be honest with themselves about what the role can be, and about whether the answer is acceptable.
A note on the argument you will be offered. When capacity is being reduced, the case will be made that the remaining team is more senior and therefore more capable, so a smaller function is a stronger one. That can be true. The way to test it is to ask what specific judgment work the senior team will now do that it could not do before, and to ask for it by name and date. If the answer is a category rather than a piece of work, the upgrade is not being planned. It is being asserted.
When the Role Is One Person, or a Fraction of One
Most of what is written about the evolving compliance role assumes a department. A large share of biotech does not have one. Quality and compliance is a director with a coordinator, or a head of quality who is also the head of regulatory, or two days a month of a consultant. The advice above does not scale down by simply doing less of it.
Start with the clause that applies most directly. EU GMP Chapter 2 states that the responsibilities placed on any one individual should not be so extensive as to present any risk to quality4. That is not aspirational language. It is a requirement, and in a small company it is the requirement most often breached without anyone noticing, because breaching it looks like a committed person working hard.
What a one-person function should automate first
The sequencing is different from a large company’s, because the constraint is different. A large function automates to increase coverage. A small function automates to survive.
Regulatory change monitoring
The highest-value automation for a small function, because the alternative is not slow monitoring, it is no monitoring. A single person cannot watch multiple agencies across multiple markets. Automated retrieval and classification converts an uncovered risk into a manageable weekly review.
Document completeness checking
Second, because it is the work that consumes the most hours for the least judgment, and because in a small company those hours are directly displacing the judgment work that only this person can do.
Audit trail screening
Third, and only with defined criteria and a documented sampling rationale. In a small company this is often the difference between a review that is genuinely performed and one that exists as a signature.
First-pass gap assessment, when it comes up
Not routine work in a small company, but when a new guidance arrives or a partner audit is coming, the first pass is the part that would otherwise not happen at all.
What never to delegate to a tool, at any size
The final disposition of a deviation. The decision that an investigation is closed. The release decision. The position taken with a regulator. The decision to accept a risk. In a one-person function these are the entire job, and they are exactly the activities that a tired person under time pressure is most tempted to accept from a system. The small-company failure mode is not over-investment in AI. It is a person with no capacity accepting machine output because there is no time to do otherwise.
What to buy from outside
EU GMP Chapter 2 also addresses consultants directly: they should have adequate education, training, and experience to advise on the subject for which they are retained, and records should be maintained of their name, address, qualifications, and type of service provided4. That clause exists because the arrangement is expected and normal, not because it is a compromise.
The parts of the advisory role that buy well from outside are the ones that are periodic rather than continuous: the regulatory difficulty ranking of a use case portfolio, an independent challenge of a proposed oversight design, a second opinion on an unsettled interpretation, and a structured review of whether the automation in place is actually being overseen. The parts that do not buy well are the ones that require standing inside the company, which is most of the rest.
The realistic small-company target. One person cannot be at level 3 in all five capability domains. Aim for level 3 in regulatory judgment and standing, level 2 in system behavior literacy and evidence design, and buy the rest periodically. That is a defensible arrangement and it can be explained to an inspector, which is the correct test. What is not defensible is a one-person function nominally responsible for all five, operating at level 1 in three of them, with nobody having said so out loud.
Conclusion
The compliance and quality leadership role is changing because a real part of its daily work has become machine-doable. That much is settled. What is not settled is whether the role that emerges is larger in influence and smaller in headcount, or simply smaller. Both are available from the same starting conditions, and the difference is made by decisions that mostly happen before any technology is selected: whether the freed hours are assigned to named work, whether anyone owns the machine part, whether the function is measured on anything besides speed, and whether a compliance refusal has ever been supported when it mattered.
The regulations will not decide this. They fix responsibility, authority, competence, and adequacy of resource, and they leave the operating model to management, which means the operating model has to be designed on purpose and defended on its merits. The compliance leaders who come out of this period with more influence than they had will be the ones who learned enough about how these systems behave to ask the questions that change a design, moved their contribution from validation back to use case selection, and were honest with their executive team about the difference between an upgraded function and a hollowed-out one while there was still time to choose.
Sakara Digital works with pharma and biotech organizations redesigning how their quality and compliance functions operate as AI takes on more of the review work. If you are working through what your compliance leadership role should become, and want an independent perspective on which parts of the current job to automate and which to protect, we are happy to have that conversation.
For Further Reading
For Further Reading
- Reskilling Your Quality Team for AI Oversight: A 6-Month Curriculum
- From Quality to AI Leadership: A Cross-Functional Career Pivot in Life Sciences
- AI Governance Framework for Pharma QA Teams
- Human-in-the-Loop Requirements for Pharma AI: What FDA and EMA Actually Expect
- The Fractional Leadership Operating Model for Series B Biotechs
References & Sources
- U.S. Food and Drug Administration. “21 CFR 211.22: Responsibilities of quality control unit.” Electronic Code of Federal Regulations, current edition. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-B/section-211.22
- International Council for Harmonisation. “ICH Q10: Pharmaceutical Quality System.” Step 4 version, 4 June 2008. https://database.ich.org/sites/default/files/Q10%20Guideline.pdf
- European Commission. “EudraLex Volume 4, Part I, Chapter 1: Pharmaceutical Quality System.” January 2013. https://health.ec.europa.eu/system/files/2016-11/vol4-chap1_2013-01_en_0.pdf
- European Commission. “EudraLex Volume 4, Part I, Chapter 2: Personnel.” Revision effective 16 February 2014. https://health.ec.europa.eu/document/download/11f4f8e6-a6e9-4897-afe3-f21e1dc56cb8_en
- International Council for Harmonisation. “ICH Q9(R1): Quality Risk Management.” Step 4 version, 18 January 2023, Section 5.3, Managing and Minimizing Subjectivity. https://database.ich.org/sites/default/files/ICH_Q9(R1)_Guideline_Step4_2023_0126.pdf
- U.S. Food and Drug Administration. “Warning Letter: Zydus Lifesciences Limited (722576).” 2 June 2026. FDA Warning Letter 722576 (2 June 2026)
- U.S. Department of Justice, Criminal Division. “Evaluation of Corporate Compliance Programs.” Updated September 2024. https://www.justice.gov/criminal/criminal-fraud/page/file/937501/dl
- Goddard K, Roudsari A, Wyatt JC. “Automation bias: a systematic review of frequency, effect mediators, and mitigators.” Journal of the American Medical Informatics Association, 2012;19(1):121-127. https://pubmed.ncbi.nlm.nih.gov/21685142/
- Parasuraman R, Manzey DH. “Complacency and bias in human use of automation: an attentional integration.” Human Factors, 2010;52(3):381-410. https://pubmed.ncbi.nlm.nih.gov/21077562/
- METR. “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.” 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Eloundou T, Manning S, Mishkin P, Rock D. “GPTs are GPTs: Labor market impact potential of LLMs.” Science, 2024;384(6702):1306-1308. https://pubmed.ncbi.nlm.nih.gov/38900883/
- World Economic Forum. “The Future of Jobs Report 2025.” January 2025. https://reports.weforum.org/docs/WEF_Future_of_Jobs_Report_2025.pdf
- European Commission. “First rules of the Artificial Intelligence Act are now applicable.” Shaping Europe’s Digital Future, 3 February 2025. https://digital-strategy.ec.europa.eu/en/news/first-rules-artificial-intelligence-act-are-now-applicable
- HMA and EMA Joint Big Data Steering Group. “Multi-annual Artificial Intelligence Workplan 2023-2028.” December 2023. EMA multi-annual AI workplan 2023-2028 (PDF)
- ISPE. “Workforce Preparedness and Organizational Readiness Take Center Stage at the 2026 ISPE AI in Life Sciences Summit.” Pharmaceutical Engineering, 13 May 2026. https://ispe.org/pharmaceutical-engineering/ispeak/workforce-preparedness-and-organizational-readiness-take-center
- ISPE. “AI in Pharma: Transforming Quality, Manufacturing, and Workforce Readiness.” Pharmaceutical Engineering. https://ispe.org/pharmaceutical-engineering/ispeak/ai-pharma-transforming-quality-manufacturing-and-workforce
- Lockovitch K, Bredin E, Meyer A, Reid L. “A Reinvented Role for Future Bank Chief Compliance Officers.” Oliver Wyman, June 2026. https://www.oliverwyman.com/our-expertise/insights/2026/jun/strategic-compliance-blueprint-for-banking-next-gen-cco.html








Your perspective matters—join the conversation.