In This Article
- Executive Summary
- Complaint Handling Is a Data Problem Wearing a Workflow Costume
- Intake Quality: The Minimum Structured Record
- The Dual Path: When One Contact Is Both a Quality Complaint and an Adverse Event
- Reporting Clocks and the Awareness Timestamp
- Investigation, Batch Linkage, and the Decision to Escalate
- Trending and Signal Detection: Rates, Not Counts
- Automation and AI in Coding and Triage
- A Modernization Sequence That Holds Up
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Complaint handling is the one routine channel where the outside world tells a pharmaceutical or biotech company that something may be wrong with a product that has already left the site. In most organizations it is run as a customer service function with a quality review bolted on the side. Calls are logged, cases are closed, cycle time is reported, and everyone treats the process as a workflow question. That framing is the root of nearly every complaint handling finding we see.
A complaint record is a data record before it is anything else. If intake does not capture a lot number, a structured defect category, and a defensible timestamp, then no amount of workflow tuning will make the record usable. It cannot be linked to a batch, it cannot be trended, it cannot support an investigation, and it cannot be reconciled against the safety database. The three failures that follow are predictable: unusable intake, a split path where a single contact containing both a product quality issue and an adverse event reaches only one process, and count-based trending that mistakes better reporting for worse product.
This article works through the five decisions that determine whether a complaint system produces signal or noise: what the minimum structured intake record must contain, how the complaint system and the pharmacovigilance system have to interoperate so a dual-content contact reaches both processes without diverging, how a complaint connects to the batch record and what should trigger a full investigation rather than a trend entry, how to normalize complaint rates against units distributed instead of counting cases, and where automation and AI assistance genuinely help in coding and triage along with the governance those tools now require under recent FDA and EU guidance.
Complaint Handling Is a Data Problem Wearing a Workflow Costume
Ask a quality leader to describe their complaint handling process and you will usually get a flow diagram. Intake, triage, acknowledgment, assessment, investigation, response to complainant, closure. Ask what the process produces and the answer is usually a set of cycle time metrics: average days to acknowledge, average days to close, percentage closed within target, backlog aging. Those numbers are real and they matter for resourcing. They tell you almost nothing about whether the product is behaving as it should.
The regulation has never described complaint handling as a service function. Under 21 CFR 211.198, written procedures for handling all written and oral complaints must exist and be followed, the quality control unit must review any complaint involving a possible failure of a drug product to meet its specifications, and a determination must be made as to whether an investigation under 211.192 is needed. The same section requires a review to determine whether the complaint represents a serious and unexpected adverse drug experience that must be reported to FDA.1 That single sentence is the dual-path obligation, written into the GMP rule rather than the safety rule, and it is the piece most often handled by a checkbox rather than a designed handshake.
The European position is framed the same way. EudraLex Volume 4 Chapter 8 covers complaints, quality defects and product recalls together, and the 2014 revision built quality risk management principles into how complaints are assessed, how quality defects are classified, and how recall decisions are made.2 ICH Q10 places complaint handling inside the pharmaceutical quality system as one of the inputs that feeds corrective and preventive action, process and product monitoring, and management review, which is to say it is expected to be a source of evidence about the state of control, not a customer relations queue.3
So the framing is settled in the regulation and unsettled in practice. Why? Because complaint handling touches the two functions least likely to share a system: commercial customer service, which owns the phone number and the inbox, and quality, which owns the consequence. Pharmacovigilance sits in a third place with its own database, its own vocabulary, and its own clocks. When three functions each own part of a record, the record fragments, and once it fragments the data quality problem becomes invisible because every function sees only its own portion and every portion looks complete.
The three failures that follow from bad framing
Every complaint handling remediation we have been asked to look at comes down to some combination of three problems, and all three are data problems.
Unusable intake. The record exists but cannot be used. There is a narrative and no lot number. There is a defect description in free text and no category. There is a date of entry and no date of receipt. Each of these is a small omission at the point of capture and a permanent limitation afterward, because you cannot go back and ask a caller from eleven months ago which lot they had.
Path divergence. A single contact contains a product quality issue and an adverse event. It is entered into one system, routed to one owner, and the other process never sees it. Or it is entered into both systems by two people from the same call notes, and from that moment the two records drift: different event dates, different seriousness assessment, different follow-up, different closure. At the next inspection someone asks for the reconciliation and there is no key to reconcile on.
Count-based trending. Complaints are counted, plotted, and compared month over month with no denominator. Distribution volume changes, a new market launches, a call center script is rewritten, an online reporting form goes live, and the count moves. The quality organization interprets a movement in reporting behavior as a movement in product performance, or misses a genuine change because a volume increase masked it.
The reframe. Stop asking whether the complaint workflow is efficient. Ask three data questions instead. Can every complaint record be joined to a batch record? Can every complaint record be joined to its safety counterpart when one exists? Can every complaint count be divided by a defensible denominator? If the answer to any of these is no, the workflow is not the problem.
Intake Quality: The Minimum Structured Record
Intake is where the value of the entire process is decided. Everything downstream is constrained by what the first person captured, and that person is usually the least trained, most time-pressured participant in the chain. They are often on a phone, often working from a script written by a commercial team, and often measured on call handling time.
A complaint recorded as “product did not work” is not a complaint record. It is a note. It cannot be trended because it does not map to a category. It cannot be linked to a batch because there is no lot number. It cannot support an investigation because there is nothing to compare against the batch record. And it cannot be assessed for safety reportability because nobody asked whether a patient took the product and what happened afterward.
The fields that have to be mandatory
21 CFR 211.198(b) sets a floor for the written record: name and strength of the drug product, lot number, name of complainant, nature of the complaint, and the reply to the complainant, along with the findings of the investigation and follow-up, or the documented reason no investigation was necessary and the name of the person who made that determination.1 That is the floor, not the design. A record that satisfies the floor and nothing else will still fail every analytical use you have for it.
| Field | Why it exists | What breaks without it |
|---|---|---|
| Lot or batch number | The join key to the batch record, distribution data, stability data, and open deviations | No batch linkage, no lot-level trending, no way to bound an investigation or a recall |
| Product, strength, presentation, pack size | Distinguishes a formulation issue from a packaging or device-component issue | Defect categories pool across presentations and hide presentation-specific patterns |
| Date of receipt (immutable) and date of event | Starts the regulatory awareness clock and separates event timing from reporting timing | Reporting clocks start late; time-to-event patterns cannot be constructed |
| Structured defect category and sub-category | The only basis on which complaints can be grouped and trended | Trending degrades into free-text search, which is neither reproducible nor complete |
| Complainant type and country or market | Separates health care professional reports from patient reports and from distributor reports | Channel effects are read as product effects; local regulatory obligations are missed |
| Any mention of patient exposure or harm | The trigger that starts the safety assessment path | Adverse events sit undetected inside quality complaint narratives |
| Sample availability and return status | Determines whether the complaint can be confirmed by testing | Investigations stall or close as unconfirmed without a documented reason |
| Source channel and original contact identifier | Enables deduplication and links the complaint to its safety counterpart | Duplicate cases inflate counts; dual-path reconciliation has nothing to match on |
| Storage and handling narrative | Separates a manufacturing defect from a distribution or in-use handling issue | Root cause work starts from the wrong hypothesis and often ends there |
Free text is evidence, not data
None of this argues against free text. The narrative is often the only place the real information lives, and a good narrative is what makes an investigation possible six months later. The point is that free text is evidence and structured fields are data, and the two do different jobs. A narrative supports a human reading one case. Structured fields support a machine comparing ten thousand.
The practical design rule is that every structured field must be derivable at the moment of capture and every narrative must be preserved verbatim. Do not let intake staff summarize the caller. Do not let a category selection replace the description. And do not let a system force a category before enough is known, because a forced early category is worse than a blank one: it is wrong and it looks confident.
Designing the controlled vocabulary
Most complaint categorization schemes fail because they mix two different questions into one list. “Broken tablet”, “wrong count”, “label illegible”, and “no effect” are not the same kind of statement. The first three describe an observable attribute of a physical unit. The last describes an outcome in a patient.
Separate the vocabulary into independent axes and code each one:
- What was observed. The physical or performance observation, described in the complainant’s terms and mapped to a controlled term.
- Where it was observed. The product component: drug product itself, primary container, closure, secondary packaging, labeling, insert, delivery component, shipper.
- When in the life of the unit. On receipt, during storage, at the point of preparation, during administration, after administration.
- Whether a patient was exposed. A yes, no, or unknown that is captured separately from the defect itself and never inferred from it.
With those four axes you can answer questions that a flat list cannot. Is the increase in “damaged product” complaints concentrated in one shipper configuration? Are labeling complaints coming from one market where a translated insert changed? Is a rise in delivery-component complaints happening only after a component supplier change? A single flat category list gives you a bar chart. Four axes give you a hypothesis.
The intake role must not carry the classification decision. The person taking the call captures what was said and selects from the controlled vocabulary. They do not decide whether the complaint is reportable, whether it is a confirmed defect, or whether an investigation is needed. Those are quality unit determinations under 211.198(a), and putting them in the intake script is how organizations end up with a documented decision made by someone who was never qualified to make it.
The Dual Path: When One Contact Is Both a Quality Complaint and an Adverse Event
This is the heart of the problem, and it is the part most complaint handling projects underestimate because it does not look like a technology issue until you try to solve it.
A single phone call arrives. A patient says the tablets in the last bottle looked different from usual, and that after taking them for three days she developed a rash. That contact contains a product quality complaint and a suspected adverse reaction. They are not the same thing, they are assessed by different people against different criteria, they carry different regulatory clocks, and they may or may not be causally related. What they share is one caller, one product, one lot, and one moment in time.
Now consider what has to happen. The quality organization needs a complaint record with the lot number, the appearance description, and a request for the returned sample. Pharmacovigilance needs an individual case safety report with an identifiable reporter, an identifiable patient, a suspect medicinal product, and a suspected adverse reaction, which are the four elements that make a case valid for reporting under GVP Module VI.4 The two records need to reference each other, follow-up needs to be coordinated so the patient is not called twice by two departments asking overlapping questions, and if the investigation finds a subpotent or mislabeled lot, that finding has to travel back to the safety assessment.
How the dual path actually fails
Single entry, single route
The contact is logged in one system by whoever answered. The intake person saw a product complaint and routed it to quality, or saw a rash and routed it to safety. The other process never learns the contact existed. This is the most common failure and the hardest to detect, because nothing looks wrong in either system.
Double entry, divergent records
Two people create two records from the same notes. Event dates differ by a day. One record captures the lot, the other does not. Follow-up is done separately and yields different answers. At reconciliation the two cases cannot be matched because no shared identifier was ever assigned.
Reconciliation as an annual event
The two systems are compared once a year for the periodic report. Gaps found then are twelve months old, the reporting clocks are long past, and the only available remediation is a retrospective explanation. Reconciliation done annually is an audit exercise, not a control.
One-way information flow
Safety tells quality about product complaints found in adverse event narratives, but quality never tells safety what the batch investigation concluded. A confirmed subpotency finding that would change the causality assessment on an efficacy failure case sits in a quality record that pharmacovigilance never reads.
The design that fixes it: one contact, two children
The pattern that holds up is straightforward to describe and takes real discipline to implement. Every inbound contact creates one canonical contact record with a unique contact identifier, an immutable receipt timestamp, and the verbatim narrative. That contact record is not a complaint and not a safety case. It is the record of the interaction.
From that contact record, the intake process creates child records: a product quality complaint if the contact contains any product quality content, a safety case if the contact contains any adverse event content, and both if it contains both. Every child record carries the parent contact identifier. That identifier is the reconciliation key, and it exists from the first second rather than being constructed later.
Capture once, into a shared contact record
One intake, one narrative, one immutable receipt timestamp, one contact identifier. The intake person answers two independent screening questions: does this contact describe anything about the physical product or its performance, and does this contact describe anything that happened to a person. Both can be yes.
Spawn the child records automatically
The screening answers drive record creation. A yes on the product question creates a complaint record. A yes on the person question creates a safety intake record. Neither creation is discretionary and neither can be suppressed by the intake role.
Carry the shared fields, not copies of them
Lot number, product, receipt timestamp, complainant identity, and country belong to the contact record. The child records reference them rather than holding independent copies that can be edited apart from each other.
Coordinate follow-up through one owner
One person owns contact with the complainant, works from a combined follow-up question set covering both quality and safety needs, and writes the answers back to the contact record. The complainant is not called twice.
Reconcile on a short cycle and treat gaps as events
Match complaint records to safety cases on the contact identifier at least monthly, and weekly for products under heightened monitoring. An unmatched record on either side is a quality event with an owner and a due date, not a line on a spreadsheet.
Close the loop back to safety
When a batch investigation confirms or excludes a product cause, that conclusion is written to the contact record and flagged to the safety case owner. A confirmed quality finding is information the causality assessment needs.
The boundary cases that decide how good your design is
Four situations separate a designed dual path from a diagram.
Lack of effect. A patient reports the medicine stopped working. Is that a quality complaint about potency, a safety case about therapeutic failure, both, or neither? Under most safety frameworks, lack of efficacy is reportable in defined circumstances, and under GMP it may indicate subpotency in a specific lot. The correct answer is that it enters both paths and each assesses it on its own criteria. Notably, ISPE excluded lack of effect from the total complaint rate metric in its quality metrics pilot, which tells you the industry has long recognized that this category behaves differently from physical defect complaints.5
An adverse event with no product quality content. A patient reports nausea. Nothing about the tablet, the pack, or the label. This is a safety case only. But the complaint system should still be able to see it, because a rise in a specific reaction concentrated in one lot is a quality signal even when no complainant ever mentioned the product’s appearance.
A quality defect with no patient. A pharmacist reports a cracked vial before dispensing. No patient exposure. This is a complaint only, and it may still trigger a Field Alert Report and a quality defect notification depending on what the investigation finds.
A medication error. A wrong dose was administered because two strengths of the same product have nearly identical cartons. This is simultaneously a labeling and packaging quality complaint, a potential safety case, and a design issue that belongs in a product risk file. It also sits explicitly within the scope of the ICH E2D framework for post-approval safety data, and the E2D(R1) revision published for consultation in 2024 addressed exactly this widening of data sources and definitions.6
The test for your dual path. Take twenty closed complaint records from the last quarter that contain any reference to a person, a symptom, or an outcome. For each one, ask whether a corresponding safety case exists, whether the two records agree on the lot and the dates, and whether either record references the other. If you cannot run that test in an afternoon because there is no shared key, you have found the real finding before an inspector does.
Reporting Clocks and the Awareness Timestamp
The dual path also carries two sets of regulatory clocks that start from different definitions of the same moment, and this is where well-run organizations still get caught.
For a US applicant holding an NDA or ANDA, a Field Alert Report is due within three working days of becoming aware of information concerning a significant quality problem with a distributed drug product, including any incident that causes a drug product or its labeling to be mistaken for another article, and any bacteriological contamination or significant chemical, physical, or other change or deterioration in a distributed product, or any failure of a distributed batch to meet its specification.7 FDA maintains dedicated guidance and submission channels for FARs, and the operative word in the requirement is awareness.8
For a serious and unexpected adverse drug reaction, ICH E2D sets the expedited reporting expectation at 15 calendar days from the point at which the marketing authorization holder receives the minimum information needed for a valid case.9 GVP Module VI carries the equivalent obligations in the EU, including electronic submission of individual case safety reports to EudraVigilance and the validation and follow-up expectations that surround them.4
On the quality defect side, a marketing or manufacturing authorization holder must notify EMA of any product quality defect, including a suspected defect, of a centrally authorized medicine that could result in a recall or an abnormal restriction on supply, and national competent authorities operate a rapid alert system for defects that present a serious risk to public health.10 EMA publishes both the risk-based approach to classifying suspected quality defect reports and its own analysis of quality product defects arising in the centralized procedure, which are two of the more useful public documents for calibrating what regulators consider a serious defect.1112
Awareness starts at first contact, not at classification. The single most common clock failure is treating the reporting clock as starting when the quality unit categorized the complaint, rather than when the organization first received the information. If a call arrives on Monday, sits in a customer service queue for four days, and reaches quality on Friday, the three working day Field Alert Report window is already spent. The immutable receipt timestamp on the contact record is not administrative tidiness. It is the evidence that determines whether a report was on time.
Investigation, Batch Linkage, and the Decision to Escalate
Once a complaint is captured properly, the next question is what to do with it, and the regulation is unusually direct here. Under 211.198(a), the quality control unit must review any complaint involving the possible failure of a drug product to meet any of its specifications and determine whether an investigation under 211.192 is needed. Where an investigation is not conducted, the record must include the reason it was found unnecessary and the name of the person who decided.1
That last clause is quietly one of the most inspected sentences in Part 211. It converts a decision not to act into a documented, attributable determination. Inspectors read those determinations in bulk, and thin reasoning shows up quickly when a hundred records carry the same three-word justification.
What batch linkage actually requires
The lot number is the join key, and everything useful about a complaint investigation depends on it resolving to something. When it does, the investigation can pull:
- The executed batch record, including any in-process results near a limit and any manual interventions
- Release testing results and the certificate of analysis for that lot
- Open and closed deviations associated with that batch, that line, and that time window
- Stability data for the lot or its representative, including any out-of-trend results
- Retained samples, and whether a retain can be examined for the reported attribute
- The complaint history for the same lot, adjacent lots, and the same product on the same line
- Distribution records showing where the lot went, in what quantity, and through which channels
- Component and material lot genealogy, particularly for containers, closures, and delivery components
When the lot number is missing, none of that is available and the investigation becomes an exercise in describing the complaint rather than explaining it. This is why lot capture belongs in the intake design conversation and not in a data cleanup project two years later.
Investigation or trend entry: making the trigger explicit
Not every complaint needs a full investigation, and pretending otherwise produces a backlog that damages the quality of the investigations that genuinely matter. The decision needs to be a written rule with named owners, not a judgment made case by case under time pressure.
| Situation | Default disposition | What escalates it |
|---|---|---|
| Single cosmetic observation, no specification implication, no patient exposure | Trend entry with documented rationale for no investigation | A second occurrence in the same lot, or any occurrence in a lot with an open related deviation |
| Complaint alleging a possible specification failure (potency, appearance outside limits, contamination, wrong content) | Full investigation under 211.192 | Already at the highest tier; assess for Field Alert Report immediately |
| Labeling or packaging complaint with mix-up potential | Full investigation plus immediate reportability assessment | Any indication the product or labeling could be mistaken for another article triggers the FAR clock |
| Complaint accompanied by any reported patient harm | Full investigation plus safety case | Serious and unexpected reaction accelerates the safety clock independently of the quality timeline |
| Repeated same-category complaints across multiple lots | Trend investigation at the product or process level, not lot by lot | Rate increase that survives denominator normalization and reporting-behavior checks |
| Complaint where no sample can be returned | Investigation proceeds on records and history; unconfirmed is a finding, not a closure reason | Pattern of unreturned samples in one channel is itself worth investigating |
Isolated event or the visible part of a pattern
The most consequential judgment in complaint handling is whether one complaint is an isolated event or the part of a pattern you can see. You cannot answer that by looking at the complaint. You answer it by defining comparison sets in advance and checking the complaint against all of them.
Useful comparison sets are: the same lot, adjacent lots from the same campaign, the same product and presentation across a rolling window, the same manufacturing line regardless of product, the same component or material supplier lot, the same market or distribution channel, and the same defect category across the whole portfolio. A complaint that is unremarkable in five of those sets and a clear outlier in one has told you where to look.
FDA warning letters make this point repeatedly and specifically. Investigations that fail to be adequately expanded to include other potentially affected batches or products are a recurring citation, as is failure to extend a complaint investigation to other products handled on the same equipment.1314 Letters have also cited firms for failing to establish and follow adequate written complaint procedures and for not investigating a complaint that included an adverse reaction.15 The pattern across those citations is not that companies lacked a procedure. It is that the scope of the investigation stopped at the lot the complainant happened to have.
What a well-scoped complaint investigation contains. A statement of what was alleged in the complainant’s words. A statement of what was confirmed, with the evidence. The comparison sets checked and the result for each. The batch record and testing evidence reviewed. An explicit conclusion on product cause: confirmed, excluded, or unresolved with a stated reason. A defined scope of affected material with the rationale for the boundary. The reportability decisions made and their dates. And the CAPA or trend disposition with an owner.
Trending and Signal Detection: Rates, Not Counts
Most complaint trending we see is a count of complaints per month by category, plotted as a bar chart, reviewed in a meeting, and filed. It is nearly useless for detecting change, and it is actively misleading when volume moves.
Normalize against what was distributed
A complaint rate is complaints divided by exposure. The FDA draft guidance on submission of quality metrics data described a product quality complaint rate as an indicator of product quality with respect to label claims and patient perception, calculated against distribution rather than reported as a raw number.16 The ISPE quality metrics pilot used total complaint rate excluding lack of effect, expressed per million packs.5 Both are telling you the same thing: the numerator alone is not a metric.
Choosing the denominator is a real decision and it depends on the product:
- Packs or units distributed works for solid oral products with a stable pack configuration and is the most common basis.
- Doses or vials is better where pack size varies widely across markets, because a per-pack rate will otherwise move whenever the market mix moves.
- Patient-days of exposure is the right basis for chronic therapies where a single pack covers very different treatment durations across strengths.
- Administrations is appropriate for products with a delivery component, where the complaint relates to use rather than to the drug substance.
Whatever you pick, three problems will bite. Distribution data lags, so the denominator for the most recent month is provisional and the rate will restate. Channel inventory means product distributed is not product used, and a launch ramp or a channel fill will suppress the apparent rate for a period. And free or sample distribution is frequently absent from the commercial data set entirely, which understates exposure and overstates the rate. None of these are reasons to skip normalization. They are reasons to document the denominator definition, hold it constant, and annotate the chart when it changes.
An increase in complaints often means better reporting
This is the point that separates a quality organization that understands its own data from one that reacts to it. A rise in the complaint rate has at least two families of explanation, and the reporting explanation is more common than the product explanation.
The pharmacovigilance literature has documented this carefully because spontaneous reporting systems are subject to exactly the same distortions. FDA safety alerts measurably increase the volume of adverse event reports submitted for the products named, an effect described as stimulated reporting.17 Notoriety bias raises reporting for specific event and product combinations after media or regulatory attention. The Weber effect describes reports rising through the first years after launch, peaking, and then declining even as prescriptions keep increasing. Analysis of spontaneous reporting in Japan documented similar structural biases in national reporting data.18 None of these effects have anything to do with the product changing.
Product quality complaints are subject to the same distortions plus several of their own.
| Observed change | Product explanation | Reporting explanation | How to distinguish |
|---|---|---|---|
| Complaint rate rises across all categories at once | Rare; would require a systemic change | New reporting channel, revised script, easier web form, market expansion, an acquired portfolio joining the system | Check whether the rise is proportional across unrelated categories and whether it starts on a system change date |
| Rise concentrated in one category and one lot range | Likely a real process or material change | Unlikely | Check batch records and component genealogy for the lot range; check adjacent lots |
| Rise concentrated in one market | Possible local packaging, transport, or storage issue | Local regulatory campaign, a local reporting requirement change, a distributor policy change | Compare the same lots distributed to other markets |
| Rate falls sharply with no intervention | Rare | A channel stopped forwarding complaints; a system migration dropped records; a contract partner changed process | Reconcile channel-level counts against prior periods and against the partner’s own records |
| Rise in a single defect category shortly after a public safety communication | Possible but must be tested | Stimulated reporting and notoriety bias | Check whether the rise is confined to the reported event, whether new reports are older events, and whether the reporter mix shifted |
The rule worth writing into the procedure. Any change in complaint rate must be tested against a reporting-behavior explanation before a product explanation is accepted, and the test must be documented. This does not delay action on a genuine defect, because a defect signal will normally be concentrated by lot, category, or line, and a reporting artifact will normally be broad. What it prevents is a CAPA opened against a manufacturing process because a new online reporting form went live.
Stratification and thresholds
Trending at the portfolio level detects almost nothing. Real signals appear when you stratify: by lot, by presentation, by manufacturing line, by market, by channel, by defect category, and by the four intake axes described earlier. The practical constraint is statistical, not technical. Slice too finely and every cell contains one or two complaints, at which point normal variation looks like a signal.
Set thresholds deliberately and document who set them. Fixed thresholds are simple and defensible but do not adjust for volume. Statistical rules on a control chart of the normalized rate handle volume properly but need enough history to establish a baseline and need a documented rationale for the rule set. For most portfolios a two-tier approach works: a low fixed threshold that triggers a review at any occurrence for high-severity categories such as contamination or mix-up, and statistical rules on normalized rates for everything else.
Whatever the mechanism, the output has to reach the pharmaceutical quality system rather than a dashboard nobody opens. ICH Q10 expects complaint data to inform process and product monitoring, CAPA, and management review, which means a complaint trend that crosses a threshold should generate a quality event with an owner and appear in the periodic product quality review and the management review pack.3
Automation and AI in Coding and Triage
Complaint intake and safety case processing are among the most credible near-term applications of AI assistance in a regulated life sciences operation, and also among the easiest to govern badly. Both things are true and the balance matters.
Where it genuinely helps
The work that consumes capacity in complaint and case processing is high volume, repetitive, language-heavy, and mostly not judgment. That is a good match.
- Structuring free text. Extracting candidate values for lot number, product, strength, dates, and reported symptoms from a narrative or an email and presenting them as suggestions for confirmation. This is the single highest-value application because it attacks the intake quality problem directly.
- Duplicate detection. Identifying that a case arriving through a distributor and a case arriving through a call center describe the same event. Duplicates inflate counts, distort rates, and are tedious to find manually.
- Coding suggestion. Proposing MedDRA terms for safety cases and controlled defect categories for complaints, with the human coder accepting, adjusting, or rejecting. Consistency of coding is what makes trending possible, and consistency is exactly what varies between human coders.
- Translation and transcription. Making non-English complaints and voice recordings available to reviewers without a delay that eats the reporting clock.
- Triage prioritization. Surfacing contacts that contain probable seriousness indicators so they reach a qualified assessor first, without making the seriousness determination.
- Reconciliation gap finding. Matching complaint records to safety cases where the shared key is imperfect, and flagging probable orphans for human review.
- Anomaly flagging in trends. Detecting rate changes across many strata simultaneously, which is genuinely hard to do manually across a large portfolio.
IQVIA’s assessment of how AI is reshaping pharmacovigilance under emerging regulatory guidance describes much the same distribution of value: material gains in case intake, coding, and duplicate handling, with regulators focused on validation, transparency, and oversight rather than on prohibition.19
Where it does not belong, at least not yet
There is a clean line and it is worth stating plainly. A model may propose. A qualified person decides, and the decision record must be separate from the model output.
The determinations that stay with people are the ones the regulation attributes to a named individual or a named function: the seriousness and expectedness assessment on a safety case, the reportability decision and its clock, the quality unit determination on whether an investigation is required under 211.192, the documented reason an investigation was not conducted, the conclusion of a batch investigation, and the scope of affected material. Each of those is a decision the regulator will ask a person to explain.
The governance this now requires
The governance question has moved from theoretical to concrete in the last eighteen months, and there are two documents any pharma or biotech organization introducing AI assistance into complaint or safety processing should have read.
FDA published draft guidance on 6 January 2025 titled “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products”, covering nonclinical, clinical, post-marketing, and manufacturing uses where an AI model produces information supporting a regulatory decision about safety, effectiveness, or quality.20 Its central construct is a risk-based credibility assessment framework: define the question the model addresses, define the context of use, assess model risk based on how much the model influences the decision and how consequential that decision is, then scale the evidence you generate to that risk. For complaint handling this is directly usable. A model suggesting a defect category that a coder confirms carries low model influence. A model that filters which contacts a human ever sees carries very high model influence, because the human cannot correct what never reaches them.
On the manufacturing side, the EU published draft Annex 22 on artificial intelligence on 7 July 2025 as part of a wider digital update to EudraLex Volume 4, with the consultation closing in October 2025.21 The draft sets expectations for intended use, validation, lifecycle management, explainability, and human oversight for AI and machine learning models embedded in critical GMP applications. It also took an initial position that dynamic models which continuously adapt during use, models returning different outputs for identical inputs, and generative AI and large language models should not be used in critical GMP applications. Following consultation feedback the drafting group has been considering whether and under what risk-based safeguards such models might be addressed, and EMA has run multistakeholder work on the guidance development.22 Published interpretation of the draft in the PDA Journal reads it as an attempt to bridge existing computerized systems validation practice and the specific behavior of learning models.23
The practical implication for complaint handling is not that language models are off limits. It is that the more a model touches a critical GMP decision, the more constrained the acceptable design becomes, and the safest architecture keeps the model on the suggestion side of a human decision boundary that is documented and auditable.
Minimum governance for AI assistance in complaint and case processing
- Write a context of use statement per model: what it does, what it does not do, which decision it supports, and what happens when it is wrong in each direction.
- Classify model risk on both influence and consequence, and scale validation evidence to that classification rather than applying one standard to everything.
- Record the model version and the model output alongside every human decision, so a later review can distinguish what the model proposed from what the person decided.
- Monitor acceptance and override rates by category and by reviewer. A category with a very high acceptance rate deserves back-testing; a very low one means the model is adding work.
- Keep a labeled evaluation set that is refreshed as products, presentations, and vocabularies change, and re-test after every model or prompt change.
- Never let a model suppress a record. Assistance may reorder a queue; it may not remove something from it.
- Do not let automation touch the denominator. Distribution data feeding your complaint rate should come from a controlled source, not from an inferred one.
Plan for the number to go up. If automation improves extraction from free text and duplicate detection is tightened, more contacts will be correctly identified as complaints and more will be correctly linked to a lot. Your complaint rate will rise, possibly sharply, and none of that rise reflects a change in product quality. Decide before deployment how you will explain this in the periodic product quality review, in management review, and to an inspector, and annotate the trend at the deployment date. Organizations that skip this step end up either opening CAPAs against a phantom quality problem or, worse, quietly tuning the tool down.
A Modernization Sequence That Holds Up
Most complaint handling modernization programs start with a system selection and end with a system that carries the same broken records into a better interface. The sequence below is deliberately the reverse: fix what is captured and how it links before changing where it lives.
Rebuild the intake record and the controlled vocabulary
Make lot number, receipt timestamp, patient exposure indicator, and structured defect coding mandatory. Split the vocabulary into observation, component, timing, and exposure axes. Preserve the verbatim narrative alongside every coded field. This work is procedural and training-driven, and it delivers value before any system changes.
Establish the single contact record and the dual-path handshake
One contact identifier, two child records, mandatory cross-reference, a combined follow-up question set, and reconciliation on a monthly or shorter cycle with unmatched records treated as quality events. This is the change that most reduces regulatory exposure.
Make the lot number resolve to a batch record
A captured lot number that cannot be joined to the batch record, distribution data, and stability data is only half useful. Fix the reference data, the format variations across markets, and the partner data feeds before building any analytics.
Define and validate denominators before building dashboards
Pick the exposure basis per product family, document it, and agree how lagging distribution data will be handled in the most recent periods. A dashboard built on an undefined denominator will be rebuilt within a year.
Write the escalation logic as an explicit decision table
Investigation versus trend entry, comparison sets to be checked, reportability triggers, and named owners for each determination. Replace case-by-case judgment under time pressure with a rule that a trained assessor applies and a reviewer can audit.
Introduce assistance where it is reversible, under governance
Start with extraction suggestions, duplicate detection, and coding proposals, all with human confirmation. Write the context of use, classify model risk, log versions with decisions, and monitor override rates. Expand only where the evidence supports it.
Questions worth asking before any of this
A short diagnostic will tell you where you actually are, and it does not need a project to run.
- What percentage of complaint records from the last twelve months contain a lot number that resolves to a batch record?
- What is the median gap between the receipt timestamp on the contact record and the timestamp on the quality record? What is the 95th percentile?
- Of complaints that reference a person or a symptom, what proportion have a matching safety case, and what key was used to match them?
- What denominator does the current complaint rate use, who defined it, and when was it last restated?
- How many complaint records in the last year carry a documented reason for not investigating, and how many distinct reasons appear across them?
- When the complaint rate last moved materially, what was the documented test that separated a reporting explanation from a product explanation?
- Which complaint trends appeared in the last two management reviews, and what decisions did they produce?
An organization that can answer all seven quickly has a functioning system regardless of what software it runs on. An organization that cannot answer the first one has a data problem that no system selection will fix.
Conclusion
Complaint handling deserves more attention than it gets, and it deserves a different kind of attention than it usually gets. The instinct to treat it as a workflow to be accelerated is understandable, because backlog and cycle time are visible and uncomfortable. But the value of the process is not in how fast records close. It is in whether the records, taken together, tell you something true about how your product is performing outside your control. That is a data quality question and a signal detection question, and both are answered at intake.
The three changes that matter most are unglamorous. Capture a structured record that can be joined to a batch and to a safety case. Design the dual path so a single contact containing both a quality issue and an adverse event reaches both processes from one shared record with one shared key and one shared clock. Trend normalized rates rather than counts, and test every movement against a reporting-behavior explanation before accepting a product explanation. Automation and AI assistance help meaningfully with the first of these and can help with the third, but only inside a governance frame that keeps the regulated determinations with named people and documents where the model stopped and the person started.
Sakara Digital works with pharma and biotech organizations building this kind of complaint and quality data capability, particularly where the complaint system, the safety database, and the batch record have grown up separately and now need to speak to each other. If you are looking at a complaint handling modernization and want an independent perspective on where to start, we are happy to have that conversation.
For Further Reading
For Further Reading
- Quality Event Management Modernization: A Framework
- Pharmacovigilance Signal Detection with AI: What Regulators Expect
- Deviation Trending Analytics: From Excel to Real-Time Dashboards
- CAPA Analytics: Moving From Reactive to Predictive
- Annex 22 Mock Inspection: What a Pharma Quality Team Should Practice Now
- The Manufacturing Data Quality Scorecard: KPIs Beyond Regulatory Submissions
References & Sources
- US Food and Drug Administration. “21 CFR 211.198: Complaint files.” Electronic Code of Federal Regulations, current edition. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-J/section-211.198
- European Commission. “EudraLex Volume 4, EU Guidelines for Good Manufacturing Practice, Chapter 8: Complaints, Quality Defects and Product Recalls.” August 2014. https://health.ec.europa.eu/system/files/2016-11/2014-08_gmp_chap8_0.pdf
- International Council for Harmonisation. “ICH Q10: Pharmaceutical Quality System.” ICH Harmonised Tripartite Guideline. https://database.ich.org/sites/default/files/Q10 Guideline.pdf
- European Medicines Agency. “Guideline on good pharmacovigilance practices (GVP) Module VI: Collection, management and submission of reports of suspected adverse reactions to medicinal products (Rev 2).” EMA/873138/2011 Rev 2. https://www.ema.europa.eu/en/documents/regulatory-procedural-guideline/guideline-good-pharmacovigilance-practices-gvp-module-vi-rev-2_en.pdf
- ISPE. “Quality Metrics Initiative: Quality Metrics Pilot Program Wave 2.” June 2016. https://ispe.org/sites/default/files/regulatory/2023/QMWAVE2DL.pdf
- Federal Register. “E2D(R1) Post-Approval Safety Data: Definitions and Standards for Management and Reporting of Individual Case Safety Reports; International Council for Harmonisation; Draft Guidance for Industry; Availability.” 14 March 2024. https://www.federalregister.gov/documents/2024/03/14/2024-05381/e2dr1-post-approval-safety-data-definitions-and-standards-for-management-and-reporting-of-individual
- US Food and Drug Administration. “21 CFR 314.81: Other postmarketing reports.” Electronic Code of Federal Regulations, current edition. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-D/part-314/subpart-B/section-314.81
- US Food and Drug Administration. “Field Alert Reports.” Surveillance and Post Drug Approval Activities. https://www.fda.gov/drugs/surveillance-post-drug-approval-activities/field-alert-reports
- International Council for Harmonisation. “ICH E2D: Post-Approval Safety Data Management: Definitions and Standards for Expedited Reporting.” November 2003. https://database.ich.org/sites/default/files/E2D_Guideline.pdf
- European Medicines Agency. “Quality defects and recalls.” Compliance, post-authorisation. https://www.ema.europa.eu/en/human-regulatory-overview/post-authorisation/compliance-post-authorisation/quality-defects-recalls
- European Medicines Agency. “Management and classification of reports of suspected quality defects for medicinal products: risk-based decision making.” https://www.ema.europa.eu/en/documents/regulatory-procedural-guideline/management-classification-reports-suspected-quality-defects-medicinal-products-risk-based-decision-making_en.pdf
- European Medicines Agency. “An analysis of quality product defects in the centralised procedure.” https://www.ema.europa.eu/en/documents/regulatory-procedural-guideline/analysis-quality-product-defects-centralised-procedure_en.pdf
- US Food and Drug Administration. “Warning Letter: Catalent Indiana LLC, MARCS-CMS 718189.” 20 November 2025. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/catalent-indiana-llc-718189-11202025
- US Food and Drug Administration. “Warning Letter: Brands International Corporation, MARCS-CMS 689983.” 17 December 2024. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/brands-international-corporation-689983-12172024
- US Food and Drug Administration. “Warning Letter: Bi Coastal Pharma International, MARCS-CMS 628196.” 30 June 2022. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/bi-coastal-pharma-international-628196-06302022
- US Food and Drug Administration. “Submission of Quality Metrics Data: Guidance for Industry (Revised Draft).” https://www.fda.gov/media/101522/download
- Hoffman KB, Demakas AR, Dimbil M, Tatonetti NP, Erdman CB. “Stimulated Reporting: The Impact of US Food and Drug Administration-Issued Alerts on the Adverse Event Reporting System (FAERS).” Drug Safety, 2014. https://pmc.ncbi.nlm.nih.gov/articles/PMC4206770/
- Hasegawa S, et al. “Bias in Spontaneous Reporting of Adverse Drug Reactions in Japan.” PLOS ONE, 2015. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0126413
- IQVIA. “How AI is reshaping pharmacovigilance through regulatory guidance.” September 2025. https://www.iqvia.com/blogs/2025/09/how-ai-is-reshaping-pharmacovigilance
- Federal Register. “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products; Draft Guidance for Industry; Availability; Comment Request.” 7 January 2025. https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug
- ECA Academy. “EU GMP Annex 22 (Draft 2025): Artificial Intelligence.” GMP Guideline summary. https://www.gmp-compliance.org/guidelines/gmp-guideline/eu-gmp-annex-22-draft-2025-artificial-intelligence
- European Medicines Agency. “Good manufacturing practice: Multistakeholder workshop on expert contributions to artificial intelligence guidance development (Annex 22).” https://www.ema.europa.eu/en/events/good-manufacturing-practice-multistakeholder-workshop-expert-contributions-artificial-intelligence-guidance-development-annex-22
- PDA Journal of Pharmaceutical Science and Technology. “Bridging Guidance and Regulation: Interpreting the Draft Annex 22 on Artificial Intelligence in GMP Manufacturing.” https://journal.pda.org/content/early/2026/02/14/pdajpst.2025-000076.1








Your perspective matters—join the conversation.