Two Vocabularies, Two Different Jobs

Most conversations about AI readiness in a regulated environment go wrong in the same place. Someone asks whether the data is good enough, and the quality organization answers with its data integrity record. Those are different questions, and answering the second with the first is how a program reaches its pilot phase before anyone notices the problem.

ALCOA+ grew out of regulatory enforcement. Its purpose is evidentiary. When an investigator asks whether a result is real, ALCOA+ is the framework that lets you demonstrate the record is what it claims to be, was created when it claims, by who it claims, and has not been altered without a trace. The MHRA, PIC/S, and the World Health Organization have all published guidance built around these attributes, and FDA covers the same ground for drug CGMP in its data integrity questions and answers.5694

FAIR came from somewhere else entirely. It was published in 2016 in Scientific Data by a group of researchers, funders, and publishers trying to solve a different problem: enormous quantities of research data existed, and almost none of it could be found or reused by anyone who had not generated it.1 The paper set out four principles aimed explicitly at making data usable by machines, not only by people.

That last part is why FAIR matters for AI. The original paper puts machine actionability at the center. The authors were not writing about compliance at all. They were describing the conditions under which a computer can locate a dataset, understand what is in it, combine it with another dataset, and do something useful with the result.

That is precisely what an AI project needs from a quality system, and it is not what ALCOA+ was designed to deliver.

The distinction in one line

ALCOA+ tells you whether one record can be trusted. FAIR tells you whether ten thousand records can be used together.

What ALCOA+ Was Built to Answer

It is worth being precise about what ALCOA+ does cover, because the argument here is not that it is insufficient as data integrity guidance. It is excellent data integrity guidance. The argument is that data integrity and fitness for machine use are different properties, and an organization can hold one without the other.

The MHRA guidance defines the attributes and, importantly, frames them around the lifecycle of a record and the organizational culture that produces it.5 PIC/S PI 041-1 goes further into the governance and system controls that make the attributes real rather than aspirational.6 Between them they describe a mature discipline.

Here is what that discipline guarantees, and what it does not.

ALCOA+ attributeWhat it guaranteesWhat it does not guarantee
AttributableYou know who created or changed the recordThat the field they filled in means the same thing as the equivalent field at another site
LegibleThe record can be read and remains readableThat it can be parsed by a system rather than read by a person
ContemporaneousIt was recorded at the time of the activityThat the timestamp format is consistent across the systems you want to combine
OriginalIt is the first capture or a certified copyThat the original is retrievable from a system you still operate
AccurateIt reflects what happenedThat it is specific enough to be useful. “Operator error” can be entirely accurate and carry no information
CompleteNothing was deleted, including repeat analysesThat the fields you need for analysis were ever captured in the first place
ConsistentThe sequence of events is intact and datedThat categories and vocabularies are consistent between sites or systems
EnduringIt survives its retention periodThat it survives in a form anything can still open and process
AvailableIt can be retrieved for review or inspectionThat it can be retrieved in bulk, on demand, by a system

Read the right-hand column as a list. Every item on it is a reason an AI project stalls, and not one of them is a data integrity failure.

What FAIR Adds, Principle by Principle

The four principles are short, and they translate into quality system terms without much strain.12

FINDABLE

Can It Be Located Without Already Knowing Where It Is?

Data has a persistent identifier and enough descriptive metadata to be located on its properties. In a document system, this is the difference between an SOP you can find by product, process, and owner, and an SOP you can only find if you already know its document number.

ACCESSIBLE

Can It Be Retrieved Through a Defined Route?

Accessible does not mean open. Controlled access with clear authentication satisfies this principle completely. What fails it is data that technically exists but cannot practically be retrieved, such as a scanned document in a folder from a system retired years ago.

INTEROPERABLE

Can It Be Combined With Data From Elsewhere?

The data uses shared vocabularies and formats. This is where quality systems struggle most. If one site classifies a deviation as “Human Error, Procedural” and another classifies the same event as “Training Related,” both records are accurate, both are compliant, and neither can be counted with the other.

REUSABLE

Can Someone Who Did Not Create It Use It Correctly?

The data carries enough context, provenance, and usage information to be interpreted by a stranger. A CAPA closed with the root cause recorded as “operator retrained” is not reusable. It records what was done and nothing about why the failure happened.

Notice that none of these four is a trust question. FAIR assumes the data is honest. It asks whether the data is usable. That assumption is exactly why the two frameworks fit together rather than competing: ALCOA+ earns the assumption FAIR makes.

Where the Two Overlap and Where They Diverge

There is genuine overlap, and an organization that has done serious data integrity work has already built part of the foundation.

ALCOA+ requires that records be available; FAIR requires that they be accessible. ALCOA+ requires attributability; FAIR’s reusability principle requires provenance. Both care that a record endures. If you have invested in data integrity remediation over the past decade, you are not starting from nothing.

The divergence is concentrated in the middle two FAIR principles, and it is worth stating plainly what is absent from ALCOA+:

  • Nothing in ALCOA+ requires that deviation categories mean the same thing across sites.
  • Nothing in ALCOA+ requires that a root cause be captured in a structured field rather than a paragraph.
  • Nothing in ALCOA+ requires that document metadata be complete enough to find a record without knowing its number.
  • Nothing in ALCOA+ requires that two systems holding related data use compatible identifiers for the same product, batch, or piece of equipment.

These are not compliance failures. A site can pass inspection with every one of them present. They are usability failures, and they are the reason AI projects stop after the proof of concept.

Why this matters for how you scope the work

If you frame data readiness as a data integrity problem, the work goes to the people who already own data integrity, and they will correctly report that the data is compliant. If you frame it as an interoperability and reusability problem, the work goes to the people who own the definitions: the categories, the controlled vocabularies, the metadata standards. Those are different people and a different project.

What the FY2025 Inspection Record Actually Shows

It is worth testing this argument against what regulators actually cite, because the answer is more useful than the industry shorthand suggests.

FDA publishes an annual inspection observations dataset counting Form 483 observations by citation.3 In fiscal year 2025, covering inspections that ended between October 2024 and September 2025, FDA recorded 2,837 drug CGMP observations across 316 cited provisions, drawn from 713 drug Form 483s.

243 Observations against 21 CFR 211.22(d), the most cited provision of the year: quality control unit procedures not in writing, or not fully followed
309 Observations against section 211.22 as a whole, roughly 11 percent of every drug observation recorded that year
153 Observations against 21 CFR 211.68, the computerized systems provision, ranking seventh by section

The top of the list is worth reading slowly. The most cited provision is about whether quality unit procedures exist in writing and are followed. Second, at 164, is 21 CFR 211.192, failure to thoroughly review discrepancies. Third, at 162, is 211.100(a), absence of written procedures. Fourth, at 121, is 211.160(b), laboratory controls that are not scientifically sound.

The pattern is that the most common findings in pharmaceutical manufacturing are not about equipment or chemistry. They are about whether procedures exist, whether they are followed, and whether the record demonstrates it.

A correction worth making before you quote this data

FDA’s citation database contains no citation for “data integrity,” none for “audit trail,” and none for “quality risk management.” Those phrases do not appear in the reference text at all. Data integrity is a concept enforced through specific CGMP provisions rather than a regulation you can be cited against directly, and ICH Q9 is guidance rather than codified United States regulation.

The closest provision is 21 CFR 211.68, covering computerized systems. Within it, 211.68(b) accounted for 110 observations, and its largest single line item, at 87, reads that appropriate controls are not exercised over computers or related systems to assure that changes to records are instituted only by authorized personnel. That is the audit trail concept. FDA does not use the phrase.

If you put inspection data in a business case, cite the provision and the count. The specifics are more persuasive than the shorthand, and they survive a challenge from someone who knows the dataset.

Now connect this back to the two frameworks. A finding that procedures are not consistently followed is a finding about the record. A finding that discrepancy investigations were not thorough is a finding about what the investigation record contains. These are the same records an AI program would train on, and the inspection record is telling you their condition.

Five Places the Gap Appears in a Quality System

1. Deviation and CAPA Classification

This is the most common and the most expensive. Categories drift over time and diverge across sites. Free text absorbs the information that should have been structured. The result is a dataset where every individual record is defensible and the aggregate cannot support trending, prediction, or any model that depends on consistent labels.

The specific failure is subtle. It is not that the data is missing. It is that two records describing the same physical event carry different labels, so any count of that event type is wrong, and no amount of model sophistication corrects a label problem.

2. Document and SOP Metadata

Orphan documents with no current owner. Review dates that passed without action. Documents referencing products, equipment, or sites that no longer exist. None of this is visible from inside any single document, which is why it persists for years. It becomes visible the moment you try to build anything that reasons across the document set.

3. Training Records and Qualification State

In FY2025, 21 CFR 211.25(a), covering training, education, and experience, drew 67 observations.3 Training data has the same structural problem as deviation data: individual records are complete, but the curriculum assignments behind them are inconsistent, so the organization cannot reliably answer who is currently qualified for a given task across sites.

4. Equipment and Product Master Data

The same physical asset carries different identifiers in the maintenance system, the manufacturing execution system, and the quality system. Each is internally consistent. Joining them requires a mapping that usually lives in one person’s spreadsheet. This is a pure interoperability failure and it blocks nearly every cross-system analysis anyone wants to run.

5. Retired System Archives

Data retained for compliance in a format nobody can process: scanned documents, exports from a system that no longer runs, or a database whose schema was never documented. It satisfies the ALCOA+ requirement to be enduring and available. It fails accessibility in the FAIR sense completely, and it is usually the dataset with the longest history and therefore the most analytical value.

The useful reframe

Each of these five has a bounded, describable fix. That is the practical advantage of naming the problem as a FAIR problem rather than as a vague concern about data readiness. Vague concerns do not get funded. “Deviation categories are inconsistent across three sites, here is the count, here is the harmonization plan” does.

The Six Dimensions That Make “Clean” Measurable

ALCOA+ and FAIR both describe properties. Neither gives you a number. For the assessment described later in this article you need something countable, and the standard data quality dimensions supply it. They are not a regulatory framework and they do not replace either of the two above. They are the measurement layer that makes an argument about data condition concrete.

Six dimensions carry almost all the useful signal in a quality system.

DimensionThe question it answersWhat to count in a QMS
AccuracyDoes the value reflect reality?Records where the recorded classification disagrees with the narrative description of the same event
CompletenessAre required values present?Deviation records with an empty or default root cause field; documents with no owner
ConsistencyDoes the same thing carry the same value everywhere?Distinct category values in use where one controlled list was intended; equipment identifiers that differ between systems
TimelinessIs the value current?Documents past their review date; training assignments referencing superseded procedure versions
ValidityDoes the value conform to its defined format or list?Free text in a field that should hold a controlled value; dates outside a plausible range
UniquenessIs anything recorded twice?Duplicate deviation records for one event; the same SOP existing under two document numbers

Two of these deserve particular attention because they are the ones that most often look acceptable and are not.

Consistency is the dimension that breaks aggregate analysis, and it is invisible from inside a single site. Every site can be perfectly internally consistent while the organization as a whole has four different vocabularies. The only way to see it is to pull the distinct values in use across sites and count them. When a field intended to hold eight categories turns out to hold sixty distinct values, that number is the finding.

Validity is where free text does its damage. A root cause field that accepts free text will collect free text, and the information it holds becomes unavailable to anything except a human reader. This is worth separating from completeness in your reporting, because the two imply different fixes. Incomplete data needs backfilling. Invalid data needs a template change and, usually, a controlled vocabulary that does not yet exist.

How this connects to the two frameworks

Accuracy and completeness are largely ALCOA+ territory, and mature organizations score well on them. Consistency, validity, and uniqueness are the measurable expression of FAIR’s interoperability and reusability principles. When an assessment comes back strong on the first pair and weak on the second three, that is the ALCOA+ and FAIR gap showing up as numbers rather than as an argument.

Where Data Quality Predictably Erodes

Data quality problems in a quality system are rarely random. They accumulate at four identifiable events, and knowing which ones an organization has been through tells you where to look before you have run a single query.

System Migrations

Migration is the single largest source, because it is the moment when field mappings are decided under schedule pressure. Fields without a clean destination get concatenated into a notes field. Historical categories get mapped to a current list with a best-fit judgment that is rarely documented. The data arrives intact by record count and degraded by structure, and because validation focuses on whether records transferred rather than whether meaning survived, nothing flags it.

This is why assessment before migration returns more value than assessment after. The findings become mapping requirements and cleanup criteria while there is still a decision to influence.

Mergers and Acquisitions

Two organizations with mature, internally consistent quality systems produce one organization with two vocabularies. The deviation classification schemes will not match. The severity scales will not match. The document numbering will not match. Integration usually addresses the systems and defers the definitions, and the deferred work is precisely the interoperability work that AI depends on.

Paper to Digital Conversion

Converting a paper process to an electronic one preserves the form and loses the structure unless someone deliberately redesigns the capture. A scanned investigation report is electronic without being digital: it satisfies retention, it is retrievable, and no system can read what is in it. Organizations that digitized early often have the longest data history and the least usable, which is a difficult finding to deliver and an important one.

Multi-Site Growth Without Central Definitions

A second site stands up its quality system by copying the first, then both evolve independently for several years. Neither has done anything wrong. The divergence is gradual, nobody owns the comparison, and it surfaces only when someone tries to report across both. By then the gap is a decade of records in two incompatible vocabularies.

Using this diagnostically

Before running any assessment, ask which of these four the organization has been through in the last ten years. Each one predicts a specific failure: migrations predict structural loss and undocumented mappings, acquisitions and multi-site growth predict vocabulary divergence, and paper conversion predicts unusable historical data. The answers tell you which dataset to assess first.

Why Annex 22 Reads Like FAIR in GMP Language

Regulators are moving toward this ground rather than away from it. EMA set out its broader position in a 2024 reflection paper covering the use of artificial intelligence across the medicinal product lifecycle, from drug discovery through post-authorization.12 The clearest and most specific evidence, though, is the draft EU Annex 22 on artificial intelligence.

The European Commission ran a stakeholder consultation on revisions to EudraLex Volume 4 covering Chapter 4, Annex 11, and a new Annex 22.78 As of this writing the annex remains in draft. The consultation has closed, EMA’s GMP and GDP Inspectors Working Group held a multistakeholder workshop on 30 June and 1 July 2026 to gather expert input, and EMA is still working through the results.10 Nothing described below is in force yet.

Read the scope before you read anything else into it

The draft annex applies to static models, meaning models that do not adapt during use, and to models with deterministic output. It states that dynamic models which continuously and automatically learn during use are not covered and should not be used in critical GMP applications. It states the same of models with probabilistic output.

It then says explicitly that the document does not apply to generative AI and large language models, and that such models should not be used in critical GMP applications. Where they are used in non-critical applications, a qualified human remains responsible for the output, which the draft names as human-in-the-loop.7

This is considerably narrower than most commentary suggests. If your AI ambition involves a large language model touching anything with direct impact on patient safety, product quality, or data integrity, the draft annex is not a compliance pathway for it.

Within that scope, look at what the draft actually requires of data. Section 5 covers test data selection, and the wording is that test data should be representative of and expand the full sample space of the intended use, should be stratified, should include all subgroups, and should reflect the limitations, complexity, and all common and rare variations within the intended use. The criteria and rationale for selection are to be documented.7

Section 6 requires that people involved in developing and training the model have never had access to the test data, that the test data be protected by access control and audit trail functionality logging accesses and changes, and that there be no copies outside that repository. It further requires a record of which data was used for testing, when, and how many times.7

Section 10 covers operation: change control over the model, the system, and the process before deployment; configuration control with measures to detect unauthorized change; regular monitoring of model performance against its defined metrics; and monitoring of whether input data remain within the model’s sample space, with metrics defined for monitoring any drift in the input data.7

Now read those requirements as data requirements rather than as model requirements. To show that test data is stratified and includes all subgroups, you must be able to describe your data’s structure. To document the rationale for selection, you must have provenance. To monitor drift in input data, you must have a stable definition of what the input is supposed to look like. Those are findability and reusability requirements written in GMP language.

The practical implication

An organization whose deviation categories differ across sites cannot demonstrate that a test dataset “includes all subgroups,” because it cannot define the subgroups consistently. The data work and the regulatory work are the same work.

A Two-Week Assessment You Can Actually Run

The instinct at this point is to launch a data governance program. Resist that for a moment. Governance programs are slow to show value and easy to stall, and the thing that gets a second phase funded is evidence from the first.

What follows fits in two weeks with one analyst and a quality lead who can answer questions.

1

Pick One Dataset With an Obvious Reuse Case

Deviation and CAPA classification, or SOP metadata. Deviation data is the better choice if you have a trending or prediction ambition. SOP metadata is better if a system migration is coming, because the findings become migration scope. Scope it to one site, one product, or one department. Not the enterprise.

2

Run Your Existing Data Integrity Assessment First

Use whatever ALCOA+ assessment you already have. Record the result. In most mature organizations this comes back largely clean, and that clean result is the point: it is the control that makes the second assessment meaningful.

3

Run a Separate FAIR Pass, and Keep the Results Apart

Four questions. Can a system locate these records on their properties rather than their identifiers? Can they be retrieved in bulk through a defined route? Do the categories and identifiers mean the same thing as the equivalent fields elsewhere? Does a record carry enough context for someone who did not create it to interpret it correctly?

4

Count. Do Not Estimate

How many deviation records in the last twelve months have a root cause only in free text? How many documents have no current owner? How many distinct category values exist where there should be one controlled list? A precise number from a narrow scope is more defensible than a rough figure from a wide one, and it is the number that goes in the business case.

5

State the Finding in One Sentence a Non-Specialist Understands

Something in the shape of: at this site, 40 percent of deviation records classified as human error carry no structured root cause, which makes trending across sites impossible. That sentence is the deliverable. Everything else is supporting evidence.

The output of this assessment is a document showing that the same dataset scores well on one framework and badly on the other. That contrast is the most persuasive artifact available, because it preempts the objection you will otherwise meet: that the data is already compliant and therefore already fine.

What to Fix First

Prioritize with a method regulators already accept. ICH Q9(R1) gives you quality risk management as a structure: assess likelihood and impact on product quality and patient safety, grade the findings, and sequence remediation accordingly.11 Using it means your prioritization is defensible rather than a matter of preference.

There is a second reason to work this way. Remediating a data quality problem in a validated system is itself a change to that system, and FDA’s computer software assurance guidance sets out a risk-based approach to sizing the assurance effort: establish the intended use, assess whether a failure would affect product quality or patient safety, and apply activities proportionate to that risk rather than scripted testing applied uniformly.13 Applied to a metadata or vocabulary correction, that reasoning usually produces a modest answer rather than a full re-validation.

Two practical rules on top of that. Fix interoperability before you chase volume, because a harmonized category list applied to two sites is worth more than a complete but inconsistent dataset across twelve. And structure what is currently free text going forward rather than retrospectively: change the capture template so next year’s data is usable, then decide deliberately how far back to remediate.

Four Objections Worth Answering

“Our Data Is Already Compliant”

It probably is, and that is not the claim in question. Compliance is a statement about trustworthiness. The question here is whether records can be combined and interpreted by a system that did not create them. Running both assessments separately is what makes the distinction concrete rather than theoretical.

“This Is an IT Problem”

IT builds and runs the systems. The definitions, categories, controlled vocabularies, and ownership have to come from the people accountable for the process. The FY2025 citation pattern points the same way: the most cited findings concern whether procedures exist and are followed, which is quality’s territory.3

“We Should Wait for the New System”

Migrating unresolved data quality problems into a new platform moves them without fixing them, and you will have paid for the move. Assessment before migration is one of the highest-return sequences available, because the findings become migration scope and cleanup criteria rather than problems discovered after go-live.

“Our AI Plans Are Years Away”

The remediation described here takes longer than the modeling does. Harmonizing a category list across sites is a change control exercise with training implications, and that is measured in quarters. The work done on data quality today determines what is available to use later, which is the whole argument.

Conclusion

ALCOA+ and FAIR are not competing frameworks and there is no choice to make between them. They answer different questions, and an AI program needs both answered.

If quality data is not ALCOA+ compliant, that is a compliance problem and it comes first. If quality data is ALCOA+ compliant but not FAIR, that is a different problem, and it is the one that stops an AI project after the proof of concept while everyone involved continues to report, accurately, that the data is compliant.

The regulatory direction reinforces this rather than complicating it. The draft Annex 22 requires organizations to characterize their data’s structure, document selection rationale, and define metrics for monitoring drift in input data.7 An organization that cannot describe its data consistently cannot meet those requirements, whatever the state of its audit trail.

Naming the problem correctly is what makes it fixable. A specific finding, with a count, from a bounded scope, gets funded. A general concern about data readiness does not.

For Further Reading