What Article 10 Says, and When It Actually Applies

Article 10 is titled “Data and data governance.” It is short, roughly one page of legal text, and it is the single most operationally demanding clause in the high-risk chapter of the EU AI Act for anyone who owns data pipelines. Risk management (Article 9), technical documentation (Article 11), record-keeping (Article 12), and human oversight (Article 14) all have recognizable analogues in a pharma quality system. Article 10 does not.

The clause in plain terms

Paragraph 1 sets the trigger: high-risk AI systems that involve training models with data must be developed on the basis of training, validation, and testing datasets that meet the quality criteria in paragraphs 2 through 5.2 That framing matters. The obligation attaches to the datasets, not to the model, and it attaches to all three splits, not only the training set.

Paragraph 2 requires that those datasets be subject to data governance and management practices appropriate to the intended purpose, and then lists eight specific practices those governance arrangements must cover: design choices; data collection processes and the origin of the data; data preparation operations such as annotation, labeling, cleaning, updating, enrichment, and aggregation; the assumptions made about what the data measures and represents; an assessment of the availability, quantity, and suitability of the datasets needed; an examination for possible biases likely to affect health and safety, negatively affect fundamental rights, or lead to discrimination prohibited under Union law; measures to detect, prevent, and mitigate those biases; and the identification of relevant data gaps or shortcomings that prevent compliance, together with how those gaps can be addressed.2

Paragraph 3 is the sentence that gets quoted most often and understood least: datasets shall be relevant, sufficiently representative, and to the best extent possible free of errors and complete in view of the intended purpose. Paragraph 4 adds that datasets must account for characteristics particular to the specific geographical, contextual, behavioral, or functional setting in which the system is intended to be used. Paragraph 5 gives providers a narrow, safeguarded legal basis to process special categories of personal data where that is strictly necessary to detect and correct bias. Paragraph 6 confirms that for high-risk systems that do not involve training a model, the data governance requirements apply to the testing datasets only.2

Recital 67 puts the intent plainly: data quality is essential for high-risk AI to perform safely, and the requirements are meant to make sure the datasets are good enough for the specific purpose the system is being placed on the market for.7 The legislator did not attempt to define an absolute quality bar. Everything is anchored to intended purpose. That is both the flexibility and the trap, because a purpose defined loosely produces acceptance criteria that cannot be tested.

The date most people still have wrong

Through 2025 and into 2026, the working assumption across life sciences was that the high-risk obligations, Article 10 among them, would apply from 2 August 2026. That is no longer correct.

The Commission proposed a simplification package in November 2025. The Council and the European Parliament reached political agreement on the AI-related part of that package in May 2026, and the Digital Omnibus on AI was formally adopted by the Council on 29 June 2026 and entered into force on 27 July 2026, just before the original date.18 The effect on timing is straightforward:

2 Dec 2027 New application date for high-risk obligations covering stand-alone Annex III systems, including Article 109
2 Aug 2028 New application date for AI that is a safety component of, or itself, a product covered by Union harmonization legislation9
8 Governance practices explicitly enumerated in Article 10(2), from design choices to data gap identification2

What did not move is worth naming, because several teams have read “delay” as “nothing happens in 2026.” The transparency obligations stayed on their original schedule, with a limited grace period into December 2026 for marking machine-generated content in systems already on the market. The prohibitions and the AI literacy duty have applied since February 2025. General-purpose AI model obligations have applied since August 2025. The governance and penalties provisions, including the role of the AI Office and the national market surveillance authorities, have applied since 2 August 2025.1 The supervisory machinery is live. Only the substantive high-risk requirements moved.

Why the deferral is less generous than it looks. The harmonized standards that would give a presumption of conformity with the high-risk requirements are not ready. CEN and CENELEC’s Joint Technical Committee 21 missed its original April 2025 target, and the standards bodies agreed in October 2025 to accelerate by allowing direct publication after a positive enquiry vote rather than a separate formal vote, aiming for availability in the fourth quarter of 2026.16 If the full suite arrives late in 2026, the practical runway between usable standards and the December 2027 date is closer to twelve months than to the sixteen months on the calendar. Data work of this kind does not compress well.

The one substantive change to Article 10 itself

The Omnibus left the data quality criteria in Article 10 alone. What it changed was paragraph 5. The permission to process special categories of personal data for the purpose of detecting and correcting bias, which previously sat only with providers of high-risk AI systems, now extends more broadly, subject to a strict necessity standard and a set of safeguards: preference for non-sensitive or synthetic data where that would work, pseudonymization, access controls, no onward transmission, and deletion once the bias work is done.1011

Read that carefully, because it cuts both ways. It removes a genuine legal obstacle that had made bias examination difficult for teams working with health data under the GDPR. It does not create an obligation to perform bias detection where none existed. And it does not relieve anyone of the Article 10(2)(f) duty where the system is high-risk. For a pharma data team, the practical reading is: the legal basis you were told you did not have for holding protected attributes long enough to test for bias is now clearer, and the safeguards you must wrap around that processing are now specified.

Which Pharma AI Systems Are Actually High-Risk

This is the question that should be answered before a single acceptance criterion is written, and in most organizations it has not been. Data teams are being handed a compliance mandate for a population of systems nobody has enumerated. The result is either paralysis or the opposite failure, where a team applies Article 10 rigor to every model in the estate and exhausts itself producing evidence for systems that were never in scope.

There are exactly two routes into high-risk classification, and a filter that can pull a system back out.

Route one: AI inside a regulated product

Under Article 6(1), an AI system is high-risk when it is a safety component of a product, or is itself a product, covered by the Union harmonization legislation listed in Annex I, and that product must undergo a third-party conformity assessment before being placed on the market.3 For life sciences this is mostly the Medical Devices Regulation and the In Vitro Diagnostic Regulation. The Machinery Regulation was moved from Section A to Section B of Annex I by the Omnibus, so AI performing a safety function on production equipment no longer triggers the substantive high-risk regime through that route.

The Omnibus narrowed the definition of “safety component” by adding that the component must have the intended purpose of preventing or mitigating risks to health and safety.9 That is a meaningful tightening. A model that optimizes throughput on a filling line is not a safety component simply because it runs on equipment that has safety functions. A model that decides when to halt a line to prevent harm probably is.

For a pharma or biotech company, this route usually bites in one of three places: a companion diagnostic or a diagnostic algorithm the company owns; software the company develops that meets the definition of a medical device; and AI performing a safety function on machinery covered by the Machinery Regulation. The Omnibus also empowers the Commission to adopt delegated acts limiting AI Act requirements where sectoral product legislation already provides equivalent protection, with the Machinery Regulation called out specifically.9

Route two: the Annex III list

Article 6(2) makes any system falling within Annex III high-risk by default. Annex III has eight headings: biometrics; critical infrastructure; education and vocational training; employment and worker management; access to essential private services and public services and benefits; law enforcement; migration, asylum, and border control; and administration of justice and democratic processes.4

Notice what is not there. There is no heading for drug discovery, clinical development, manufacturing, supply chain, pharmacovigilance, or regulatory affairs. Annex III is organized around AI that makes consequential decisions about individual people, not around AI that supports the development of a product.

Three Annex III entries do reach into life sciences organizations, and two of them catch companies that assume they are outside the Act entirely:

  • Employment and worker management. AI used to place job advertisements, filter applications, evaluate candidates, allocate tasks, monitor performance, or influence promotion and termination decisions is high-risk.4 Almost every large pharma company runs one of these. It is usually owned by human resources, procured from a vendor, and entirely absent from the AI inventory that IT or quality maintains.
  • Life and health insurance risk assessment and pricing. Relevant to any group in the organization doing this, and directly relevant if the company operates a health benefits or patient support arrangement with underwriting characteristics.4
  • Emergency call classification, dispatch priority, and patient triage systems. This is the entry that catches patient-facing services. A medical information line that classifies incoming contacts by urgency, or a patient support program that triages symptom reports, needs a careful read against this heading.4

The Article 6(3) filter, and the paperwork it still creates

Article 6(3) lets a system that falls under an Annex III heading escape high-risk status if it does not pose a significant risk of harm to health, safety, or fundamental rights, and if it meets one of four conditions: it performs a narrow procedural task; it improves the result of a previously completed human activity; it detects decision-making patterns or deviations from prior patterns without replacing or influencing the earlier human assessment absent proper human review; or it performs a preparatory task to an assessment relevant to the Annex III use case. A system that performs profiling of natural persons is always high-risk regardless.3

The filter is real and it is useful. It is also not free. Article 6(4) requires the provider to document that assessment before the system is placed on the market, and to make that documentation available to national authorities on request. In other words, deciding that a system is out of scope is itself a documented decision with an evidence trail. Teams that quietly conclude “this one is fine” without recording the reasoning have created an audit finding, not an exemption.

Provider or deployer changes what you owe

Most pharma and biotech companies are deployers, not providers. They buy AI systems and put them into use rather than developing and placing them on the market under their own name. Article 10 in full applies to providers. Deployers get a shorter but pointed obligation in Article 26(4): where the deployer exercises control over the input data, it must ensure that input data is relevant and sufficiently representative in view of the intended purpose of the high-risk AI system.5 Deployers also have to use the system according to the instructions, monitor its operation, suspend use and notify the provider where they suspect a risk, and keep the system’s automatically generated logs for at least six months.

“Relevant and sufficiently representative” appears in both Article 10(3) and Article 26(4). The words are the same, the scope is narrower for the deployer, and the measurement problem is identical. A deployer that fine-tunes a purchased model on its own data, or that supplies the reference data the model reads at inference time, is squarely inside this obligation. A deployer that materially modifies a high-risk system, or puts it on the market under its own name, can become a provider and inherit the whole of Article 10.

A working scoping table

The table below is a starting position, not a legal opinion. Classification depends on the specific intended purpose written into the system’s documentation, and reasonable people disagree at the margins. Industry has pressed this point directly: EFPIA has argued that the majority of AI used in medicines research and development is neither regulated under Annex I nor listed in Annex III, and therefore cannot qualify as high-risk under the Act.1213 Commentators have noted that the sector has been waiting some time for clearer guidance on exactly where the boundary sits.14

AI system in a pharma or biotech settingLikely classificationReasoning
Target identification or compound screening modelNot high-risk under the AI ActNo Annex III heading covers drug discovery, and the model is not a safety component of an Annex I product. GxP expectations still apply where the output feeds a regulatory decision.
Clinical trial patient recruitment or site selection modelUsually not high-riskNot an Annex III use case. Fairness exposure is real but arises under clinical and data protection law rather than the AI Act.
Manufacturing process control or batch release support modelUsually not high-risk under the AI ActFalls outside Annex III. Draft Annex 22 of the EU GMP guide is the governing expectation here, not Article 10.
AI safety function on production machineryPotentially high-risk (Annex I)Machinery Regulation route, subject to the narrowed safety component definition and pending Commission clarification.
Pharmacovigilance signal detection or case triageUsually not high-risk under the AI ActNot an Annex III use case. Governed by GVP expectations and, where it supports a regulatory decision, by agency guidance on AI credibility.
Diagnostic or companion diagnostic algorithmHigh-risk (Annex I)Medical device or IVD requiring third-party conformity assessment.
Clinical decision support offered to prescribersFrequently high-risk (Annex I)Classification turns on whether the software qualifies as a medical device under the MDR.
Recruitment screening or performance evaluation toolHigh-risk (Annex III, point 4)Employment and worker management is an explicit heading. Commonly missed because it sits outside IT and quality.
Patient support program symptom triagePotentially high-risk (Annex III, point 5)Triage and urgency classification is named in the essential services heading.
Internal knowledge assistant over SOPs and study documentsNot high-riskNo Annex III heading. Transparency duties may still apply, and GxP controls apply if outputs inform regulated decisions.

Article 10 Clause by Clause: From Legal Text to Acceptance Criteria

Here is the practical problem. Article 10 is written in the register of European product safety law. It says datasets shall be “relevant,” “sufficiently representative,” and “to the best extent possible, free of errors and complete.” A data engineer cannot run “sufficiently representative.” A quality reviewer cannot approve it. An inspector cannot verify it. Somebody has to convert each phrase into a number, a threshold, and a document.

That conversion is the work. The table below does it clause by clause. For each provision it names the data quality dimension being measured, a testable acceptance criterion of the kind a team can actually write into a data validation plan, and the evidence artifact that demonstrates the criterion was met. The thresholds shown are illustrative placeholders. The right values come from the intended purpose and the risk assessment for the specific system, and setting them is a decision that should be made and recorded before testing begins, not fitted afterward.

Article 10 provisionWhat the text requiresDimensionTestable acceptance criterionEvidence artifact
10(2)(a) Relevant design choices for the datasets Traceability Every dataset design decision (inclusion rules, split strategy, sampling method, unit of observation) is recorded with a named owner and date before data collection begins. Zero undocumented design changes after the plan is approved. Approved dataset design specification under change control
10(2)(b) Data collection processes and origin of the data; for personal data, the original purpose of collection Traceability, provenance 100 percent of records trace to a named source system, extraction job, and time window. For personal data, the original collection purpose and lawful basis are recorded per source. Data lineage record and source register with consent and purpose mapping
10(2)(c) Data preparation: annotation, labeling, cleaning, updating, enrichment, aggregation Accuracy, consistency Labeling instructions are version controlled. Inter-annotator agreement meets a preset threshold on a defined sample. Every transformation step is reproducible from raw data by re-running a versioned pipeline. Labeling protocol, agreement report, reproducible transformation code with commit reference
10(2)(d) Formulation of assumptions about what the data is supposed to measure and represent Validity Each feature and each label carries a written statement of what it is assumed to measure, plus at least one stated condition under which that assumption fails. Assumptions are reviewed by someone who did not write them. Assumption register with reviewer sign-off
10(2)(e) Assessment of the availability, quantity, and suitability of the datasets needed Sufficiency Minimum sample size per defined subgroup is calculated in advance and met, or the shortfall is declared. Suitability is assessed against the written intended purpose, not against convenience of access. Data sufficiency assessment with per-subgroup counts against target
10(2)(f) Examination for possible biases affecting health and safety, fundamental rights, or leading to prohibited discrimination Bias A named set of attributes is tested for distributional skew against a reference population, and model performance is measured separately per subgroup. Performance disparity between the best and worst performing subgroup stays within a preset limit or triggers a documented decision. Bias examination report covering dataset composition and subgroup performance
10(2)(g) Appropriate measures to detect, prevent, and mitigate identified biases Bias Every bias finding above the preset limit has a linked mitigation with a stated method (reweighting, targeted collection, threshold adjustment, restriction of intended purpose) and a re-test result showing the effect. Bias mitigation log with before and after measurements
10(2)(h) Identification of relevant data gaps or shortcomings preventing compliance, and how they can be addressed Completeness A gap register lists every known shortfall, its effect on the intended purpose, whether it is closable, and the plan and date if it is. An empty gap register is treated as an incomplete assessment, not a clean result. Data gap register with remediation plan and residual risk statement
10(3): relevant Datasets relevant in view of the intended purpose Relevance Every field in the dataset maps to a stated element of the intended purpose. Fields with no mapping are removed or justified. No target leakage: no feature is derived from information unavailable at prediction time. Field-to-purpose mapping and leakage test result
10(3): sufficiently representative Datasets sufficiently representative of the intended use population Representativeness The reference population is defined in writing. Distribution distance between dataset and reference population on named strata stays within a preset limit. Every stratum meets its minimum count. Splits preserve stratum proportions. Representativeness statement with distribution comparison per stratum
10(3): free of errors To the best extent possible free of errors Accuracy Error rate measured against an independently verified gold-standard sample stays below a preset threshold. Duplicate rate, out-of-range rate, and referential integrity failures each stay below their own thresholds. Data quality test report with measured rates against preset thresholds
10(3): complete To the best extent possible complete Completeness Missingness measured per field and per subgroup. Fields exceeding the threshold are either imputed with a documented method or excluded. Missingness that differs materially by subgroup is escalated as a bias signal, not treated as a cleaning task. Completeness profile by field and by subgroup
10(4) Account for characteristics of the geographical, contextual, behavioral, or functional setting of intended use Contextual fit Deployment settings are enumerated (country, site, instrument or platform, care setting, operating conditions). Each setting is either represented in the data or explicitly excluded from the intended purpose. Performance is reported per setting. Deployment setting matrix with per-setting performance and exclusions
10(5) Conditions for processing special categories of personal data for bias detection and correction Lawfulness, security Strict necessity is documented, including why non-sensitive or synthetic data would not work. Pseudonymization, access restriction, no onward transmission, and deletion after the bias work are all evidenced. Necessity assessment and processing record with deletion confirmation
10(6) For systems not trained on data, requirements apply to testing datasets only Scope control A recorded determination of whether the system involves model training, with the resulting scope of data governance stated explicitly. Scope determination memo referenced in the technical documentation

Three rules for using this table well.

  • Set thresholds before you measure. An acceptance criterion chosen after seeing the result is not an acceptance criterion. Draft Annex 22 makes the same point for GMP models, requiring metrics and criteria to be defined ahead of testing rather than fitted to outcomes.15
  • Apply every criterion to all three splits. Article 10 names training, validation, and testing datasets together. A representativeness statement that covers only the training set fails the clause on its face.
  • Write the intended purpose first. Nine of the fifteen rows above resolve against “in view of the intended purpose.” A vague purpose statement makes every one of them untestable, which is why purpose definition, not data work, is usually the real bottleneck.

Annex IV, which sets out what the technical documentation must contain, closes the loop. It requires a description of the training, validation, and testing datasets, their provenance, scope and main characteristics, how the data was obtained and selected, and the labeling and cleaning methods used.6 The artifacts in the right-hand column are not extra work invented for the sake of tidiness. They are the source material for a documentation pack you will have to produce anyway.

Where Article 10 Meets ALCOA+, and Where It Does Not

Pharma data teams do not start from zero here. They start with two decades of muscle memory built on ALCOA+ and 21 CFR Part 11. The instinct to map Article 10 onto that existing frame is correct, and it will get a team most of the way. It is the remainder that causes trouble, because the gap is not small and it is not the kind of gap that closes by working harder at things you already do well.

The overlap is real and worth claiming

ALCOA is attributable, legible, contemporaneous, original, and accurate. The plus adds complete, consistent, enduring, and available, and the MHRA’s GxP data integrity guidance sets out what each means in a regulated setting.17 A pharma organization that genuinely meets these across its data estate has already built much of what Article 10(2)(a) through (e) asks for, though it has usually built it for records rather than for datasets.

ALCOA+ attributeNearest Article 10 requirementWhat still has to change
Attributable10(2)(b) data collection processes and originAttribution has to extend to the dataset and the transformation, not only to the record. Who assembled this training set, from what, and when.
Legible10(2)(c) data preparation operationsLegibility becomes interpretability of labels and encodings. A label schema nobody can decode two years later fails this in substance.
Contemporaneous10(2)(b) collection process documentationExtends to recording when each extract was taken, because a dataset assembled from snapshots at different times introduces skew that is invisible later.
Original10(2)(b) and (c) origin and preparationThe raw source must remain reachable and the path from raw to model-ready must be reproducible, not just described.
Accurate10(3) free of errors to the best extent possibleAccuracy has to become a measured rate against a verified reference, with a preset threshold. “Accurate” as a qualitative assertion is not testable.
Complete10(2)(h) and 10(3) completeness and gapsGxP completeness means no data was discarded. Article 10 completeness also asks whether the data you have covers the population you will deploy against. Different question, same word.
Consistent10(2)(c) and (d) preparation and assumptionsExtends to consistency of definitions across sources, which is where multi-site and multi-system datasets usually break.
EnduringAnnex IV documentation retentionThe dataset itself, its version, and its documentation must endure alongside the records, so a past model decision can be reconstructed.
AvailableAnnex IV and Article 11 documentation on requestAvailable to market surveillance authorities and notified bodies, not only to GxP inspectors. A different audience with different questions.
No ALCOA+ analogue10(3) sufficiently representativeRequires defining a reference population and measuring distance from it. Nothing in ALCOA+ asks whether the data resembles the world the system will operate in.
No ALCOA+ analogue10(2)(f) and (g) bias examination and mitigationRequires identifying protected and clinically meaningful subgroups, measuring performance separately for each, and acting on disparity. Data integrity is silent on fairness.
No ALCOA+ analogue10(4) setting-specific characteristicsRequires enumerating deployment settings and showing the data reflects them. GxP asks whether the record is trustworthy, not whether it generalizes.

The two demands ALCOA+ never made

Strip the table down and two genuinely new obligations remain. Both are about the relationship between the data and the world, and neither can be satisfied by better record-keeping.

Representativeness. ALCOA+ is a set of properties of a record. A record can be perfectly attributable, contemporaneous, original, and accurate, and still describe a population that looks nothing like the population the model will be used on. A dataset assembled from three European sites can be flawless by every data integrity measure and still be unfit for a system intended for use across twenty countries. Article 10(3) and 10(4) ask a question data integrity has never asked: compared to what? Answering it requires naming a reference population, which is a scientific and commercial judgment, not a data engineering task. This is the single most common reason a representativeness statement stalls. Nobody wants to own the definition.

Bias examination. Article 10(2)(f) requires examining datasets for biases likely to affect health and safety, negatively affect fundamental rights, or lead to prohibited discrimination. There is no GxP analogue at all. Twenty-one CFR Part 11 governs electronic records and electronic signatures: it is concerned with authenticity, integrity, and, where appropriate, confidentiality.18 A system can be fully Part 11 compliant and systematically underperform for a subgroup, and nothing in Part 11 would surface that. The examination Article 10 asks for is a different discipline with different methods and, usually, different people.

The trap in “we already do data quality.” Traditional data quality asks whether values are correct, complete, and consistent. Every one of those questions can be answered by looking only at the dataset. Representativeness and bias cannot be answered from inside the dataset at all. They require an external reference: a population definition, a subgroup taxonomy, a performance comparison across groups. A team that scores well on conventional data quality metrics can still fail Article 10 comprehensively, and will not see it coming, because the measurement instrument does not point in that direction.

Regulators outside the AI Act are converging on the same point

This is not only a European product safety concern. The EMA’s reflection paper on the use of AI in the medicinal product lifecycle flags data quality and representativeness for small populations directly, noting that the need to over-sample rare populations should be considered and warning of risks of bias and other discrimination against non-majority genotypes and phenotypes, and expecting data sources and processing to be documented in a way that supports traceability consistent with GxP.1923 Draft Annex 22 of the EU GMP guide asks for a characterization of the model’s input data sample space, an assessment of limitations and possible biases, and division into subgroups where relevant so that each is sufficiently represented in the test data.1520

The FDA has taken a parallel route through its draft guidance on using AI to support regulatory decision-making for drug and biological products, which sets out a risk-based credibility assessment framework tied to a specific context of use, informed by the agency’s review of more than 300 submissions containing AI or machine learning components.2124 Different legal instrument, different vocabulary, same underlying demand: state what the model is for, show the data fits that purpose, and produce the evidence.

The Evidence Package a Data Team Can Actually Produce

Compliance conversations tend to end at principles. What survives an inspection is a set of documents that exist, are versioned, and can be produced within an hour. Below are the six artifacts that carry most of the weight for Article 10. They are deliberately small. Six documents that are current and honest are worth more than a thirty-document framework that is eighteen months stale.

ARTIFACT 1

Dataset specification and design record

Written before collection. States the intended purpose, the unit of observation, inclusion and exclusion rules, the split strategy, the sampling method, and who approved each choice. Covers Article 10(2)(a) and anchors everything downstream.

ARTIFACT 2

Data lineage and provenance record

Traces every field back to a named source system, extraction job, and time window, with the original collection purpose and lawful basis for personal data. Covers Article 10(2)(b) and is what an inspector will ask for first.

ARTIFACT 3

Preparation and labeling protocol

Version-controlled labeling instructions, the qualifications of annotators, the agreement measurement, and reproducible transformation code. Covers Article 10(2)(c) and is the artifact most often missing entirely.

ARTIFACT 4

Representativeness statement

Defines the reference population in writing, compares the dataset against it on named strata, states the acceptance limits set in advance, and reports the result. Covers Article 10(3) and 10(4). Requires a business decision, not just a query.

ARTIFACT 5

Bias examination and mitigation report

Names the attributes examined and why, reports dataset composition and per-subgroup model performance, records the disparity limit set beforehand, and links each finding to a mitigation with a re-test. Covers Article 10(2)(f) and (g).

ARTIFACT 6

Data gap register

Lists every known shortfall, its effect on the intended purpose, whether it can be closed, and the plan if it can. Covers Article 10(2)(h). An empty register is a warning sign, not a pass.

Two practices make these documents hold up better than they otherwise would.

The first is independence between the people who build the training data and the people who evaluate the test results. Draft Annex 22 is unusually specific about this, describing access controls, audit trails, and a requirement that staff involved in developing and training the model have never had access to the test data, with a four-eyes review by a colleague who has not had that access where separation is not possible.15 Article 10 does not spell it out, but the requirement that testing datasets meet the same quality criteria as training data points the same way. If the same person assembles the training set and decides whether the test result is acceptable, the test is not independent and the evidence is weak.

The second is treating dataset versions as controlled items. A dataset is not a file. It is a specific extraction, at a specific time, transformed by a specific version of a pipeline. If a regulator asks in 2029 why a model behaved the way it did in 2027, the answer requires reconstructing the exact dataset. Teams that version code but not data discover this late and cannot recover.

A structural shortcut worth taking. ISO/IEC 5259 is a published multi-part standard specifically on data quality for analytics and machine learning, covering data quality measures, management requirements, a process framework, and a governance framework.22 Building the artifact set on that vocabulary means the definitions are already agreed, already international, and already familiar to the standards bodies drafting the harmonized standards for the AI Act. It is a much easier position to defend than a set of internally invented terms.

One Evidence Set, Two Regulators

The strongest practical argument for doing this work now, rather than waiting for December 2027, is that almost none of it is unique to the AI Act.

A pharma organization deploying AI in a GMP setting is already heading toward draft Annex 22, which asks for a characterized input sample space, predefined metrics, independent test data, subgroup performance, change control, and ongoing monitoring of both performance and the input space.1520 The same organization submitting AI-derived evidence to the FDA is heading toward the credibility assessment framework, which asks it to state the context of use, assess model risk, and gather credibility evidence proportionate to that risk.21 The same organization operating a high-risk system in the EU is heading toward Article 10.

Three regulatory instruments, three vocabularies, one underlying evidence set. The dataset specification, lineage record, preparation protocol, representativeness statement, bias report, and gap register serve all three. What differs is the wrapper and the audience, not the substance.

Where the instruments genuinely diverge is worth knowing so nothing is missed:

  • Fundamental rights. Article 10(2)(f) explicitly names fundamental rights and prohibited discrimination as harms to examine for. Annex 22 and the FDA framework are concerned with product quality, patient safety, and the reliability of evidence. A bias examination built only for GMP purposes will look at clinical subgroups and miss protected characteristics.
  • Scope of models covered. Draft Annex 22 is aimed at static, deterministic models in critical GMP applications, and the draft itself states that it does not apply to generative AI and large language models, and that such models should not be used in critical GMP applications.15 The AI Act draws no such line. A high-risk system built on a general-purpose model is fully in scope.
  • Who owes the evidence. GxP obligations sit with the regulated company. AI Act obligations split between provider and deployer, and the split is not intuitive. A company can be a GMP-regulated manufacturer and an AI Act deployer at the same time, owing different evidence under each.

A note on where this lands organizationally. Article 10 does not fit neatly into any existing function. Legal can interpret it but cannot measure anything. Quality can govern it but does not own the pipelines. Data engineering can measure it but has no standing to define a reference population or approve a residual risk. In practice the assessments that hold up are the ones with a named owner in the data function, a defined review path through quality, and a documented decision point where the business owner accepts what the data cannot cover.

A Sequence for the Runway You Now Have

The deferral to December 2027 is genuinely useful, but only for organizations that use it. The sequence below is ordered by dependency, not by ease. Each step produces something the next one needs.

1

Inventory before anything else

Build a complete list of AI systems in use or in development, including the ones procured outside IT. Human resources tooling, patient-facing services, and vendor-embedded features are the three places systems hide. Record for each: intended purpose in one sentence, whether the organization is provider or deployer, and who owns it. Most organizations discover their inventory is between two and five times longer than they expected.

2

Classify, and document the classification

Run each system against Article 6(1), Annex III, and the Article 6(3) filter. Record the reasoning for every determination, including the out-of-scope ones, because Article 6(4) makes that documentation an obligation rather than a courtesy. Expect the classification exercise to force the intended purpose statements to get sharper, which is a benefit in itself.

3

Set the definitions that everything else depends on

For each high-risk system, define the reference population, the subgroup taxonomy, and the deployment setting matrix. This is the step teams skip, and skipping it is why representativeness and bias work stalls later. These are business and scientific decisions with regulatory consequences, and they need a named approver.

4

Write acceptance criteria before you measure

Convert each Article 10 provision into a threshold using the translation table as a starting point, tuned to the risk of the specific system. Get the criteria reviewed and approved before any measurement runs. Criteria set after the fact do not survive scrutiny and everyone in the room knows it.

5

Run a full pass on one system end to end

Pick one high-risk system, ideally one that is uncomfortable rather than easy, and produce all six artifacts against the approved criteria. A single completed pass teaches more about where the organization’s data actually breaks than a year of framework design, and it produces a template the rest of the estate can follow.

6

Fix the pipeline, not just the document

The first pass will surface structural problems: missing provenance, subgroup fields that were never collected, transformations that cannot be reproduced. These take quarters to fix, not weeks. Starting them in 2026 rather than 2027 is the entire value of the deferral.

7

Wire in monitoring and periodic re-examination

Representativeness decays. Populations shift, sites change, coding practices drift. Build the re-measurement into periodic review so the representativeness statement and bias report are refreshed on a defined cycle rather than reconstructed under pressure when someone asks.

One practical note on standards. Because none of the harmonized standards supporting the high-risk requirements have been published and cited in the Official Journal yet, no organization can rely on the presumption of conformity today.16 That is not a reason to wait. It is a reason to build on the closest published international standards, ISO/IEC 5259 for data quality among them, so that when the harmonized standards arrive the mapping is a translation job rather than a rebuild.

Failure Modes Worth Naming Early

The patterns below come up repeatedly in organizations that are doing serious work and still going wrong. Naming them early is cheaper than finding them in an assessment.

Treating the deferral as a pause

The most common response to the December 2027 date has been to move the work down the priority list. The pipeline remediation that Article 10 requires runs on a multi-quarter clock. Provenance that was never captured cannot be captured retrospectively. Subgroup attributes that were never collected cannot be recovered from a warehouse. The organizations that will be ready in December 2027 are the ones that started in 2026.

Scoping by anxiety instead of by rule

Two opposite errors, both driven by the same underlying uncertainty. Some teams decide everything is high-risk and try to produce Article 10 evidence for every model in the estate, which guarantees that nothing gets done properly. Others decide that because pharma is not named in Annex III nothing applies, and miss the recruitment tool and the patient triage service entirely. The rule is narrow and readable. Apply it, and write down what you concluded.

Defining intended purpose too broadly

“Support clinical decision-making” is not an intended purpose. It cannot be tested against, which means no acceptance criterion derived from it can be tested either. A usable purpose statement names the population, the setting, the decision the output feeds, and the boundaries of use. Broad purpose statements feel safer commercially and are far riskier in an assessment, because they widen the population the data has to be representative of.

Confusing a bias policy with a bias examination

Article 10(2)(f) requires an examination. An examination produces numbers: dataset composition by subgroup, model performance by subgroup, and a comparison against a limit set in advance. A commitment to fairness, however sincerely written, produces none of those. This distinction is where the largest volume of remediation work usually sits.

Forgetting the deployer obligation

Companies that buy rather than build often assume Article 10 is entirely the vendor’s problem. Article 26(4) puts the input data obligation on the deployer wherever the deployer controls that data, using the same “relevant and sufficiently representative” language.5 If your organization supplies the reference data, the lookup tables, the site master data, or the fine-tuning set, that obligation is yours.

Leaving the reference population undefined

Every representativeness assessment needs a comparator. Defining it means someone has to state, in writing, which population the system is meant to serve. That statement has commercial and clinical consequences, so it tends to circulate without anyone signing it. Assign the decision to a named owner with a date, and treat a missing reference population as a blocking issue rather than an open item.

Conclusion

Article 10 is a data engineering requirement wearing the clothing of a legal obligation. Most of the effort it demands cannot be delegated to legal, cannot be resolved by policy, and cannot be produced in the final quarter before a deadline. The clause asks a data team to state precisely what a system is for, prove that the datasets behind it match that purpose, demonstrate that they represent the population the system will meet in the real world, and show that they were examined for bias against a limit set in advance. Then it asks for the documents.

The good news for pharma and biotech is that the foundation is already there. ALCOA+ and 21 CFR Part 11 have built genuine discipline around provenance, traceability, and record integrity, and Article 10(2)(a) through (e) sits comfortably on top of that. The honest news is that representativeness and bias examination are new. They ask a question data integrity has never asked, which is not whether the data is trustworthy but whether it resembles the world. That question needs a reference population, a subgroup taxonomy, and a preset limit, and none of those come out of a data warehouse. They come out of a decision somebody has to own. The deferral to December 2027 gives organizations time to make those decisions carefully, and to fix the pipeline problems the first honest assessment will expose. That time only helps the teams that start using it now.

Sakara Digital works with pharma and biotech organizations building the data foundations that AI governance requirements depend on, from scoping which systems are actually in scope through to the acceptance criteria and evidence artifacts that hold up under assessment. If you are working out where Article 10 lands in your estate and want an independent perspective on where to start, we are happy to have that conversation.

For Further Reading