In This Article
- Executive Summary
- Why This Problem Is Still Open After Thirty Years
- The Affiliation Problem Is the Real Difficulty
- Three Matching Approaches, Compared Honestly
- Choosing an Approach and Budgeting the Stewardship Effort
- Survivorship and the Golden Record
- Where Resolution Errors Become Compliance Findings
- Measuring Whether Resolution Is Actually Working
- A Practical Sequence for Fixing an Existing Estate
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Resolving healthcare professional (HCP) and healthcare organization (HCO) records into one trustworthy identity is among the oldest problems in pharma commercial data, and it is still one of the least well solved. Almost every commercial organization has a customer master. Very few have one that field teams, medical affairs, and compliance all agree with. The reason is rarely the matching engine. It is that the underlying reality changed shape: 82% of practicing US physicians were employed by hospitals or other corporate entities as of January 1, 2026, up from 25.8% in 2012, and the affiliation between a person and an institution is now the most important and most volatile part of the record.7
The argument in this article is that name matching is the easy part and affiliation is the hard part. An HCP is associated with several HCOs at once, those associations change, and most master data models flatten a many to many, time varying relationship into a single current employer field. Every downstream number that depends on who a prescriber works for, from territory alignment to aggregate spend, inherits that flattening error. The article compares deterministic rules, probabilistic matching, and machine learning based resolution, and states plainly when each is the right fit and what each demands from data stewards.
It then covers survivorship and golden record design, which decides which source wins for each attribute and why a single trusted source policy usually fails. Finally it grounds the consequences in transparency reporting under the Physician Payments Sunshine Act and the EFPIA Disclosure Code, where a wrong resolution stops being a data annoyance and becomes a misattributed public spend report with a named physician on it.
Why This Problem Is Still Open After Thirty Years
Pharma commercial organizations have been buying HCP reference data, building customer masters, and running match and merge routines since the 1990s. The tooling is mature. The vendor market is well developed. Data stewardship teams exist in almost every mid-cap and large commercial organization. And yet in most companies, if you ask three different functions how many prescribers they engage in a given specialty, you get three different answers, and none of them can be reconciled without a week of manual work.
That persistence is worth taking seriously rather than treating as an execution failure. The problem stays open because the thing being modeled keeps moving, and because the model most organizations use is structurally too simple for what it is trying to represent.
The identity data itself is unstable
Consider the reference sources that a commercial customer master typically draws on. The National Plan and Provider Enumeration System (NPPES) publishes National Provider Identifier records, and CMS makes downloadable files and monthly and weekly update files available for public use.19 Those files are the backbone of most US HCP masters. But NPPES is self-reported and self-maintained, which means practice addresses, taxonomy codes, and credentials go stale between the moment a provider moves and the moment that provider gets around to updating a federal registry that has no direct bearing on their day.
The best public evidence for how bad this gets comes from CMS itself. In the third round of its Medicare Advantage online provider directory reviews, conducted between November 2017 and July 2018, CMS examined 5,602 providers at 10,504 listed locations across 52 Medicare Advantage organizations. It found that 48.74% of listed provider locations had at least one inaccuracy, with the rate for individual organizations ranging from 4.63% to 93.02%. Inaccuracies with the highest likelihood of preventing access to care appeared in 41.75% of all locations.8
The same report cites an earlier study of Medicare Advantage dermatology listings in which, among 4,754 dermatologists listed in the largest plans across twelve metropolitan areas, 45.5% were duplicates within the same plan directory, and among the remaining unique listings only 48.9% were reachable, accepted the listed plan, and offered an appointment.8 Nearly half of the entries in a curated, regulated, commercially important directory were duplicates of each other. That is an entity resolution failure in its purest form, in an industry with a direct financial incentive to get it right.
The people who could fix the source data are the least able to
There is a second reason the upstream data stays poor, and it explains why waiting for public sources to improve is not a strategy. CAQH surveyed 1,240 physician practices about the administrative work of responding to health plan requests to verify and update directory information. The average practice holds 20.2 plan contracts, each with its own platform, format, and schedule. Practices reported spending at least one full staff day per week on directory maintenance, at an average cost of $998.84 per month, rising to $2,525.31 per month for practices with more than 25 providers.9
The practices carrying the burden of keeping provider data current have no incentive tied to a pharmaceutical company’s customer master. Every commercial organization is therefore building identity on top of a source layer that is maintained under duress by people with no stake in the outcome. Any design that assumes the reference data will get cleaner is starting from a false premise.
The structure is consolidating faster than the data model
The third reason is the one that matters most, and it is the subject of the next section. Physician practice has consolidated dramatically. The Physicians Advocacy Institute and Avalere Health found that 82.0% of practicing physicians were employed by hospitals or other corporate entities as of January 1, 2026, split between 59.7% employed by hospitals and 22.3% by corporate entities such as insurers, private equity backed groups, and health technology firms. Between 2018 and 2026, hospitals and corporate entities acquired 85,000 additional physician practices, leaving 157,200 practices (63.9% of the total) under non-physician ownership.7
When most prescribers were independent, a single employer field on an HCP record was a reasonable simplification. It is no longer one. A prescriber today may hold a hospital appointment, practice at two outpatient sites owned by a different corporate entity, sit on a health system committee, and hold an academic title. All four are affiliations. All four matter to a different function inside the commercial organization. A single field can hold one of them.
The Affiliation Problem Is the Real Difficulty
Most conversations about HCP entity resolution get stuck on name matching. Is Dr. Robert Chen the same person as Dr. Bob Chen, R. Chen MD, and Robert C. Chen? That question is real, and the techniques for answering it are well understood and covered in the next section. But name matching is a solved research problem with a mature toolset. Affiliation is neither.
Industry practitioners have been saying this for years. A review of master data management in pharma described affiliation data, meaning which organizations an HCP belongs to, is employed by, or is otherwise connected to, as one of the most important current factors in pharmaceutical master data management, precisely because compliance functions handling global aggregate spend and commercial functions running account based selling are both pulling on the same relationships.17
What the relationship actually looks like
The relationship between an HCP and an HCO has four properties that a single employer field cannot represent:
- It is many to many. One prescriber has several affiliations at the same time. One institution has thousands of affiliated prescribers, many of whom are also affiliated elsewhere.
- It is typed. Employment is not the same relationship as admitting privileges, which is not the same as a teaching appointment, which is not the same as membership on a formulary committee. Treating them all as “affiliation” throws away the distinction that determines which function cares.
- It is time varying. Every affiliation has a start, sometimes an end, and a period of overlap with its successor. A prescriber who moved practices in March had two valid affiliations in the same calendar year.
- It is weighted. A prescriber who spends four days a week at one site and half a day at another has two true affiliations of very different commercial and clinical significance.
A model that stores one current employer per HCP record captures none of these. It captures a single point on a moving line, refreshed whenever the master data feed happens to update, with no history and no way to answer a question about a past state.
The flattening error, stated plainly. When a many to many, typed, time varying, weighted relationship is stored as one current employer field, three things become impossible: you cannot reproduce a past reporting period, you cannot attribute an interaction to the site where it happened, and you cannot count an institution’s prescribers without either double counting people who are affiliated in more than one place or losing the ones whose primary affiliation sits elsewhere.
What breaks downstream
The consequences of the flattening error are specific, and they surface in different functions in ways that look unrelated until you trace them back.
Territory alignment and field targeting
Territories are drawn on where prescribers are. If the master holds one address per prescriber and that address is the hospital rather than the outpatient clinic where they see most patients, the prescriber lands in the wrong territory. Two representatives call on the same person. A third assumes someone else has the account. The field force notices this long before the data team does, and their response is to build a private spreadsheet, which is how a commercial organization ends up with a customer master that nobody uses.
Account based engagement
Selling to an integrated delivery network requires knowing which prescribers sit under which parent organization. That requires an HCO hierarchy that models parent and child institutions, and an HCP to HCO relationship that can attach a person to a child site while still rolling up to the parent. Where the HCO hierarchy is thin or the affiliations are single valued, account level reporting either double counts prescribers who appear under several sites or undercounts an institution’s real reach. Both errors are invisible in the report and obvious in the field.
Medical affairs and speaker programs
Medical teams care about a different affiliation type than commercial teams. A key opinion leader’s academic appointment and society roles matter more than their billing address. If the master models only employment, medical affairs maintains its own separate list, and the two lists drift apart until nobody can produce a single view of the organization’s interactions with a given physician.
Transparency reporting
This is the one that turns a data problem into a compliance problem, and it is covered in detail later in this article. Whether a transfer of value is reported against an individual physician or against an institution depends on the affiliation in place at the moment of the interaction. If the master holds only the current affiliation, and the reporting run happens eight months later, the system attributes the payment to wherever the prescriber is now.
How to model it properly
The fix is not exotic. It is a separate affiliation entity rather than an attribute on the HCP record. In practice that means a relationship table with, at minimum, the HCP identifier, the HCO identifier, the relationship type, a start date, an end date that is allowed to be open, a source system, a confidence indicator, and a flag for which affiliation is treated as primary for a given purpose.
| Element | What it holds | Why it is needed |
|---|---|---|
| Relationship type | Employment, privileges, academic appointment, committee role, group membership | Different functions need different types. Collapsing them forces every function to filter on the wrong thing. |
| Valid from and valid to | Effective dating on every relationship, with open ended current records | Makes a past reporting period reproducible. Without it, you cannot defend last year’s numbers. |
| Source and confidence | Which system asserted the relationship and how strongly | A claims derived affiliation and a rep entered affiliation are not equally reliable and should not be treated as such. |
| Primary flag, by purpose | Primary for field alignment may differ from primary for medical engagement | Avoids forcing one function’s definition of primary onto every other function. |
| Weight or intensity | Share of practice, visit volume, or a similar measure where available | Lets targeting and account roll-ups distinguish a main site from an occasional one. |
The objection to this design is always the same: it is more complex, and complexity has to be maintained. That objection is fair, and it is why the stewardship section of this article matters as much as the matching section. But the alternative is not simplicity. It is the same complexity, pushed into spreadsheets maintained by field teams and into manual reconciliation before every compliance submission. The complexity does not disappear when you refuse to model it. It just stops being governed.
A useful test for any HCP master. Ask the system to reproduce the affiliation state of a named prescriber as it stood on a date eighteen months ago. If the answer requires restoring a backup, the model is flattened and every historical report the organization has produced is unreproducible.
Three Matching Approaches, Compared Honestly
With affiliation modeling established as the harder half of the problem, the matching question becomes tractable. There are three families of approach in production use, and the honest position is that all three are legitimate and the choice depends on the quality of your input data, the scale of your record volumes, and how much stewardship capacity you actually have.
One thing applies to all three. Comparing every record against every other record is quadratic, so at pharma scale no approach compares all pairs. Every production system reduces candidate pairs first, a step the research literature calls blocking or filtering, then applies its matching logic within those candidate sets.16 A good blocking strategy matters more to real world recall than the choice of matching method, and a blocking key built only on last name plus ZIP code will silently lose every record where the prescriber moved.
Deterministic rules
Deterministic matching applies explicit rules. If the NPI matches, it is the same person. If last name, first initial, date of birth, and state license number all match, it is the same person. Rules are ordered, each produces a match or no match, and there is no ambiguity in the output.
What it is good at. Deterministic rules are fast, cheap to run, and completely explainable. When a steward asks why two records merged, the answer is a rule number. That explainability is worth more in a regulated environment than it usually gets credit for. When the input data carries reliable unique identifiers and has low rates of missing values and errors, deterministic matching performs close to the alternatives. A 2026 comparative study of linkage methods across electronic health record systems reported an average F-score of 97.2% for deterministic rules with run times of under one second per rule.11 A simulation study of linkage strategies concluded that with low rates of missing data and error, deterministic linkage performed not significantly worse than probabilistic.12
Where it fails. Deterministic rules are brittle exactly where HCP data is weakest. A missing NPI, a transposed license number, a maiden name, a hyphenated surname entered two ways, an address recorded as Suite 300 in one system and Ste 300 in another. Each of these makes a rule miss. The failures are silent, because a rule that does not fire produces two separate records rather than an error message. Deterministic estates tend to accumulate duplicates quietly for years.
Probabilistic matching
Probabilistic matching descends from the framework Fellegi and Sunter published in 1969. For each compared field, the method estimates the probability that the field agrees given that the records are a true match (the m-probability) and the probability that it agrees given that they are not (the u-probability). The ratio of those probabilities produces a weight, weights are summed across fields, and the total is compared against two thresholds. Above the upper threshold the pair is a link, below the lower threshold it is a non-link, and in between it is a possible link routed to human review.10
That third outcome is the important one. Probabilistic matching is the only one of the three families that was designed from the start around the assumption that some pairs cannot be decided automatically and must be sent to a person. Everything about production stewardship workflow follows from that design choice.
What it is good at. Probabilistic matching handles the messy middle of HCP data well: partial names, missing fields, address variation, and the ordinary noise of records assembled from claims feeds, congress attendee lists, sample signature capture, and rep entered contacts. The simulation study cited above found that probabilistic linkage uniformly outperformed deterministic linkage in the trade-off between sensitivity and positive predictive value, regardless of data quality, and was clearly the more accurate method in poorer quality data.12 A study of tuberculosis record linkage that compared both approaches reported sensitivity in the range of 87.2% to 95.2% and specificity of 99.8% to 99.9% across the methods it evaluated.13
Where it fails. Probabilistic weights depend on parameter estimates that have to be tuned and periodically retuned. When the mix of source systems changes, the parameters drift and nobody notices until the review queue length changes. Explaining an individual match to an auditor requires explaining a weight calculation rather than pointing at a rule, which is a harder conversation. And the possible link band is only useful if someone works it. A probabilistic system with an unstaffed review queue is a deterministic system with worse explainability.
Machine learning based resolution
Machine learning approaches treat matching as a classification task. Pairs of records are represented as feature vectors, a model is trained on labeled examples of matches and non-matches, and the trained model predicts on unseen pairs. More recent work uses pre-trained language models, which handle the semantic side of matching considerably better than feature engineering: recognizing that “Mass General” and “Massachusetts General Hospital” refer to the same institution, or that two differently worded department names describe the same unit.15
What it is good at. On messy textual attributes, machine learning outperforms rule based and classical approaches. A design space study of deep learning for entity matching found that these methods outperform a strong classical baseline on textual matching tasks, where the records consist of a few attributes that are essentially free text.14 The 2026 EHR linkage comparison reported that its machine learning classification approach reached an F-score of 99.8%, the highest of the methods tested.11 For HCO name matching in particular, where institution names carry abbreviations, legal suffixes, and colloquial forms, this is a genuine advantage.
Where it fails. The requirement is labeled training data, and that requirement is often understated. The same design space study found that the classical baseline significantly outperformed the deep learning models when only a small number of labeled examples were available, with transfer learning and active learning needed to close the gap.14 Labeled HCP pairs do not exist off the shelf. Someone in your organization has to produce them, which means a steward adjudicating thousands of pairs before the model is worth deploying.
The second issue is computational. The 2026 study reported run times for its machine learning approaches ranging from 13 seconds to 16,936 seconds, against under one second per rule for the deterministic approach and 0.1 to 5 seconds for probabilistic.11 That is a four order of magnitude spread, and it changes what a nightly refresh looks like.
The third issue is explainability, and in a regulated commercial environment it is not a minor one. If a merged record produces a transparency report that a physician disputes, “the model scored the pair at 0.94” is a weaker answer than a rule reference or a weight breakdown. Explainable entity matching is an active research area precisely because this gap is real.
Choosing an Approach and Budgeting the Stewardship Effort
The comparison above should make clear that there is no dominant approach. What follows is a direct statement of when each fits and what each demands from the people who maintain it.
| Approach | Right fit when | Stewardship demand | Main failure mode |
|---|---|---|---|
| Deterministic rules | Source records carry reliable NPI or equivalent identifiers, missing rates are low, volumes are large, and every merge must be explainable in one sentence | Low ongoing effort but concentrated: rule authoring and periodic rule review by someone who understands both the data and the business meaning | Silent under-matching. Duplicates accumulate because a rule that fails to fire produces no signal |
| Probabilistic matching | Records arrive from many sources with inconsistent completeness, identifiers are frequently missing, and you have or can build a standing stewardship function | High and continuous: a staffed review queue, service levels for clearing it, and periodic parameter retuning | An unworked review queue, or threshold drift that nobody detects because nobody watches queue volume as a metric |
| Machine learning | Matching depends heavily on messy text such as HCO names and department descriptions, and you can fund an initial labeling exercise plus ongoing model governance | Front loaded and specialized: labeled training pairs, model validation, monitoring for drift, and a documented rationale for regulated use | Model decay after a source system changes, combined with weak explainability when a specific merge is challenged |
The layered pattern most mature estates converge on
In practice, the organizations that get this right rarely pick one. They layer.
Deterministic first pass on strong identifiers
Match on NPI, state license plus state, and other reliable keys. This resolves the large majority of records at effectively no computational expense and with complete explainability. Everything it catches is removed from the harder work downstream.
Probabilistic pass on the residual
Run weighted matching against what the deterministic pass could not resolve. Set the upper threshold conservatively, because a false merge of two prescribers is far more damaging and far harder to unwind than a duplicate that survives another week.
Machine learning where text dominates
Apply learned models specifically to HCO name and address resolution, where semantic variation is highest and where a small labeling exercise produces a large gain. Keep the model scoped to that task rather than replacing the whole pipeline.
Human adjudication on the band in between
Route the undecidable pairs to stewards with enough context to decide, and capture every decision as labeled data. This is the step that turns a stewardship function from a treadmill into an asset, because those adjudications are the training set for step three.
The compounding benefit of capturing adjudications. Every steward decision on a possible link is a labeled pair. Organizations that log those decisions in a structured form build the training data for machine learning resolution as a byproduct of running probabilistic matching. Organizations that let stewards resolve pairs in a user interface without capturing the outcome pay for the same labeling work twice.
Survivorship and the Golden Record
Matching decides which records refer to the same entity. Survivorship decides what the resulting single record says. These are different problems and they fail in different ways, but they are routinely discussed as one thing, which is why survivorship design tends to get about a tenth of the attention it needs.
Once a cluster of records is agreed to describe one prescriber, the cluster typically contains conflicting values for the attributes that matter most: specialty, primary address, degree, credentials, contact details, and affiliation. Survivorship rules decide which value survives into the golden record.
Why single trusted source policies fail
The most common survivorship policy is also the weakest: nominate one source system as authoritative and take everything from it. It is attractive because it is simple to explain and simple to implement. It fails for a reason that is obvious once stated. No source is best at everything.
A third party reference file is usually the best source for credentials, specialty taxonomy, and license status, because that is what the compiler invests in verifying. It is often a poor source for current practice location, because it depends on registry data with the staleness problems described earlier. A field CRM is frequently the best source for current practice location and contact preference, because a representative was physically there last month, and among the worst sources for specialty, because it was entered once at record creation and never revisited. A claims derived feed may be the best source for where a prescriber actually practices by volume, and carries no reliable contact information at all.
A single trusted source policy forces the organization to take that source’s worst attributes along with its best. What works instead is survivorship defined per attribute.
| Attribute | Typical winning source | Rule pattern |
|---|---|---|
| NPI and license | Federal or state registry | Source priority. There is one right answer and the registry holds it. |
| Primary specialty | Verified reference file | Source priority, with a stewardship exception when the field disputes it and provides evidence. |
| Practice address | Field CRM or claims derived location | Most recent verified value, with a recency window after which the value is flagged as unconfirmed rather than trusted. |
| Contact details | Field CRM | Most recent, with a completeness check so that a blank does not overwrite a populated value. |
| Affiliation | No single winner | Do not collapse. Keep all asserted affiliations with source and dates, and derive a primary rather than overwriting. |
| Consent and preference | Consent system of record | Never survivorship driven. Consent is a legal state, not a data quality judgment, and must come from one system. |
The four survivorship rule patterns worth knowing
Source priority
A ranked list of sources per attribute, highest available wins. Simple and explainable. Right for attributes with one objectively correct value that a specific source is responsible for maintaining.
Most recent
The newest value wins regardless of source. Right for attributes that genuinely change, such as address. Dangerous without a completeness guard, because a low quality recent feed can overwrite a good value with a blank.
Quality score
Score each candidate value on completeness, format validity, and verification status, and let the highest score survive even if it is not the newest. More work to configure and considerably more resilient to a bad feed.
Steward override
A human decision that pins a value and survives subsequent refreshes. Essential, and the one that most often goes wrong, because pinned values decay silently unless they carry an expiry and a review trigger.
Three survivorship design decisions that matter more than the rules
Keep the contributing values, not just the winner. A golden record that stores only the surviving value cannot answer the question every dispute eventually raises: what did the other sources say, and when did this value change? Retaining the full set of contributing values with source and timestamp is what makes a golden record defensible rather than merely convenient.
Decide what happens when a steward and a rule disagree. If a steward corrects a specialty and the next refresh overwrites it, the steward stops correcting things. If the override never expires, the record slowly fills with values that were true three years ago. The workable middle is an override that persists but carries a review date, so that a human decision is honored and then revisited rather than honored forever.
Treat unmerge as a first class operation. Every matching system eventually merges two prescribers who are not the same person, usually a father and son with the same name at the same practice, or two physicians with a common surname in the same specialty and city. If unmerging requires a support ticket and a database restore, the organization will avoid unmerging and will instead work around the bad record, which spreads the error into every downstream system that consumed it. Unmerge needs to be a supported, logged, reversible operation with defined propagation to downstream systems.
The consent exception. Consent and communication preference should never be subject to survivorship logic. If two records for the same prescriber carry different consent states, the correct answer is not the newest or the highest priority source. It is the most restrictive state, held pending a documented resolution. Survivorship rules optimize for the most likely value. Consent requires the defensible one.
Where Resolution Errors Become Compliance Findings
Everything above can be argued as a matter of commercial efficiency, and reasonable people can disagree about how much precision is worth funding. Transparency reporting removes that latitude, because the output of the customer master becomes a public record with a named physician attached to it.
The scale of what gets published
The Physician Payments Sunshine Act requires applicable manufacturers of covered drugs, devices, biologicals, and medical supplies to report payments and other transfers of value made to covered recipients, and CMS publishes that data through the Open Payments program.2 This is not a small file. In its Fiscal Year 2025 Report to Congress, CMS reported that for Program Year 2024, reporting entities collectively submitted $13.18 billion in publishable payments and ownership and investment interests, comprising 16.16 million records attributable to 651,977 physicians, 338,340 non-physician practitioners, and 1,288 teaching hospitals. That data was published on June 30, 2025.1 Across all active program years, the published data set covers 88.25 million records.4
Four ways entity resolution errors become reporting errors
Wrong person. A false merge attributes a payment to the wrong physician. The record publishes under a name that never received the payment. This is the failure mode that turns into a dispute, and it is the direct output of a matching threshold set too loosely.
Split identity. A missed match splits one physician’s interactions across two records. The published totals for that physician are understated in one record and understated in the other, and internal spend monitoring against a per physician threshold never triggers because neither half crosses it. Under-matching is quieter than over-matching and, for compliance monitoring, often more dangerous.
Wrong entity type. A payment made to an institution gets attributed to an individual, or the reverse. The EFPIA Disclosure Code addresses this directly with a non-duplication rule: where a transfer of value is made to an individual HCP indirectly through an HCO, it is disclosed once, and to the extent possible on an individual named basis.5 Applying that rule correctly requires knowing whether the recipient organization is an HCO in scope and whether the individual behind the transaction is identifiable. Both are entity resolution questions.
Wrong affiliation at the time of the interaction. This is the failure mode that follows directly from the flattening error described earlier, and it is the one that a good matching engine will not save you from. If the master holds only current affiliation, and reporting runs months after the event, the system attributes an interaction to wherever the prescriber is now rather than where they were when it happened. In a year when 82% of physicians are employed by institutions and practice acquisitions run at scale,7 a meaningful share of a year’s interactions involve a prescriber who changed affiliation mid-year.
Why the review window does not rescue you
Open Payments includes a pre-publication review and dispute period. For Program Year 2024, that window ran from April 1 to May 15, 2025, followed by an additional 15 day correction period for reporting entities.1 Covered recipients may review data attributed to them and, if they believe a record is inaccurate or incomplete, initiate a dispute and work directly with the reporting entity toward a resolution.3
Three features of that process are worth understanding before treating it as a safety net. First, covered recipient review is voluntary, and CMS itself notes this while encouraging review to help ensure accuracy.1 Most physicians do not review their records, which means most attribution errors are never surfaced by the person best placed to spot them. Second, CMS does not mediate disputes. Resolution happens between the covered recipient and the reporting entity, which puts the burden of investigation, correction, and resubmission on the manufacturer’s compliance and data teams inside a short window.3 Third, CMS holds authority to impose civil monetary penalties for late, inaccurate, or incomplete reporting.1
The operational point. A dispute arriving in April is a data investigation with a hard deadline. If the customer master cannot show which records were merged, when, on what evidence, and what the affiliation state was on the date of the interaction, the team cannot resolve the dispute inside the window. The audit trail that makes a dispute answerable has to exist before the dispute arrives, which means it is a design requirement of the master, not a reporting feature.
The European picture adds a consent dimension
The EFPIA Disclosure Code requires member companies to disclose transfers of value to HCPs and HCOs, and directs that where possible companies should identify and publish at the individual HCP level rather than the HCO level, provided this can be done with accuracy and consistency and in compliance with applicable law.5 European data protection law makes individual named disclosure dependent on the HCP’s consent, which introduces a second identity requirement: the consent has to be reliably tied to the right individual.6
That combination is unforgiving. A false merge in a European market does not only misattribute a payment. It can attach one physician’s consent state to another physician’s disclosure, which is a data protection issue on top of a transparency issue. It is also why the survivorship guidance above treats consent as outside survivorship logic entirely.
For organizations reporting in both the United States and Europe, the two regimes pull in the same direction on identity even though the mechanics differ. Both require that the entity named in a disclosure is the entity that received the value, that institution level and individual level transfers are distinguished correctly, and that the attribution reflects the relationship as it stood at the time.
Measuring Whether Resolution Is Actually Working
Most organizations measure the wrong things about their customer master. Record counts, duplicate rates, and completeness percentages are easy to produce and tell you almost nothing about whether resolution is correct. A master with a low duplicate rate might be over-merging. A master with high completeness might be full of confidently wrong addresses.
The measures worth reporting to a governance forum fall into four groups.
Match quality, measured against truth
Precision and recall on a maintained gold standard sample. This requires someone to build and periodically refresh a set of record pairs whose true match status has been established by human review, then run the pipeline against it. It is unglamorous work and it is the only way to know whether a threshold change helped or hurt. Report the two numbers separately, never as a single accuracy figure, because the trade-off between them is exactly the decision the governance forum needs to make.
Affiliation currency
Distribution of the age of affiliation records: what share of active affiliations were verified in the last 90 days, the last year, and more than a year ago. This is the single most useful metric for the problem this article is about, and almost nobody reports it. It tells the organization how much of its territory alignment and spend attribution rests on relationships nobody has confirmed recently.
Stewardship throughput
Review queue volume, average age of an open item, and clearance rate. A growing queue is the earliest available signal that something upstream changed: a new source, a feed format change, or parameter drift. Treating queue growth as a capacity problem rather than a diagnostic signal is a common and expensive mistake.
Downstream trust
Volume of manual corrections submitted by field teams, and the share of field users who report working from a local list rather than the master. Both are measures of whether the organization believes its own data. A technically sound master that the field routes around has failed at the only thing that matters.
One question worth asking every quarter. Of the prescribers who received a transfer of value in the last quarter, what share had an affiliation change during that same quarter, and were those interactions attributed to the affiliation in place on the interaction date? If nobody can answer, the reporting process is relying on assumptions that have not been checked.
A Practical Sequence for Fixing an Existing Estate
Very few organizations get to build this from nothing. The realistic situation is an existing master, a matching configuration nobody fully remembers designing, several source feeds of varying quality, and a field organization that has quietly stopped trusting the output. The sequence below is ordered by what produces the most durable improvement per unit of effort.
Establish a gold standard sample before changing anything
Have stewards adjudicate a few thousand record pairs, stratified across easy and hard cases, and freeze the result. Without this, every subsequent change is a guess and no improvement can be demonstrated. This is the highest value week of work available and it is almost always skipped.
Measure current precision and recall, and publish the numbers
Run the existing configuration against the gold standard. The result is usually worse than expected on recall and better than expected on precision, because most configurations were tuned conservatively at go-live and never revisited. Publishing the numbers converts an argument about opinions into a discussion about a trade-off.
Separate affiliation from the HCP record
This is the structural change and the one with the longest tail of downstream work, which is why it should start early. Build the relationship entity with type, effective dates, source, and confidence. Backfill what history is recoverable and mark the rest as unknown rather than inventing it.
Rewrite survivorship per attribute
Replace any single trusted source policy with attribute level rules, add completeness guards so blanks cannot overwrite populated values, and give steward overrides an expiry with a review trigger. Document the winning source per attribute in a table the business can read.
Staff the review queue and capture every decision
Set a service level for clearing possible links and log every adjudication in a structured form. This makes probabilistic matching work as designed and builds the labeled data that any future machine learning step will need.
Add learned matching only where text dominates
Once labeled pairs exist, apply machine learning to HCO name and address resolution first, where the gain is largest and the regulatory exposure of an individual level error is lowest. Keep it scoped, validated, and monitored rather than letting it replace the whole pipeline.
Two notes on sequencing. Step three is the one most likely to be deferred because it touches downstream consumers, and deferring it is the reason most remediation efforts produce a temporary improvement in duplicate rates and no improvement in the numbers the business actually disputes. And step one is genuinely a prerequisite. An organization that begins at step four is tuning rules without any way to know whether it made things better.
What good looks like at the end. A named prescriber’s identity can be traced to the records that formed it and the evidence for each merge. Their affiliations are held as dated relationships with sources, not as one overwritten field. Each attribute in the golden record can be traced to the source that won and the rule that decided it. And a compliance question about a payment made fourteen months ago can be answered from the system in an afternoon rather than reconstructed from exports.
Conclusion
HCP and HCO entity resolution has stayed difficult not because matching algorithms are inadequate but because the data model most organizations use was designed for a world of independent physician practices that no longer exists. With more than four in five US physicians now employed by hospitals or corporate entities, the relationship between a person and an institution has become the central fact about a prescriber, and it is many to many, typed, dated, and constantly changing. Storing it as a single current employer field is a modeling decision that quietly limits everything built on top of it.
The matching approach question has a defensible answer, and it is usually a layered one: deterministic rules on strong identifiers, probabilistic matching on the residual, learned models where free text dominates, and human adjudication on the band in between with every decision captured. Survivorship deserves the same rigor, defined per attribute rather than by nominating one source as authoritative, with contributing values retained and consent held outside the rules entirely. The consequence of getting these wrong is not an inconvenient report. It is a published transparency disclosure with the wrong physician’s name on it, inside a program that published 16.16 million records in a single program year and gives the manufacturer, not the regulator, responsibility for resolving the dispute.
Sakara Digital works with pharma and biotech organizations on the data foundations that commercial, medical, and compliance functions all depend on, including customer master design, affiliation modeling, and the stewardship operating model that keeps them current. If you are working through where your HCP and HCO resolution is breaking down and want an independent perspective on what to change first, we are happy to have that conversation.
For Further Reading
For Further Reading
- Master Data Management for Life Sciences: Creating a Single Source of Truth Across Global Operations
- MDM Vendor Selection for Mid-Cap Pharma: A 2026 Comparison Framework
- Choosing the Right Data Governance Model for Pharma: Centralized, Decentralized, or Federated?
- Digital Transformation Roadmap for Pharma Commercial Operations
- Fixing the Foundations: How Pharma Can Remediate Common Data Quality Issues
References & Sources
- Centers for Medicare & Medicaid Services. “Report to Congress, Fiscal Year 2025: Annual Report on the Open Payments Program.” March 2026. https://www.cms.gov/files/document/open-payments-fy-2025-report-congress.pdf
- Centers for Medicare & Medicaid Services. “What is Open Payments?” CMS Key Initiatives. https://www.cms.gov/priorities/key-initiatives/open-payments
- Centers for Medicare & Medicaid Services. “Review and Dispute for Open Payments Covered Recipients.” https://www.cms.gov/openpayments/program-participants/covered-recipients/review-and-dispute
- Centers for Medicare & Medicaid Services. “About the Open Payments Data.” Open Payments Data website. https://openpaymentsdata.cms.gov/about
- European Federation of Pharmaceutical Industries and Associations. “EFPIA HCP/HCO Disclosure Code.” https://www.efpia.eu/media/25837/efpia-disclosure-code.pdf
- European Federation of Pharmaceutical Industries and Associations. “The EFPIA Disclosure Requirements for Healthcare Professionals: Your Questions Answered.” June 2025. https://www.efpia.eu/media/i0yiz30d/disclosure-qa-june-2025.pdf
- Physicians Advocacy Institute and Avalere Health. “PAI-Avalere Health Report on Physician Employment Trends and Practice Acquisitions: 2018-2026.” https://www.physiciansadvocacyinstitute.org/PAI-Research/PAI-Avalere-Health-Report-on-Physician-Employment-Trends-and-Practice-Acquisitions-2018-2026
- Centers for Medicare & Medicaid Services. “Online Provider Directory Review Report” (Round 3). November 28, 2018. https://www.cms.gov/medicare/health-plans/managedcaremarketing/downloads/provider_directory_review_industry_report_round_3_11-28-2018.pdf
- CAQH. “The Hidden Causes of Inaccurate Provider Directories: How administrative burdens on physician practices may be undermining the accuracy of provider directories.” https://www.caqh.org/hubfs/43908627/drupal/explorations/CAQH-hidden-causes-provider-directories-whitepaper.pdf
- Fellegi, I.P. and Sunter, A.B. “A Theory for Record Linkage.” Journal of the American Statistical Association, 64(328), 1969, pp. 1183-1210. https://courses.cs.washington.edu/courses/cse590q/04au/papers/Felligi69.pdf
- Almadani, O., Albogami, Y., Alrwisan, A. “Linking Electronic Health Records for Multiple Sclerosis Research: Comparative Study of Deterministic, Probabilistic, and Machine Learning Linkage Methods.” JMIR Medical Informatics, 2026. https://pmc.ncbi.nlm.nih.gov/articles/PMC12872214/
- Zhu, Y. et al. “When to conduct probabilistic linkage vs. deterministic linkage? A simulation study.” Journal of Biomedical Informatics, 2015. https://pubmed.ncbi.nlm.nih.gov/26004791/
- “Accuracy of probabilistic and deterministic record linkage: the case of tuberculosis.” Revista de Saude Publica, 2016. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4988803/
- Mudgal, S. et al. “Deep Learning for Entity Matching: A Design Space Exploration.” Proceedings of SIGMOD 2018. https://pages.cs.wisc.edu/~anhai/papers1/deepmatcher-sigmod18.pdf
- Li, Y., Li, J., Suhara, Y., Doan, A., Tan, W. “Deep Entity Matching with Pre-Trained Language Models.” arXiv:2004.00584. https://arxiv.org/abs/2004.00584
- Papadakis, G., Skoutas, D., Thanos, E., Palpanas, T. “Blocking and Filtering Techniques for Entity Resolution: A Survey.” arXiv:1905.06167. https://arxiv.org/abs/1905.06167
- Pharmaceutical Commerce. “Master data management (MDM) takes center stage.” https://www.pharmaceuticalcommerce.com/view/master-data-management-mdm-takes-center-stage
- Yaraghi, N. et al. “Facilitating accurate health provider directories using natural language processing.” BMC Medical Informatics and Decision Making. https://pmc.ncbi.nlm.nih.gov/articles/PMC6448184/
- Centers for Medicare & Medicaid Services. “Data Dissemination: National Plan and Provider Enumeration System (NPPES) NPI files.” https://www.cms.gov/medicare/regulations-guidance/administrative-simplification/data-dissemination








Your perspective matters—join the conversation.