The Question Your Policy Does Not Answer

Picture a scientist in clinical development at 4:40 on a Thursday. She has a 60-page document open. She has to summarize it for a governance meeting the next morning. There is an approved AI assistant on her desktop and a browser tab open to a general-purpose model she uses at home. The document has a footer that says Confidential. She opens the company data classification policy, searches it, and finds that Confidential information must be encrypted at rest and in transit, shared only on a need-to-know basis, and not disclosed externally without an executed confidentiality agreement.

None of that answers her question. She is not storing it, sending it, or printing it. She wants to know whether pasting it into a text box is allowed, and if so, into which text box. The policy does not have a word for what she is about to do.

This is not a training failure. It is a design failure, and it is nearly universal. Classification schemes in life sciences were written against a threat model of lost laptops, misdirected email, and departing employees carrying files. The tiers were designed to drive storage controls, access reviews, and retention schedules. Most of them predate the arrival of tools that accept arbitrary text and return something useful, and that may retain, log, or learn from what they receive. The control set has not caught up with the interaction pattern.

Two failure modes, both expensive

When a policy does not answer the question, people resolve it themselves, and they resolve it in one of two directions.

The first is over-restriction. Quality or legal, reading the absence of a rule as prohibition, issues a memo saying no company information may be entered into any AI tool. This is easy to write and impossible to hold. It does not stop the work; it moves the work off the managed path. Traffic analysis of enterprise AI usage found that roughly a third of ChatGPT activity and well over half of Claude and Perplexity activity ran through personal accounts rather than corporate ones.1 A personal account is the worst of every option: no contract, no retention control, no logging, and no ability to answer an auditor’s question about what left the building.

The second is under-restriction. In the absence of a rule, a reasonable person applies a reasonable heuristic, which is usually some version of “this does not feel that sensitive.” That heuristic is fine for a meeting agenda and serious trouble for an unblinded listing. It is also invisible. Nobody files a deviation for pasting the wrong document into a chat window.

39.7% of AI interactions in enterprise environments involve sensitive data, according to traffic analysis across corporate deployments1
97% of breached organizations that experienced an AI-related security incident said they lacked proper AI access controls2
63% of organizations studied had no AI governance policy in place to manage AI use or unsanctioned tools2

Why four tiers is usually three too many

The standard scheme in large pharmaceutical companies has four levels. The names vary: Public, Internal, Confidential, Restricted; or Public, Internal Use Only, Company Confidential, Highly Confidential. Ask ten people in the same building to name all four in order and you will get a wide spread of answers. Ask them to state the practical difference between the top two and you will usually get a pause.

That pause is the whole problem. A classification scheme is used at the moment of action, from memory, by somebody who is trying to finish something else. If it cannot be recalled and applied in the time it takes to decide whether to hit paste, it is not a control. It is a document that exists so an auditor can be shown it.

Practitioners who run these programs report the same failure pattern from the inside: labels multiply over time, exceptions accumulate, definitions blur, and users faced with too many options pick whichever one gets them back to work fastest.3 The recommended range that keeps showing up in that literature is three to five levels. In our experience with pharmaceutical and biotech clients, three is the number that survives contact with a laboratory.

Three Tiers, Named for What They Mean

The scheme below uses three tiers. Each name states what the tier is, so the name itself does most of the teaching. Each has a one-line test that produces an answer rather than a debate.

TIER 1

Public

Test: Is this already outside the company, or formally cleared to go outside? Published papers, approved labeling, press releases, registry entries, posted job descriptions, your own website. If it is not already out and nobody has cleared it, it is not Public.

TIER 2

Internal

Test: If a competitor read this, would you be annoyed, or would you have a problem? Annoyed is Internal. Most SOPs, most training material, most meeting notes, most project plans, most validation deliverables.

TIER 3

Protected

Test: If this got out, would you have to tell somebody? A regulator, a partner, a patient, an exchange, an insurer, or a lawyer. If the answer is yes, it is Protected.

DESIGN NOTE

Consequence, not content

All three tests ask what happens next, not what the document contains. Content tests generate arguments about whether a given field is sensitive. Consequence tests generate answers, because most people know instinctively whether an exposure would trigger a phone call.

Why the “would you have to tell somebody” test works

The top tier is where classification schemes usually break down, because organizations try to define it by subject matter. They write a list: formulations, clinical results, patient data, deal terms, and so on. Lists are always incomplete, and they push people into a matching exercise they are not equipped to perform.

The notification test is better because it maps directly onto the thing you are actually trying to prevent. Regulatory frameworks are built around notification duties. A personal data breach carries a reporting obligation. Selective disclosure of material information about a listed company carries one. A breach of a confidentiality agreement usually carries a contractual notice requirement. An unblinding event carries a documented reporting path under good clinical practice. If a disclosure would force a notification, the harm is real, external, and not something you can fix internally. That is exactly the population you want in the top tier.

The test is also stable over time. Subject-matter lists go out of date every time the portfolio changes. The question “would I have to tell somebody” does not.

What formal frameworks contribute, and where they stop

None of this is a rejection of the established standards. ISO/IEC 27002 control 5.12 requires that information be classified according to confidentiality, integrity, availability, and relevant interested-party requirements, and control 5.13 requires a labeling procedure that reflects the scheme.45 The US federal approach in FIPS 199 and NIST SP 800-60 assigns low, moderate, or high impact to each of confidentiality, integrity, and availability, and applies a high-water-mark rule where the highest single rating drives the overall category.67 NIST is currently revising SP 800-60, with a public draft of Revision 2 released for comment.8

These frameworks tell you how to think about categorization at the level of a system or an information type. They do not tell a scientist at 4:40 on a Thursday what to do with the document in front of her. That last step, from category to a decision about a specific tool, is the part every company has to build itself, and it is the part almost nobody has written down.

The Four Places Data Can Go

A tier on its own is still not an answer. The scheme becomes usable when you pair it with a short, closed list of destinations. Four is enough to cover the real world.

DestinationWhat it isWhat happens to your data
D1: Open external A consumer or free-tier service. Personal sign-in, click-through terms, no negotiated agreement. Terms typically permit retention and may permit use for service improvement. You have no enforceable control and no usable record of what was sent.
D2: Contracted external The same class of service under a negotiated enterprise agreement, with a data processing agreement, a no-training commitment, and defined retention. Data leaves your boundary but under terms you can enforce, with logging you can request and a deletion path you can name.
D3: Your tenant A model running inside your own cloud subscription, private endpoint, or on-premises infrastructure. Data stays inside your security and residency boundary. Access is governed by your identity controls, not the vendor’s.
D4: No model The information stays in its system of record. Only named humans with existing entitlements see it. Nothing is submitted anywhere. This is the correct answer for a narrow but real set of information.

Classify by the sign-in, not by the logo. The single most common error we see is treating the destination as a property of the vendor brand. It is not. The same provider will offer a free consumer tier (D1) and a contracted enterprise tier (D2), and may also offer a deployment inside your own cloud subscription (D3). Three destinations, one logo. Train people to look at which account they are signed in with, because that is what determines the terms that apply.

The matrix

With three tiers and four destinations, the whole scheme fits in one table that a person can hold in their head.

TierD1 Open externalD2 Contracted externalD3 Your tenantD4 No model
PublicAllowedAllowedAllowedAllowed
InternalNot allowedAllowedAllowedAllowed
ProtectedNot allowedOnly with a recorded, named exceptionAllowed, subject to any special category flagAllowed

Three rows. Four columns. One sentence of explanation per cell if anyone asks. That is a scheme somebody can actually follow, and more importantly it is a scheme you can implement in a gateway as a routing rule rather than as an aspiration.

Notice what the matrix does for the middle tier. Internal information, which is the overwhelming majority of what a pharmaceutical company holds, is cleared for contracted external tools. That clearance is where nearly all of the practical value of AI in a regulated business is found. A scheme that puts everything in the top tier does not produce safety. It produces a refusal, and refusals get routed around.

This article is about classifying the data. The separate question of what the vendor agreement should say, how prompt leakage and model memorization work, and how to hold a supplier accountable for the behavior of a generative system, is covered in our companion piece on generative AI risk controls. Treat the two as a pair: the classification scheme decides what may be sent, and the vendor controls decide whether the receiving end is trustworthy enough to be a D2 at all.

Special Categories That Cut Across Tiers

Some information carries a duty that does not follow from how sensitive it feels. These are the categories that cause a well-meaning scheme to go wrong, because a document can look ordinary and still be governed by a statute, a license, or somebody else’s contract.

The design decision that matters here is this: special categories are flags, not tiers. A flag attaches to a document alongside its tier and adds named handling rules on top. Making them tiers is how a three-tier scheme becomes a nine-tier scheme within eighteen months.

Personal data

Personal data adds a legal basis question and a location question that the tier alone does not answer. It is worth being precise about a point that is widely misunderstood. The European Data Protection Board’s Opinion 28/2024 concluded that AI models trained on personal data cannot automatically be treated as anonymous, and that demonstrating anonymity requires showing the likelihood of identification is negligible for every data subject whose data contributed to the model.9 Legal analysts reading the opinion drew the same practical conclusion: anonymity has to be evidenced on a case-by-case basis rather than assumed from the architecture.11 For a classification scheme, that means “we de-identified it” is a claim that needs evidence, not a checkbox.

Where personal health information is involved, the two recognized paths under the HIPAA Privacy Rule remain removal of the enumerated identifiers or a documented expert determination that the re-identification risk is very small.10 Both are defensible. Neither happens by accident, and a dataset that has merely had the name column deleted has done neither.

Unblinded clinical data

This is the category where the tier-and-destination logic on its own gives the wrong answer, and it is worth understanding why. For every other category, the risk is that information travels outside the company. For unblinded data from an ongoing blinded study, the risk is that it travels inside the company, to people who are supposed to remain blinded. A model running in your own tenant does not help. If the output is visible to a study team member, the harm has already occurred, and the integrity of the study is affected regardless of where the compute ran.

The rule for this flag is therefore not about destination at all. It is about audience. Unblinded data goes to D4 unless there is a specific, documented, access-controlled analysis environment whose users are already unblinded.

Safety data

Individual case safety reports and adverse event narratives carry three things at once: personal data, often special category health data, and a reporting duty with a clock attached. The classification consequence people miss is the clock. If a tool reads safety narratives and surfaces something a qualified person would recognize as a reportable event, awareness has been created. Whether and when that starts a reporting timeline is a question for your pharmacovigilance function, and it should be answered before the tool is switched on rather than during the first inspection that asks about it.

Material nonpublic information

Topline results before an announcement, an unexpected safety signal, a pending regulatory action, an unannounced transaction, a manufacturing failure that will affect supply. In a listed company these are governed by securities law, and the governing question is not where the data is processed but who can see the output.

The handling rule for this flag is an audience rule, like unblinding. It restricts who may receive output, and it prohibits any configuration in which a shared retrieval index could surface the content to a user who is outside the information barrier. This is the specific failure mode to design against when one enterprise assistant serves commercial, clinical, and corporate development from the same knowledge base. Each of those populations is meant to be separated from the others, and a shared index reconnects them without anybody deciding to.

Third-party data under license

Prescription data, claims data, literature databases, reference standards, competitive intelligence subscriptions, curated chemistry sets. Every one of these arrives under a license that says what you may do with it. Most of those licenses were drafted before generative models were a normal part of the workflow. They grant internal use, and they say nothing about whether submitting content to a hosted model is internal use.

Silence in a license is not permission. A grant of “internal business use” was written when internal meant inside your firewall. Submitting licensed content to an external service, even a contracted one, is at minimum a question for legal and at worst a breach that terminates the license and voids the analysis built on it. Flag licensed data at ingestion, not at the point somebody wants to use it in a hurry.

Data under a confidentiality agreement you did not write

This is the flag that surprises people most, and it is everywhere in pharmaceutical operations. Incoming confidential disclosure agreements from potential partners. Contract research organization master service agreements. Supplier quality agreements. Site agreements. In each case the confidentiality terms were drafted by the other side or negotiated to their template, and they frequently say something stricter than your own policy: no disclosure to any third party, without exception, absent prior written consent.

Read literally, that language covers a hosted model under an enterprise agreement, because the model provider is a third party. Whether a court would agree is a question nobody wants to be the test case for. The workable position is to treat incoming-agreement material as carrying a flag that removes D1 and D2 entirely and permits only D3 or D4, unless somebody has actually read the agreement and confirmed otherwise in writing.

The regulatory overlay on commercially confidential information

Two regulatory regimes are worth knowing because they shape what your organization treats as confidential in the first place. Under 21 CFR 20.61, trade secrets and confidential commercial or financial information submitted to FDA are not available for public disclosure, with trade secret defined narrowly as a commercially valuable plan, formula, process, or device used in making a product.12 In Europe, EMA Policy 0070 publishes clinical data proactively, and applicants who want commercially confidential information redacted must justify each proposed redaction, with personal data anonymized before publication.13 Independent review of published Policy 0070 packages found the resulting reports retained substantial utility for secondary research, which tells you how narrow the accepted redactions tend to be.14

Neither of these regulations tells you how to classify a document internally. What they do is set a useful outer boundary. Information you have already accepted will be published under Policy 0070 is a strange thing to be treating as your most sensitive category two years earlier.

The Ten-Second Decision

A scheme is only as good as the decision it produces under time pressure. The flow below is written to be completed from memory, standing up, in about ten seconds. Ten seconds is a design constraint on the policy, not a performance target for the person. If a competent scientist cannot get through it in ten seconds, the scheme is wrong and needs to be simplified.

1

Is it already public?

Already published, already on a registry, already in approved labeling, already on your website. If yes, the tier is Public and you are finished. Any destination is permitted. This branch comes first because it is the cheapest and it resolves a surprising share of real requests.

2

Does it carry a flag?

Personal data, unblinded, safety, material nonpublic, licensed, or under somebody else’s confidentiality agreement. If any flag applies, follow the flag’s rule first. Flags override the tier matrix. They never relax it.

3

If it got out, would you have to tell somebody?

A regulator, a partner, a patient, an exchange, an insurer, or a lawyer. If yes, the tier is Protected. Your tenant or no model, with contracted external available only under a recorded exception.

4

Otherwise it is Internal.

Contracted external or your tenant. Not the open consumer tier. This is where most of the work is and where most of the value is.

5

If you cannot decide, treat it as Protected and ask.

There has to be a named route for “I do not know” that returns an answer the same day. Without one, uncertainty resolves as a guess, and guesses trend permissive under deadline.

Notice the ordering. Public is checked first because it terminates immediately. Flags are checked second because they override everything downstream. The consequence test comes third, and the default falls out at the end. Reversing any of these makes the flow slower without making it safer.

Defaults, Labeling, and the Combination Problem

The unlabeled majority

Whatever your policy says, most of the documents in your environment carry no label at all. A file share that has been accumulating since 2009 contains hundreds of thousands of files created before anybody thought about classification, by people who have left, in folder structures that no longer reflect the organization. This is the normal state, not a sign of a badly run company.

So the most consequential single line in the whole scheme is the one that says what an unlabeled document is. And the answer organizations reach for, because it avoids friction, is the most permissive tier available. Unlabeled means Internal or, in the worst cases, unlabeled means unrestricted.

The default must never be the most permissive tier. A permissive default means every document nobody got around to labeling is automatically cleared for the widest set of destinations. It converts an administrative gap into an authorization. And because the largest population of documents is exactly the population nobody looked at, the default is not an edge case. It governs the bulk of your data.

Our recommendation is a two-part default. Unlabeled defaults to Internal, which permits contracted and in-tenant tools and blocks the open consumer tier. And Public is never a default, in any circumstance, for any repository. Public has to be an affirmative act by a named person, because publication is the one tier you cannot reverse. Everything else you can tighten later. Once something is genuinely out, it is out.

The trade secret problem with unenforced marking

There is a legal reason to care about labeling discipline that has nothing to do with AI, and it deserves a place in this conversation because AI programs are often what finally forces the issue.

Trade secret protection turns on whether the owner took reasonable measures to keep the information secret. Counsel who litigate these cases note that confidentiality marking is a double-edged practice: courts have declined to find a duty of secrecy where a company’s own policy required confidential material to be marked and the material at issue was not marked.15 A policy that mandates labeling and is broadly ignored can be worse than a policy that describes protection through access controls and agreements instead.

The practical reading for a classification program is to be careful about writing absolute marking requirements you cannot meet. Write what you will actually do, apply it to new work, and describe your protections in terms of the controls you genuinely operate.

Labeling at the source or inspecting at the point of use

There are two places to attach a classification, and they fail in opposite directions.

AT THE SOURCE

Label on creation

The label is applied when the document is made, travels with the file as metadata, and can be read by downstream controls. Information protection platforms support a default label for new documents and can require a label before a file is saved or an email sent.1620 Strong for governance and defensibility. Weak on coverage, because it does nothing for the back catalog.

AT THE POINT OF USE

Inspect on submission

A gateway or endpoint control inspects the content at the moment somebody tries to submit it and decides based on what it finds. Data loss prevention policies work this way, matching sensitive information types and applying actions.17 Strong on coverage, including unlabeled legacy content. Weaker as evidence, because the decision is inferred rather than declared.

The answer is not to choose. It is to assign each one the job it is good at. Use source labeling for everything created from the day the scheme goes live, with a default label and mandatory labeling where the platform supports it. Use point-of-use inspection as the safety net for the back catalog. Do not launch a project to retroactively label twenty years of a file share. Those projects consume a year, deliver partial coverage, and are out of date on completion. Inspect the old material where it tries to leave instead.

When two ordinary documents make one sensitive one

The hardest case in any classification scheme is the one where the parts are fine and the whole is not. This has a name in the privacy and intelligence literature: the mosaic effect, where fragments that are individually harmless enable a sensitive inference once combined.18

Pharmaceutical examples are easy to generate. A site list is Internal. An enrollment table by month is Internal. Together they tell you which sites are failing, which is competitively useful and commercially sensitive. A de-identified subject listing may be genuinely de-identified in a large indication and re-identifiable in a rare disease program with forty patients worldwide. A supplier list is Internal. A supplier list joined to a shortage risk assessment is a different animal.

Attempts to solve this by enumerating dangerous combinations fail immediately, because the combinations are unbounded and the list is out of date the week it is published. The workable approach is a single rule stated in the scheme itself.

The aggregation rule. Classification attaches to the request, not only to the file. When you assemble content from more than one source into a single submission, a knowledge base, or a retrieval index, you classify the assembled package, and the package inherits the highest tier and all flags of its parts. If the combination is more sensitive than any of its parts, classify it higher. The person assembling the package makes the call, and they make it before the assembly is used, not after.

This is the same high-water-mark logic FIPS 199 applies across security objectives, applied instead across sources.6 It is one sentence, it is memorable, and it puts the judgment with the only person who knows what went into the package. It matters most for retrieval-augmented deployments, where a corpus is assembled once and then queried by many people who never see what is in it. The corpus is the classified object, and it needs an owner.

Five Worked Examples From Real Pharma Work

A scheme is easy to agree with in the abstract and hard to apply. The five documents below appear in every pharmaceutical company, and running the scheme against them exposes where the real judgment is.

DocumentTierFlagsPermitted destinations
Final protocol, ongoing study Internal Third-party agreement, if partner content is included D2 and D3, or D3 only where a partner agreement flag applies
Unblinded subject listing, ongoing blinded study Protected Unblinded, personal data D4, or a named unblinded analysis environment
Supplier audit report Protected Third-party confidentiality agreement D3, D4
Draft response to a regulatory information request Protected Possible material nonpublic information D3, D4
SOP for a marketed product Internal None, unless it contains trade secret process parameters D2, D3

The protocol

Protocols are routinely over-classified, and it is worth asking why. A final protocol for an ongoing study is a serious document, but a great deal of its content is already public. The registry entry carries the design, the primary endpoint, the population, the site list, and the estimated enrollment. What is not public is the detailed statistical analysis plan, any unpublished background pharmacology, and whatever a partner contributed.

Internal is the right tier for the ordinary case. The judgment moves when the protocol contains a genuinely novel endpoint or adaptive design that has not been disclosed, in which case a competitor reading it is a problem rather than an annoyance, and the tier moves to Protected. The judgment also moves when a development partner contributed content, in which case the incoming agreement flag applies and the permitted destinations narrow regardless of the tier.

The unblinded listing

This is the clearest case in the set and the one most likely to be handled wrongly, because the tooling gives the wrong intuition. A subject-level listing with treatment assignment from an ongoing blinded study is Protected, and it carries both the unblinded flag and the personal data flag. Its permitted destination is D4.

The point that gets missed is that running the model in your own tenant does not solve it. In-tenant deployment addresses whether data leaves the company. Unblinding is about who inside the company sees it. If the assistant is available to a study team member, the fact that it runs on your own infrastructure is irrelevant to the harm. The only acceptable configuration is one where the audience is already unblinded and that access is documented.

The supplier audit report

This one catches people out because it feels like internal quality documentation. It is not. It is a document about a third party, containing findings that could damage that third party’s other customer relationships, and it almost always exists under a quality agreement or supplier confidentiality agreement your company did not draft.

The tier is Protected, because disclosure would require telling the supplier. The flag is third-party confidentiality, which removes contracted external destinations unless somebody has read the agreement and confirmed that submission to a processor is permitted. In practice most companies settle on D3 and D4 for the whole category, which is the correct conservative answer, and it is a reasonable place to use in-tenant tooling for genuinely useful work such as trending findings across a supplier base.

The draft regulatory response

A draft response to an agency information request is Protected for two independent reasons. First, it may carry material nonpublic information. The existence of the request, its subject, and how close the answer is to satisfactory can all be material to a listed company. Second, the draft reveals your regulatory position and your internal assessment of your own weaknesses, in a form far more candid than the final letter.

There is a common misconception worth correcting here. People point to 21 CFR 20.61 and conclude that because FDA will not disclose their confidential commercial information, the material is somehow protected in general.12 That regulation governs the agency’s disclosure obligations. It says nothing about your own handling, and it offers no protection at all against a disclosure you caused yourself.

The marketed product SOP

This is the example that matters most for the value case, and it is the one companies most often get wrong in the restrictive direction. A standard operating procedure for a marketed product describes how your organization performs a task. It is procedure, not product. Much of it reflects regulatory expectations that are themselves public, expressed in your house style.

Internal is right, and the practical consequence is significant. Clearing SOPs for contracted external tools opens the highest-volume, lowest-risk AI use cases in a quality organization. Comparing versions across sites. Finding conflicts between two procedures that both claim authority over the same step. Drafting a first pass at a revision. Checking that a procedure actually reflects a change control approved eight months ago. This is where a large share of the real return in a pharmaceutical quality function is found, and a classification scheme that blocks it has failed on the value side even if it never leaks anything.

The exception is narrow and real. An SOP that embeds process parameters constituting a trade secret, in the sense 21 CFR 20.61 uses the term, moves to Protected. That is a small number of documents in a small number of areas, and it should be identified deliberately rather than assumed across the whole document class.

Read the pattern in the table. Two of five documents are Internal and cleared for contracted tools. That ratio is roughly right for a real pharmaceutical environment, and it is the test of whether your scheme is calibrated. If you run this exercise across twenty of your own documents and everything comes out Protected, you have not built a control. You have built a policy that says no, which people will read as advice rather than a rule.

Enforcement: Training, Tooling, or a Gateway

A scheme nobody enforces is a scheme that describes intentions. There are three enforcement mechanisms available, they differ enormously in strength and in what they produce as evidence, and most organizations should end up using all three in a specific order.

Training

Training is the cheapest, the fastest to deploy, and the only mechanism that reaches a person working on an unmanaged device or a personal phone. It is also the weakest on its own, and the shadow AI numbers say so plainly. Nearly two-thirds of organizations studied had no AI governance policy at all, and among organizations that suffered an AI-related security incident, 97 percent reported they lacked proper AI access controls.2

Training earns its place when it teaches the flow rather than the policy. Three tiers, four destinations, one aggregation rule, one route for “I do not know.” Anyone can hold that. A forty-slide deck on the information classification standard cannot be held by anybody, and the completion record it generates is evidence of attendance rather than capability. Our companion article on training effectiveness covers how to build assessment that measures whether the decision changed.

Tooling

Endpoint and browser data loss prevention, label-aware blocking, and sensitive information type detection give you a real control on managed devices. Policies match content patterns and take an action at the moment somebody tries to paste or upload.17 This is a genuine improvement over training alone because it does not depend on anybody remembering anything.

Its limits are worth stating honestly. It covers managed devices and managed browsers. It struggles with screenshots and photographs of screens. It generates false positives that erode goodwill if the tuning is left undone. And detection-based controls infer a classification rather than reading a declared one, which makes them a weaker evidentiary story than a label somebody applied on purpose.

A gateway

The strongest control is a single governed path for all model traffic: one sign-in, a short allowlist of permitted destinations, per-tier routing that implements the matrix as a rule rather than a hope, inspection at submission, and a retained record of what was sent and what came back. It is the only one of the three mechanisms that produces a record you can put in front of an inspector and say, here is every submission of company information to a model in the last two years, and here is which tier each one carried.

International cybersecurity authorities publishing joint guidance on securing AI data have converged on the same theme: control the data supply chain, know your provenance, and treat data protection as a lifecycle discipline rather than a perimeter question.19 A gateway is what makes that operationally possible on the submission side.

A gateway you can bypass is worse than no gateway. It creates a control record saying the control existed, alongside a population of people who routed around it. If you build the governed path, block the alternatives in the same change window: block the consumer endpoints at the network and browser layer, and give people a working alternative on day one. Building the path and leaving the alternatives open produces an incomplete record, which is a harder position to defend than no record at all.

The order that works

Deploy the gateway first, because it is the only mechanism that makes the other two meaningful. Then tooling, tuned against the tiers the gateway is already routing on. Then training, which is now teaching a flow the environment actually enforces rather than a policy the environment ignores. Companies that reverse this order spend a year training people on rules nothing checks, and then discover behavior did not change.

What a defensible package looks like. The scheme itself, on one page. The tier-to-destination matrix. A written default that is not the most permissive tier. The aggregation rule. The gateway record, retained for the same period as the underlying records. An exception log with named approvers and stated reasons. A periodic review that examines the exception log and asks whether the tiers are calibrated. That set of artifacts answers essentially every question an inspector or an auditor will ask about AI data handling, and it fits in a single binder.

One further point about who classifies. Classification decisions cannot all route through a central function, because that becomes a queue and the queue becomes the reason people bypass the scheme. But they also cannot be entirely distributed to whoever happens to be building something, which is the failure mode self-service and low-code development introduce into regulated environments. Our article on citizen developers covers the governance model for that population in detail, and the classification scheme is the input it needs.

Why Classification Schemes Fail

After enough of these programs, the failure modes repeat. Almost every one traces back to the same two root causes: too many tiers, and no default. The rest are variations.

Too many tiers

Four tiers is the common case and it is already one too many. Schemes drift upward over time as each new concern gets a level of its own, and practitioners describe the same slow accumulation: more labels, more rules, more exceptions, until the structure is one that hardly anybody understands.3 The tell is easy to check. Ask five people to name the tiers in order. If they cannot, the tiers are not doing any work.

No default, or a permissive one

The second root cause. Either the policy never says what unlabeled means, in which case each person invents an answer, or it says unlabeled is unrestricted, in which case the largest population of documents in the company is cleared for anything. Both are the same failure wearing different clothes.

Names that describe feeling instead of consequence

“Confidential” and “Highly Confidential” are not a distinction anybody can draw reliably, because the difference is one of degree and degrees are arguable. Names that state a consequence, or a permitted destination, remove the argument. This is why we prefer Public, Internal, and Protected. Two of the three names tell you where the information may go, and the third tells you it must not go anywhere without a decision.

The scheme is a policy, not a decision

Many classification standards define the tiers beautifully and never say what a tier permits. They describe encryption requirements and access review frequencies, which are the concerns of the security function, and stop short of the concern of the person holding the document. A tier that does not resolve into a permitted action is a label without a consequence, and people learn to ignore labels without consequences.

Enforcement lives in a different document

The classification standard is owned by information security. The AI use policy is owned by legal or the AI governance committee. The data loss prevention rules are owned by IT operations. Each is internally coherent. Nobody has reconciled the three, and the tier names in one do not map cleanly to the categories in the next. The person in the middle is asked to perform the reconciliation in real time, and they will not.

No route for “I do not know”

If the only path for an uncertain case is to email a data owner who may be on leave and wait three days, the case does not get escalated. It gets guessed, under deadline, and deadline pressure biases guesses toward permissive. Every scheme needs a named route with a same-day service level. In smaller biotechs this can be one person. In large organizations it needs to be a function, but it needs to be fast either way.

Classification treated as a project rather than a property

Programs that begin with a discovery and classification exercise across the existing estate almost always stall, because the estate is larger than expected and the results decay while the work is in progress. Classification is a property of new work and a filter on old work. Label what is created from today. Inspect what is old at the point somebody tries to use it. Resist the project.

The diagnostic is the exception log. Six months after launch, look at how many exceptions were requested. An empty log does not mean the scheme is perfect. It means nobody is using it, because real work always generates edge cases. A log with hundreds of entries in one category means that category is misclassified and the tier assignment should change. The exception log is the only honest feedback the scheme will give you, and it belongs as a standing item at the governance meeting rather than an audit artifact nobody reads.

Conclusion

The classification policy most pharmaceutical and biotech companies are operating today was built to answer a question nobody asks any more. It tells you how to store a document and who may open it. It does not tell a scientist whether the document in front of her can go into the tool on her desktop, and that is now the question that comes up dozens of times a day in every function. The gap does not stay empty. It gets filled with individual judgment, and individual judgment under deadline pressure trends toward whatever gets the work finished.

The change that fixes this is not more governance. It is less, arranged better. Three tiers whose names say what they mean. Four destinations that describe where data actually goes. One matrix connecting them. Six flags for the categories where a statute or somebody else’s contract overrides your judgment. One aggregation rule. One default that is never the most permissive option. One route for uncertainty that answers the same day. That set fits on a page, can be recalled from memory, and can be implemented in a gateway as an actual routing decision rather than a stated intention. It also produces a clean set of records, which is what turns a good idea into something you can show an inspector.

Sakara Digital works with pharmaceutical and biotech organizations building the data governance foundation that makes AI usable rather than merely permitted. If you are rewriting a classification scheme for the AI question, or you have a scheme on paper that nobody follows and you want an independent read on why, we are happy to have that conversation.

For Further Reading