In This Article
- Executive Summary
- What Shadow AI Actually Looks Like in Life Sciences
- Why Blocking Alone Reliably Fails
- Discovery Methods and What Each One Cannot See
- Triage: Sorting Findings by GxP and Data Sensitivity
- The Exposures That Actually Matter in Pharma
- Remediation That Leads With a Sanctioned Alternative
- Running the Program: Cadence, Ownership, and Evidence
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Every pharma and biotech company already has shadow AI. The only open question is whether anyone has looked. Microsoft and LinkedIn found that 78 percent of people who use AI at work bring their own tools, and that more than half hesitate to tell anyone which uses matter most to them.2 IBM’s 2025 breach research put a number on the consequence: 20 percent of breached organizations were compromised through unsanctioned AI, and those breaches ran about $670,000 higher than average.1 In a regulated environment the exposure is not only security. It is data integrity, patient privacy, trade secrets, and GxP records built on output no one reviewed.
This article is a program, not a warning. It describes what shadow AI actually looks like in life sciences (which is far broader than a consumer chatbot in a browser tab), the discovery methods that find it along with the honest limits of each one, a triage model that ranks findings by GxP and data sensitivity rather than treating everything as equally severe, and a remediation approach that leads with a sanctioned alternative. The argument running through all of it: a pure ban converts visible usage into invisible usage, which is strictly worse than what you had.
You will find a five-method discovery plan with coverage gaps stated plainly, a four-tier triage table you can adapt, the specific pharma exposures worth naming to an executive committee, and a remediation sequence with an approval speed target. The single most overlooked finding type gets its own treatment: the vendor who added a model to a system you already validated, and told you in a terms of service update.
What Shadow AI Actually Looks Like in Life Sciences
Most shadow AI programs start with the wrong mental picture. The picture is a scientist pasting text into a free consumer chatbot on a personal laptop. That happens, and it matters, but if it is the only thing you look for you will run a discovery exercise, find a modest number of chatbot sessions, block a few domains, and declare the problem handled while the larger exposures continue untouched.
Shadow AI is better defined as any AI capability processing company data without an assessment, an owner, and a record of the decision to use it. That definition catches five distinct patterns, and only the first is the one everybody pictures.
1. Consumer tools on the open web
This is the visible layer. Someone opens a general purpose assistant and pastes in a protocol section, an investigation summary, a deviation narrative, a batch record excerpt, or a regulatory response draft. Netskope’s threat research found organizations averaging 223 monthly attempts by employees to put sensitive data into generative AI prompts or uploads, with regulated data accounting for roughly 35 percent of data loss prevention violations in that context.34 Personal accounts make this harder to see: research on enterprise browsers found that a large majority of AI usage happens through personal rather than corporate logins.5
2. AI features switched on inside tools you already license
This is the pattern most governance programs are structurally unprepared for. Your document management platform, your electronic lab notebook, your CRM, your ticketing system, your meeting platform, and your office suite have all added AI features over the past two years. Many arrive enabled by default. Many are covered by the master agreement you signed before the feature existed.
From the user’s point of view there is nothing to approve. The summarize button was simply there one morning. From a compliance point of view, a new processing activity started, possibly with a new subprocessor, possibly with data leaving a region, and possibly inside a system that carries a validated status. Nobody filed a change request because nobody experienced it as a change.
3. AI notetakers in meetings
An AI notetaker joins a study team call where unblinded safety data is discussed. It joins a manufacturing investigation call where a root cause is still contested. It joins a partnering discussion under a confidentiality agreement. The transcript, the summary, and the audio now live with a third party, often under terms nobody in quality or legal reviewed.
Legal analysis of these tools has moved quickly. Mayer Brown’s 2026 review notes that AI notetakers turn conversations that used to be temporary into permanent, searchable records reachable well beyond the original participants, with consent, privilege, and cross-border transfer consequences that vary sharply by jurisdiction.18 In life sciences the additional problem is scientific. An imperfect summary of a technical discussion, circulated as the record of what was decided, can propagate an error into a document that later gets filed.
4. Browser extensions
Extensions are the quietest category. A person installs something that summarizes pages, rewrites email, or explains a PDF. Cloud Security Alliance research describes AI browser extensions as an attack surface that conventional data loss prevention tooling does not see, because the content is read and transmitted inside the browser process rather than crossing the network in a form the inspection point recognizes.6 Browser security research has found that nearly all enterprise users have at least one extension installed, and that around half have granted at least one extension permissions broad enough to read page content, cookies, or credentials.5
5. A vendor adding a model to a validated system
Here is the case that most programs miss entirely, and the one worth raising with your executive committee first.
You validated a system. You have a validation summary report, a traceability matrix, a supplier assessment, and a periodic review schedule. Eighteen months later the supplier releases a version that adds a model to a function inside that system: anomaly flagging, deviation classification, document summarization, search ranking, or coding suggestions. The release note calls it an enhancement. The terms of service update mentions AI processing and, in some cases, model improvement using customer content. Your validated state now includes a component that was never in scope of your validation, never risk-assessed, and never covered by your supplier assessment.
This is the exact seam the European regulators are moving to close. The July 2025 EudraLex Volume 4 consultation package revised Chapter 4 and Annex 11 and introduced an entirely new Annex 22 dedicated to artificial intelligence.12 The draft Annex 11 expands from roughly five pages to nineteen and adds explicit supplier and service management requirements, including mandatory contract elements covering change notification.12 Draft Annex 22 sets expectations for intended use, data quality, performance metrics, change control, ongoing model performance monitoring, and human review for AI in critical applications affecting patient safety, product quality, or data integrity.1314 A model that arrives inside a validated system through a supplier release is precisely the situation those requirements describe, and precisely the situation most change control processes will not catch.
The finding to look for first. Ask your application owners a single question: which of our validated or GxP-relevant systems have released an AI feature in the last twenty-four months, and did that release go through change control as an AI change? In most organizations the honest answer is that nobody has checked. That answer is more useful than any network report, because it is the finding an inspector is most likely to reach on their own.
Why Blocking Alone Reliably Fails
The instinct after a first discovery exercise is to block. Add the domains to the deny list, issue a policy stating that AI tools require approval, run a training module, and move on. It is fast, it is defensible in a meeting, and it produces a clean chart showing usage dropping to near zero.
The chart is the problem. Usage did not drop. Visibility did.
The demand is real and it does not evaporate
People are not using these tools to be difficult. In the healthcare setting, where a December 2025 survey of more than 500 workers found 17 percent admitting to using unauthorized AI tools and more than 40 percent aware of colleagues doing so, the most common reason given was not curiosity.78 Close to half said there was no approved tool available that did what they needed.9 That is not a discipline problem. That is a supply problem, and a deny list does not change supply.
Blocking is technically porous
A network deny list on a corporate device does not stop a personal phone on cellular data, a personal laptop at home, a mobile app, an AI feature embedded in a licensed application, a browser extension that never touches a blocked domain, or an assistant reached through a different vendor’s white-labeled front end. The controls you can enforce cover a shrinking share of the ways a person can reach a model.
Blocking without an alternative destroys your reporting channel
This is the real damage. When usage is technically permitted but ungoverned, people will tell you about it. They will ask questions, request features, and raise their hand when something looks wrong. Once usage is prohibited, every one of those conversations stops, because a person cannot report a problem with a tool they were not supposed to be using.
Microsoft’s research found that more than half of workplace AI users already hesitate to disclose their most important AI uses, and roughly the same share worry that relying on AI makes them look replaceable.2 A prohibition adds a disciplinary consequence on top of an existing reluctance. In a GxP environment that is a serious problem, because your quality system depends on people escalating things. A culture where nobody will admit how a document was drafted is a culture where you cannot reconstruct how a document was drafted, which is a data integrity failure whatever the policy says.
What blocking is actually for
None of this means never block. Blocking is a legitimate control for a specific, narrow purpose: stopping a named tool with a known unacceptable data handling practice, at a moment when you have a sanctioned alternative to point at. Blocking as a targeted response to a triaged finding works. Blocking as the whole program does not.
The rule worth stating out loud to leadership: a prohibition without a sanctioned alternative does not reduce usage. It relocates usage to devices, accounts, and networks you cannot see, and it closes the channel through which you would otherwise have learned about it. You end up with the same exposure and less information about it.
Discovery Methods and What Each One Cannot See
Discovery only produces a credible picture when you run several methods together and are honest about what each one misses. Any single method oversells its coverage. Run all five and you get a defensible inventory, plus a documented statement of residual blind spots, which is itself a useful artifact when an inspector asks how you know.
Method 1: Network and secure web gateway egress analysis
Pull outbound traffic from your gateway, proxy, or cloud access security broker and classify destinations against a maintained list of AI service endpoints. This gives you volume, frequency, which business units are involved, and, where inline inspection is in place, some visibility into what is being sent.
What it cannot see: anything off the corporate network. Personal phones on cellular data, home devices, and unmanaged laptops are invisible. It also misses AI features inside applications you already allow, because the traffic goes to a domain you have classified as your document management system or your CRM, not as an AI service. And it misses model calls made server to server by a licensed application on your behalf.
Method 2: SSO and OAuth grant review
This is the highest yield method available to most organizations, and the most consistently skipped.
Export the full list of applications registered in your identity provider, along with every third-party OAuth consent grant your users have approved. Every time someone clicked “continue with Microsoft” or “continue with Google” to start using an AI tool, your identity provider recorded it. More importantly, the grant record shows the scopes: whether that tool can read mail, read and write files, access calendars, or read the contents of a shared drive.
Two properties make this method valuable. First, it is retrospective. It shows you tools adopted months or years ago, not just current traffic. Second, it captures the durable exposure rather than the momentary one. A grant persists until it is revoked, so a tool someone tried once in 2024 and abandoned may still hold a live token to a document library today. Sort the export by scope sensitivity rather than by user count, and the top of that list is usually the most alarming page in the entire discovery exercise.
What it cannot see: tools people signed up for with an email and password instead of federated login, tools reached through a personal account, on-device software, and browser extensions that do not use OAuth. It also will not tell you what data actually moved, only what the tool was permitted to reach.
Method 3: Expense and procurement records
Search corporate card transactions, expense reimbursement records, and accounts payable for AI vendor names, and separately for small recurring charges in the twenty to two hundred dollar per month range that no one has mapped to a contract. Individually expensed subscriptions are how AI enters an organization at the team level, one manager at a time, below every procurement threshold that would have triggered a security or privacy review.
Pair this with a review of any renewal or amendment in the last two years where a vendor introduced AI functionality. That review is what surfaces the validated-system case described earlier.
What it cannot see: free tiers, which is where a large share of the riskiest use lives. It also misses anything paid for personally and never expensed, which is common precisely when someone suspects the purchase would not be approved.
Method 4: Endpoint and browser extension inventory
Query your endpoint management platform for installed applications and, critically, for browser extensions and their granted permissions. Most organizations have never produced this list. When they do, the finding is rarely the number of extensions. It is the permission scope: extensions able to read and change data on all sites, running on machines where people open clinical documents and manufacturing records.56
What it cannot see: unmanaged and personal devices, and web-only tools that install nothing. It also cannot tell you whether an extension was actually used on a sensitive page, only that it had the ability to be.
Method 5: Anonymous survey
Ask people. Anonymously, with a credible guarantee, and with questions about what they are trying to accomplish rather than only what they used.
The survey is the only method that reaches personal devices, free tiers, and intent. It is also the only one that tells you which unmet needs are driving the behavior, which is the input you need to build the sanctioned alternative. Ask what task they were doing, what tool they reached for, what they put into it, and what would have to be true for an approved option to replace it.
What it cannot see: anything people are unwilling to disclose, which is a large category given that more than half of AI users already hold back about their most significant uses.2 Treat survey results as a floor, never a measurement. If the survey and the technical methods disagree, the technical methods are closer to the truth on volume and the survey is closer to the truth on motive.
| Method | Best at finding | Primary blind spot | Typical owner |
|---|---|---|---|
| Network and gateway egress | Volume and frequency of consumer AI use on managed networks | Off-network devices; embedded AI in allowed applications | Security operations |
| SSO and OAuth grant review | Durable data access held by third-party AI tools, including dormant ones | Password-based signups; personal accounts; non-OAuth tools | Identity and access management |
| Expense and procurement | Team-level paid subscriptions below procurement thresholds | Free tiers; personally funded tools never expensed | Finance and procurement |
| Endpoint and extension inventory | Browser extensions with broad page-content permissions | Unmanaged devices; purely web-based tools | End user computing |
| Anonymous survey | Personal-device use, free tiers, and the unmet need behind the behavior | Anything people choose not to disclose | Quality or a neutral function, not the manager |
| Supplier release and terms review | Models added to systems you already validated | Suppliers who do not disclose model components in release notes | Validation and vendor management |
Write down what you did not find
The output of discovery is two documents, not one. The first is the inventory. The second is a short, signed statement of coverage: which methods were run, over what period, across which populations, and which categories of use remain outside the reach of all of them. That second document is what turns a discovery exercise into evidence. It also prevents the most common failure mode, which is a leadership team treating an inventory built from network logs alone as a complete picture.
Triage: Sorting Findings by GxP and Data Sensitivity
A discovery exercise across a mid-size biotech typically produces somewhere between forty and two hundred distinct findings. If you treat them as equally serious you will do two things, both bad. You will spend your limited remediation effort on the wrong items, and you will lose credibility with the business the first time you send an urgent notice to a marketing team about a headline generator.
Sort on two axes only. Everything else is detail.
Axis one: does the output touch a GxP decision or record? Not whether the tool is used by a GxP function, but whether what comes out of it can end up in a regulated record, influence a quality decision, or support a regulatory submission.
Axis two: what data class went in? Public or general business information, confidential business information, trade secret or unpublished clinical data, or personal and patient data.
| Tier | Pattern | Example finding | Response and timing |
|---|---|---|---|
| Tier 1: Critical | Regulated personal or patient data, or unblinded clinical data, in an unassessed tool | Subject-level safety narratives pasted into a consumer assistant with a personal account | Stop immediately. Privacy incident assessment within 24 hours. Named executive owner. Revoke access. |
| Tier 2: High | AI output reaching a GxP record or decision without the review the quality system requires | A model inside a licensed system classifying deviations; AI-drafted investigation text approved without documented human review | Quality assessment within 5 business days. Decide: retrospective assessment, change control, or removal. Consider a deviation. |
| Tier 3: Moderate | Trade secret or unpublished research data in an unassessed tool, no GxP record involvement | Formulation notes or unpublished screening results summarized by a browser extension | Assess within 30 days. Contract and data-use review. Sanctioned alternative offered before restriction. |
| Tier 4: Low | Public or general business information, no GxP or personal data | Marketing drafting external copy; a manager summarizing a public conference agenda | Log it. Route to the sanctioned path at the next convenient point. No escalation, no notice. |
The discipline that makes this model work is refusing to escalate Tier 4. It is tempting, because escalation feels like rigor. It is not. Every Tier 4 item you treat as urgent teaches the business that your severity ratings carry no information, and the next time you send a genuine Tier 1 notice it will sit in an inbox alongside the memo about the conference agenda.
Two judgment calls worth deciding in advance
Where does anonymized or de-identified data sit? Lower than identified data, but not automatically at Tier 4. The EDPB’s Opinion 28/2024 sets a demanding standard for when data associated with an AI model can be treated as genuinely anonymous, and treats the question as case by case rather than settled by a label.17 If your de-identification has not been tested against re-identification risk in combination with the other information the tool holds, do not assume it drops a tier.
What about drafting assistance where a human rewrites everything? The input still matters. If someone pastes a confidential document in to get help restructuring it, the exposure already occurred regardless of how much of the output survives. Classify on what went in, not on what came out.
The Exposures That Actually Matter in Pharma
Generic shadow AI material talks about data leakage in the abstract. In pharma and biotech the exposures are specific, and naming them plainly is what gets a governance program funded.
Personal and patient data
Adverse event narratives, investigator correspondence, patient support program records, and medical information inquiries all carry personal data. Sending them to a tool with no data processing agreement creates a processor relationship nobody authorized, potentially an international transfer with no mechanism, and under US health privacy rules a disclosure to an entity with no business associate agreement in place.
Unpublished clinical data
Interim results, unblinded safety information, and statistical outputs before database lock. Disclosure risks the integrity of the trial itself, not only confidentiality. An AI notetaker sitting silently on a data monitoring committee call is a version of this exposure that no data loss prevention rule will catch.
Trade secrets and manufacturing know-how
Process parameters, cell line details, formulation work, analytical methods, and yield data. Trade secret protection depends on demonstrating reasonable measures to keep the information secret. Uncontrolled disclosure to a third-party service undermines that argument in a way that is difficult to repair after the fact.
AI output inside a GxP record
A deviation narrative, a CAPA effectiveness summary, a validation rationale, or a submission section drafted by a model and approved without the review the quality system requires. The record is now attributable to a person who did not author its reasoning, and the organization cannot reconstruct how the conclusion was reached.
Why the fourth exposure is the one that changes an inspection
The first three are serious and largely familiar. Privacy, confidentiality, and trade secret exposure are risks your legal and security functions already understand, even if the AI route is new.
The fourth is different because it goes to the credibility of your records. Regulators have been consistent that AI used to produce information supporting regulatory decisions needs a defined intended use, a risk assessment proportionate to influence and consequence, and documented evidence that the model is fit for that use. The FDA’s January 2025 draft guidance builds exactly that framework, requiring sponsors to state the question the model addresses, define its context of use, assess risk from model influence and decision consequence, and execute a credibility assessment plan against that risk.1011 Draft Annex 22 does the parallel work on the manufacturing side, with expectations for intended use, data quality, performance monitoring, change control, and human oversight for models in critical applications.1314
Shadow AI, by definition, has none of that. There is no intended use statement, no risk assessment, no performance evidence, and no record of human review. If AI-generated content is sitting in a GxP record and you cannot show which parts came from a model or what review it received, the deficiency is not that you used AI. It is that you cannot answer a basic question about the provenance of a regulated record.
The regulatory floor is rising while this sits unaddressed
Two things are worth putting in front of a leadership team. The EU AI Act’s transparency obligations under Article 50 became applicable and enforceable on 2 August 2026, covering systems that interact with people or generate content, whether or not they are classified as high risk.15 Separately, the AI literacy obligation under Article 4 requires providers and deployers to take measures ensuring a sufficient level of AI literacy among staff operating these systems on their behalf.16 An organization that does not know which AI systems its staff are using cannot demonstrate either.
A note on retrospective exposure. The EDPB has addressed what happens when personal data was processed unlawfully in the development of a model, and whether that affects later operation of the model.17 The practical consequence for a shadow AI program is that a finding is not closed by turning the tool off. If regulated data went into a service under terms permitting model improvement, the remediation question includes what happens to what was already sent, and whether the vendor’s terms give you any mechanism to ask.
Remediation That Leads With a Sanctioned Alternative
Remediation is where most programs quietly fail. The findings are real, the triage is sound, and then the response is a policy reminder and a deny list, which returns you to the situation described earlier.
The organizing principle is simple to state and difficult to fund: for every category of use you intend to restrict, there must be an approved way to do the same work, available before the restriction lands. When close to half of unauthorized use is driven by the absence of an approved option, the absence is the finding.9
What a credible sanctioned path looks like
Credible means it survives contact with the person who was using the shadow tool. It has to be good enough that switching is not a downgrade, because if it is a downgrade people will use both: the approved tool for anything anyone might audit, and the shadow tool for the work.
Comparable, not degraded
A model generation or two behind, with tight output limits and no file upload, will not displace a current consumer tool. If the approved option cannot do the task, the restriction will simply move the task somewhere you cannot see.
Contracted properly
Enterprise agreement with no training on your content, defined retention, named subprocessors, a data processing agreement, a transfer mechanism, and audit rights. This is what makes the approved path defensible, and it is the part that takes the longest.
Available without a project
If getting access requires a business case and a steering committee slot, the approved path does not exist in any practical sense. Access should be a request handled in days, gated by training and an acknowledged acceptable use standard.
Clear rules people can apply
Say what data classes may go in, what output may go into a GxP record and under what review, and what must be recorded. Rules stated as concrete examples get followed. Rules stated as principles get interpreted generously.
How fast is fast enough
Speed is the variable that decides whether the sanctioned path wins. A person who wants to summarize a two-hundred-page document today will not wait a quarter for governance to catch up. They will use the tool they already have, once, and then keep using it.
Two targets are worth committing to publicly, because publishing them changes behavior more than the policy does:
- Access to an already-approved tool: five business days or less. Training completion and an acknowledged acceptable use standard, then access. No project, no business case.
- Assessment of a newly requested tool: an initial answer in fifteen business days. The answer may be no, or approved for Tier 4 use only, or approved pending contract terms. What matters is that a definite answer arrives inside the window where the person still has the option to wait.
If your current process cannot meet those numbers, that is the first remediation item, ahead of any individual finding. A governance function that takes four months to answer a request is not a control. It is the reason the shadow tool was adopted.
The remediation sequence
Contain Tier 1 immediately, before anything else is built
Regulated personal data and unblinded clinical data in unassessed tools does not wait for a program. Revoke the OAuth grants, run the privacy incident assessment, notify the accountable owner, and document the decision. This is the only category where restriction precedes an alternative.
Group the remaining findings into use cases, not tools
Twelve findings involving five different tools may all be one use case: summarizing long technical documents. Solve the use case once. Tool-by-tool remediation produces an endless list, because a new tool appears every month; use-case remediation converges.
Stand up the sanctioned path for the top three use cases
Rank by user count multiplied by data sensitivity. Three is enough to cover most of the volume in a typical organization, and few enough to deliver in a quarter. Contract the terms properly, publish the data class rules, and open access.
Bring embedded and supplier-added AI into change control
Add an explicit AI question to your change control and periodic review templates for GxP systems, and an AI change notification clause to supplier agreements at renewal. Perform a retrospective assessment on any validated system where a model arrived without one.
Restrict, narrowly, and only now
With alternatives live, block the specific tools with unacceptable data handling. Announce the restriction and the alternative in the same message, from the same sender, on the same day. A restriction announced alone reads as a prohibition; announced with a replacement it reads as a migration.
Reopen the reporting channel deliberately
State plainly that disclosing past use of an unapproved tool carries no disciplinary consequence, and mean it. Then give people a standing, low-friction way to request a tool. The volume of requests you receive afterward is the best single measure of whether the program is working.
What good looks like at twelve months. Discovery runs on a schedule rather than as a project. New AI capability requests arrive through a front door and get answered inside the published window. Every GxP system change record has an explicit AI question answered. Supplier agreements carry AI change notification clauses. And the number of people willing to tell you what they are using has gone up, not down. That last one is counterintuitive and it is the most reliable signal you have.
Running the Program: Cadence, Ownership, and Evidence
A one-time discovery exercise has a short shelf life. The AI capability landscape changes faster than an annual review cycle, and most of the change arrives through vendors rather than through employees, which means the inventory decays even if nobody in your organization does anything new.
Cadence
| Activity | Frequency | Output |
|---|---|---|
| OAuth grant and SSO application review | Quarterly | New grants since last review, sorted by scope sensitivity; revocation list |
| Network and endpoint discovery sweep | Quarterly | New destinations and extensions; trend against previous quarter |
| Supplier release and terms review for GxP systems | Semiannual, plus on every major release | AI components introduced; change control records raised |
| Expense and procurement scan | Semiannual | Unmapped recurring charges; individually expensed subscriptions |
| Anonymous use and unmet need survey | Annual | Motive data; input to the sanctioned tool roadmap |
| Inventory and coverage statement refresh | Annual | Signed statement of what was searched and what remains outside coverage |
Ownership
The most common structural mistake is giving this to security alone. Security can run the technical discovery, but it cannot make the GxP call on whether an AI-assisted deviation narrative needs a retrospective assessment, and it has no standing to commit the organization to a sanctioned tool roadmap.
The workable arrangement puts three roles on the same standing agenda. Security and IT own discovery and technical remediation. Quality owns the triage decision for anything touching a GxP record and owns the change control integration. A business-side owner, usually whoever runs the digital or data function, owns the sanctioned tool roadmap and the approval speed targets. Privacy and legal join for Tier 1 and Tier 3 findings. If the roadmap has no owner, the program becomes a restriction exercise by default, because restriction is the only action a security function can take alone.
Evidence
Frameworks already in use are good scaffolding here rather than something to invent. The NIST AI Risk Management Framework provides a structure for governing, mapping, measuring, and managing AI risk that maps cleanly onto discovery, triage, and remediation.19 GAMP 5 Second Edition supplies the risk-based validation approach and the critical thinking principles that determine how much evidence a given use actually needs.20 Use what you have. A shadow AI program that produces its own bespoke vocabulary will not survive an inspection question, because nobody will be able to map it to the quality system.
The evidence set worth maintaining is short: the current inventory, the coverage statement, the triage decisions with rationale, the change control records for AI introduced into GxP systems, the list of sanctioned tools with their contracted terms, and the training and acceptable use acknowledgments. That set answers most of what an inspector or an auditor will ask, and it is small enough to actually keep current.
The question to ask at the next quality council
Not “do we have a shadow AI policy” and not “have we blocked the consumer tools”. Ask instead: for each of our GxP-relevant systems, can we state whether an AI component is present, when it was introduced, and what assessment covers it? If the answer is no for even a handful of systems, you have found the work. That question is also, increasingly, the one a regulator is positioned to ask, given how explicitly the draft Annex 11 and Annex 22 revisions address supplier change notification and AI in critical applications.1213
Conclusion
Shadow AI in life sciences is not a discipline problem and it is not solved by a policy. It is a supply problem sitting on top of a visibility problem, in an environment where the consequences reach further than they do in most industries, because the output can end up in a record a regulator will read. The organizations handling it well are not the ones with the strictest prohibition. They are the ones that looked properly using several methods, told the truth about what those methods could not see, sorted the findings by GxP and data sensitivity instead of treating them all as emergencies, and put a genuinely usable approved option in front of people before restricting anything.
Two things are worth carrying out of this. First, the finding most programs miss is not the employee with the consumer chatbot. It is the model a supplier added to a system you validated, disclosed in a release note and a terms of service update, and never routed through change control. Start there, because it is the finding an inspector can reach without your help. Second, resist the reflex to measure success by usage of unapproved tools falling to zero. In the early months of a healthy program, disclosed usage usually goes up, because people finally have a reason to tell you. That is the program working, not failing.
Sakara Digital works with pharma and biotech organizations building AI governance that holds up in a regulated setting: discovery that is honest about its limits, triage that a quality unit will stand behind, and a sanctioned path fast enough to actually displace the shadow tool. If you are starting a shadow AI discovery exercise, or you have run one and are not sure what to do with the findings, we are happy to have that conversation.
For Further Reading
For Further Reading
- AI Governance Framework for Pharma QA Teams
- Generative AI Governance in Regulated Life Sciences
- Establishing an AI Policy: A Comprehensive White Paper for Life Sciences Teams
- How to Build an AI Change Control Process in Regulated Systems
- The AI Model Risk Assessment for Pharma: A Structured Checklist
- AI Vendor Selection Guide for Regulated Life Sciences Environments
- Patient Data Privacy in Clinical Research: Navigating GDPR and FDA Expectations
References & Sources
- IBM Security. “Cost of a Data Breach Report 2025.” IBM, 2025. https://www.ibm.com/reports/data-breach
- Microsoft and LinkedIn. “AI at Work Is Here. Now Comes the Hard Part.” 2024 Work Trend Index Annual Report, May 2024. https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part
- Netskope Threat Labs. “Cloud and Threat Report: 2026.” Netskope, 2026. https://www.netskope.com/resources/cloud-and-threat-reports/cloud-and-threat-report-2026
- Help Net Security. “Gen AI data violations more than double.” 7 January 2026. https://www.helpnetsecurity.com/2026/01/07/gen-ai-data-violations-2026/
- LayerX Security. “AI Is Now the #1 Data Exfiltration Vector in the Enterprise, and Nobody’s Watching.” LayerX, 2025. https://layerxsecurity.com/blog/ai-is-now-the-1-data-exfiltration-vector-in-the-enterprise-and-nobodys-watching/
- Cloud Security Alliance Labs. “AI Browser Extensions: Shadow AI’s Hidden Attack Surface.” CSA Research Note, 2026. https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-browser-extension-attack-surface-202604/
- Wolters Kluwer. “Shadow AI: Providers are using unapproved tools to improve workflow.” Expert Insights, 2026. https://www.wolterskluwer.com/en/expert-insights/shadow-ai-providers-are-using-unapproved-tools-to-improve-workflow
- Healthcare Dive. “‘Shadow AI’ use is widespread in healthcare: survey.” 2026. https://www.healthcaredive.com/news/shadow-unauthorized-ai-/810191/
- CIO Dive. “Shadow AI use is widespread in healthcare: survey.” 2026. https://www.ciodive.com/news/shadow-unauthorized-ai-healthcare/810421/
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products.” Draft Guidance for Industry, January 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological
- Federal Register. “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products; Draft Guidance for Industry; Availability.” 7 January 2025. https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug
- European Commission, DG Health and Food Safety. “Stakeholders’ Consultation on EudraLex Volume 4 Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and New Annex 22.” 7 July 2025. https://health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en
- ECA Academy. “EU GMP Annex 22 (Draft 2025): Artificial Intelligence.” GMP Guideline summary, 2025. https://www.gmp-compliance.org/guidelines/gmp-guideline/eu-gmp-annex-22-draft-2025-artificial-intelligence
- Stassen, M., Schmucki, M., Valero, F., and Manzano, T. “Bridging Guidance and Regulation: Interpreting the Draft Annex 22 on Artificial Intelligence in GMP Manufacturing.” PDA Journal of Pharmaceutical Science and Technology, 2026. https://pubmed.ncbi.nlm.nih.gov/41698693/
- EU Artificial Intelligence Act. “Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems.” https://artificialintelligenceact.eu/article/50/
- EU Artificial Intelligence Act. “Article 4: AI Literacy.” https://artificialintelligenceact.eu/article/4/
- European Data Protection Board. “Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models.” 17 December 2024. https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- Mayer Brown. “AI Notetakers: Productivity Tool or Emerging Legal Risk?” 3 June 2026. https://www.mayerbrown.com/en/insights/publications/2026/06/ai-notetakers-productivity-tool-or-emerging-legal-risk
- National Institute of Standards and Technology. “AI Risk Management Framework.” NIST. https://www.nist.gov/itl/ai-risk-management-framework
- ISPE. “GAMP 5 Guide: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition).” https://ispe.org/publications/guidance-documents/gamp-5-guide-2nd-edition
- Pharmaceutical Executive. “What is Shadow AI and How is it Impacting Pharma Companies?” PharmExec. https://www.pharmexec.com/view/shadow-ai-impacting-companies
- Stratokey. “Your SaaS is Adding AI Faster Than Compliance Can Keep Up.” https://www.stratokey.com/blog/saas-ai-features-outpacing-compliance-checks








Your perspective matters—join the conversation.