When a Prompt Stops Being Personal Scratch Work

The first six months of generative AI inside a life sciences organization look like individual experimentation. One person in regulatory affairs finds a way to get usable first drafts of module summaries. Someone in manufacturing works out how to turn shift notes into a structured deviation narrative. A clinical data manager builds a query that pulls the sense out of a monitoring report. Each of these is private work. Nobody governs it, and at that scale nobody should.

Then something changes, and it changes fast. The person in regulatory affairs sends the prompt to two colleagues. One of them adds it to a team wiki. It gets demonstrated in a lunch session. Within a quarter, a dozen people are using a lightly modified version of the same instruction to produce content that flows into controlled documents. Nobody made a decision to standardize on it. It spread because it worked.

This is the moment the governance question arrives, and most organizations miss it entirely. The prompt is no longer one person’s scratch work. It is a shared instruction that shapes how a group of people produce regulated output. It has an author who may have left the team, no stated boundaries, no record of what model it was written against, and no evidence that anyone checked whether it still behaves the way it did when it was written.

Why this specific gap is so easy to miss

Quality organizations in pharma are good at controlling two kinds of things: documents and software. A prompt is neither, and it falls between the two control systems in a way that is almost designed to be invisible.

It is not a document, because nobody reads it as content. It has no approval workflow, no effective date, no periodic review, and it does not appear on any document index. It is not software, because it is not compiled, not deployed, not part of a release, and not covered by any change control record that a validation lead would recognize. It lives in a chat history, a shared note, a spreadsheet tab, or the saved instructions field of a commercial AI assistant.

And yet it does exactly what a controlled artifact does. It determines, in a repeatable way, how a task gets performed and what the output looks like. Change one sentence in a prompt and the structure of every downstream document changes with it. Nobody signs anything. Nobody is notified.

The practical failure mode. An organization discovers during an audit readiness review that four different versions of a “CAPA effectiveness summary” prompt are circulating. Two of them instruct the model to propose a conclusion when the evidence is ambiguous. Nobody can say which version produced which historical record, because nothing was captured alongside the output. The finding is not that AI was used. The finding is that the organization cannot describe the instruction that produced the record.

The regulators have started asking about the instruction, not just the tool

European regulators moved on this earlier than most companies expected. In September 2024, the European Medicines Agency and the Heads of Medicines Agencies published guiding principles on the use of large language models in regulatory science and medicines regulatory activities.12 The document is short and deliberately practical, and one of its recurring themes is that the user has to understand the application well enough to adapt their prompting to the degree of control the organization actually has over that application. That framing is worth sitting with. It treats the prompt as part of the control environment, not as an incidental keystroke.

That is not a one-off document either. The EMA and HMA workplan for 2025-2028 names guidance on AI use across the medicine lifecycle as a strategic area, and the supporting governance includes annual observatory reporting on AI activity across the network and joint work with the FDA on shared principles for AI across the medicines lifecycle.19 Every one of those mechanisms depends on an organization being able to say what was asked of the model. In a generative system, what was asked is the prompt.

The Argument: A Reused Prompt Is Functionally a Procedure

Here is the central claim, stated plainly. A prompt that is used repeatedly to produce content that lands in a GxP record is functionally a procedure. Not metaphorically. Functionally. It has every property that makes a procedure a procedure.

A procedure specifies how a task is to be performed. It is written down so that different people performing the task get comparable results. It is stable over time unless deliberately changed. It has a scope, an owner, and consequences when it is wrong. A reused prompt has all of these. The only difference is that the entity following the instruction is a language model rather than a person, and that difference cuts in the direction of more control rather than less, because the model will follow an ambiguous instruction confidently and without flagging the ambiguity.

The regulatory foundation for this is not exotic. In the United States, 21 CFR 211.186 requires that master production and control records be prepared, dated, and signed by one person, then independently checked, dated, and signed by a second.14 Section 211.188 then requires each batch production record to be an accurate reproduction of the appropriate master record, which is another way of saying that the instruction and the record it produced have to be traceable to each other.15 The principle underneath both is the same: instructions that shape regulated output are themselves controlled, someone other than the author confirms they are fit for the purpose, and the finished record can be tied back to the instruction that governed it.

ICH Q10 describes a pharmaceutical quality system that manages knowledge and process performance across the product lifecycle, with change management applied to changes that could affect product quality.13 A prompt that generates the narrative portion of a batch investigation is squarely inside that scope. It does not become exempt because it is written in English.

The reframe that makes this workable. The question is not “does this prompt need to be validated?” That question produces a binary answer and the binary answer is almost always “no, that would be absurd.” The better question is “what does the output of this prompt touch, and what level of control is proportionate to that?” That question has useful answers across the whole range, from none at all to full change control with independent review.

Why natural language makes this harder, not easier

There is a comfortable assumption that because a prompt is written in plain English, its behavior is self-evident and anyone reading it can tell what it does. The research says otherwise, and this is the point where a lot of governance conversations turn from theoretical to concrete.

Work on prompt sensitivity has repeatedly shown that changes to a prompt that carry no semantic content at all can move model behavior substantially. One study measuring sensitivity to formatting choices in prompt design found performance spreads of up to 76 accuracy points across prompt formats that a human reader would consider equivalent.2 A separate evaluation across plain text, Markdown, JSON, and YAML templates found performance on a code translation task varying by as much as 40 percent for one model purely on the basis of the template used.3

The practical implication for a regulated organization is uncomfortable. Two prompts that look interchangeable in a review meeting can produce meaningfully different output quality. A reviewer approving a prompt by reading it is not doing the same thing as a reviewer approving an SOP by reading it. The prompt needs evidence attached, not just a signature.

76 pts Maximum accuracy spread observed across semantically equivalent prompt formats in one benchmark study [2]
84% to 51% Change in one model’s accuracy on the same prime number identification task between two versions three months apart [1]
60 days Minimum notice one major model provider commits to before retiring a publicly released model [16]

A Tiering Model That Keeps Governance Proportionate

Every AI governance program that dies does so for the same reason. It applies the heaviest available control to the lightest available risk, people conclude the program is unserious about their actual work, and they route around it. If a scientist has to raise a change request to adjust the phrasing of a prompt that summarizes conference abstracts, the scientist will simply stop telling anyone about their prompts.

So the tiering has to be real, and the bottom tier has to be genuinely light. The variable that determines the tier is not how clever the prompt is or how many people use it. It is what the output touches.

Tier What the output touches Examples Control applied
Tier 0
Personal
Nothing that leaves the individual. Thinking aid only. Summarizing a paper for your own reading, rephrasing an internal email, brainstorming meeting agendas. None beyond the organization’s acceptable use policy. Do not register these. Do not ask people to.
Tier 1
Shared, non-regulated
Work products used by a team but outside any GxP record or external communication. Internal project status drafts, first-pass literature triage, meeting note structuring. Named owner, a one-line intended use, and a place in the library so people find it instead of rewriting it. No approval workflow.
Tier 2
Regulated input
Content that a qualified human reviews and then places into a GxP record or a regulatory submission. Deviation narrative drafts, protocol section drafts, CAPA effectiveness summaries, safety narrative first drafts. Full asset record (see next section), documented evaluation against a defined set, second-person review before release, change history, periodic review.
Tier 3
Regulated output with limited human filtering
Content or decisions that flow into a regulated process at volume, where realistic human review cannot catch every error. High-volume case triage and coding support, automated document classification driving retention decisions, agent workflows that write to validated systems. Everything in Tier 2 plus formal qualification against the specific model version, monitoring in production, defined failure and escalation handling, and inclusion in the validated system’s change control.

Three things make this tiering work in practice rather than on paper.

The tier is assigned by the use, not by the prompt. The same instruction can sit in Tier 1 for one team and Tier 2 for another because of where the output goes. That means the registration record captures the use, and a new use of an existing prompt is a new registration, not a silent reuse. This sounds like bureaucracy until the first time someone takes a “summarize this document” prompt qualified for internal reading and points it at a regulatory commitment tracker.

Tier 0 must be genuinely free. If your policy says every prompt must be registered, you have written a policy that guarantees non-compliance and teaches people that the AI rules are theater. Say explicitly that personal-use prompting is not registered and not reviewed, and spend the credibility you gain on the tiers that matter.

Promotion is the event that triggers control. The governance moment is not authorship. It is the moment someone shares a prompt with the intent that others use it for regulated work. That is a discrete, observable event, and it is the right place to put the gate.

A useful test for tier assignment. Ask the person requesting registration one question: if this prompt silently started producing subtly worse output tomorrow, who would notice, and how long would it take? If the answer is “me, immediately,” it is Tier 0 or 1. If the answer is “the reviewer, probably,” it is Tier 2. If the answer is “possibly nobody for months,” it is Tier 3 regardless of what anyone would prefer.

What a Governed Prompt Asset Actually Carries

A prompt library that stores prompt text and nothing else is a text file with ambitions. The value is in the metadata, and the metadata is what makes the difference between an asset you can defend and a string you found in a wiki.

Here is what a Tier 2 or Tier 3 prompt asset should carry. None of these fields are exotic. All of them are things a validation lead would expect for any other controlled artifact, and the absence of any one of them creates a specific, predictable problem later.

FIELD 1

Named owner and function

A person, not a team mailbox. The owner answers questions about intent, approves changes, and is the one who gets the review reminder. Ownership transfers explicitly when people move roles. An asset with no owner is a candidate for retirement, not a candidate for continued use.

FIELD 2

Stated intended use

One or two sentences describing the task the prompt is qualified for, written concretely enough that someone can tell whether their situation matches. “Drafts the investigation summary section of a Category 2 manufacturing deviation from an approved evidence package” is useful. “Helps with deviations” is not.

FIELD 3

Explicit out-of-scope uses

The field most often skipped and most often needed. Name the adjacent uses this prompt must not be applied to, especially the plausible ones. This is what stops a summarization prompt qualified for internal reading from being pointed at a submission document because the two tasks sound similar.

FIELD 4

Model and version qualified against

The specific model identifier and version snapshot, plus any inference settings that were fixed during evaluation. Without this the evaluation evidence is unattached to anything and cannot be reproduced. This is the field that makes the next section possible.

FIELD 5

Evaluation evidence

What was tested, against what inputs, judged by whom or by what, and what the result was. For most Tier 2 assets this is a modest artifact: twenty to fifty representative inputs, expected characteristics of a good output, and a reviewer’s assessment. It does not need to be a validation package. It does need to exist.

FIELD 6

Review date and change history

A next-review date that someone is accountable for, and a versioned history showing what changed, who changed it, why, and what evaluation was rerun. Change history is what lets you answer the question “what instruction produced this record in March?” which is the question you will eventually be asked.

The out-of-scope field deserves more attention than it gets

Most prompt libraries that do carry metadata carry an intended use and stop there. That is a mistake, and it is worth understanding why.

People do not misuse a prompt because they ignored the intended use. They misuse it because their situation looks close enough to the intended use that the stretch feels reasonable. The person applying a deviation drafting prompt to a supplier complaint is not being careless. They are reasoning by similarity, which is normally a good instinct. The out-of-scope field is what interrupts that reasoning at the moment it matters, and it only works if it names the specific adjacent cases rather than offering a general caution.

Write it as a list of concrete prohibitions with a short reason attached to each. “Do not use for supplier complaint investigations: the evidence structure differs and the prompt will assume a manufacturing root cause taxonomy that does not apply.” That sentence does more governance work than a page of policy.

How much evaluation evidence is enough

This is where organizations either build something sustainable or build something they abandon. The instinct in a regulated environment is to over-specify, and over-specified evaluation is evaluation that does not happen.

For a Tier 2 asset, a defensible evaluation set is small: a representative set of inputs covering the normal case, two or three known edge cases, and at least one case where the correct behavior is to decline or flag rather than produce output. The evaluator can be a qualified human. It does not have to be automated. What matters is that the set is written down, that it is rerun when the prompt or the model changes, and that the results are compared rather than just observed.

Where automated evaluation is used, be careful about the reliability of the judge. Research on using language models to evaluate other models’ output found meaningful agreement with human preference in some settings but also documented position bias, verbosity bias, and limited ability on tasks requiring reasoning or exact grading.7 An automated judge is a screening tool that reduces human review effort. It is not a substitute for a qualified reviewer on a regulated output.

A note on prompt structure as a quality control. Some of the variability in prompt behavior comes from unstructured, ad hoc phrasing. The published work on prompt patterns and prompting techniques provides a shared vocabulary for the structures that reliably do specific jobs, such as constraining output format, requiring the model to ask clarifying questions, or forcing an explicit refusal path.56 Standardizing on a small set of patterns inside a library reduces the review burden, because reviewers stop assessing novel prose every time and start assessing whether a known pattern was applied correctly.

The Coupling Problem: A Prompt Without a Model Version Means Nothing

This is the part of prompt governance that most organizations have not thought through, and it is the part that makes prompt governance genuinely different from document control.

A prompt has no behavior on its own. It has behavior only when paired with a specific model. Change the model and you have changed the thing you evaluated, even though every character of the prompt is identical. This is not a theoretical concern that might matter someday. It is the observed behavior of commercial models over the periods that pharma document lifecycles operate on.

The most widely cited demonstration compared two versions of the same commercially available models three months apart on identical tasks. On a prime number identification task, one model dropped from 84 percent accuracy in the March version to 51 percent in the June version.1 Other tasks moved in the other direction. The point is not that models get worse. The point is that they change, in ways the consumer of the model does not control and is not notified about at the level of individual task performance.

Model retirement is scheduled, and the schedule is not yours

The second half of the coupling problem is that the model you qualified against will eventually be withdrawn, on a timetable set by the provider.

The published policies are specific enough to plan around. One major provider commits to at least six months of notice for generally available models, at least three months for specialized variants such as chat-tuned or code-focused builds, and states that preview models may be retired with as little as two weeks of notice.17 Another commits to at least 60 days of notice before retiring a publicly released model, and publishes a table of tentative retirement dates alongside recommended replacements.16 That provider’s published history shows models moving from deprecation announcement to retirement in windows of roughly two to four months.16

Set that against a pharma periodic review cycle, which is commonly annual or biennial. If your prompt assets are reviewed once a year and your model provider retires models on a two to six month notice, your review cycle cannot be the mechanism that catches model change. Something else has to.

The mechanism almost nobody has. Ask your organization a simple question: when the model version behind your regulated AI use changes, what fires? In most organizations the honest answer is nothing. The platform team updates an endpoint or accepts a provider default, the change is invisible to quality, and the prompt assets qualified against the previous version continue to be used with no reassessment. The prompt register and the model register exist in different systems, owned by different functions, with no link between them.

Building the link between model change and prompt revalidation

The fix is not complicated, but it does require someone to own it. Four elements make it work.

1

Pin the version, never the alias

Regulated use pins a specific model snapshot identifier, not a floating alias that points to whatever the provider considers current. Floating aliases are appropriate for Tier 0 and Tier 1 use. For Tier 2 and Tier 3, an alias means your qualified configuration can change without any action on your side, which is the definition of an uncontrolled change.

2

Maintain a model-to-prompt dependency map

One list showing which prompt assets are qualified against which model versions. This can start as a column in the prompt register. Its only job is to answer, in seconds, the question “which of our regulated prompts are affected if this model is retired?” Without it, every provider deprecation notice becomes a manual investigation.

3

Treat a model version change as a change control trigger

Write it into the procedure explicitly: a change to the model version behind a Tier 2 or Tier 3 prompt asset requires the evaluation set to be rerun and the results compared to the qualified baseline before the new version is used for regulated work. This is the sentence that connects two control systems that are currently disconnected.

4

Subscribe someone to the deprecation feed

Provider deprecation notices are published and, for named customers, emailed. Someone in the organization needs to receive them, and that person needs to be connected to the dependency map. This is a small assignment that prevents a specific and otherwise certain failure: discovering a retirement after the endpoint has already stopped responding.

There is a version of this problem that is subtler and worth naming. Providers sometimes change the behavior of a model without changing its identifier, through updates to safety layers, system-level instructions, or serving infrastructure. Version pinning reduces exposure to this but does not eliminate it. For Tier 3 uses, the answer is periodic re-execution of a small regression set against the pinned version and comparison to the stored baseline. It is the same logic as any periodic verification of a validated system, applied to a component that happens to be operated by someone else.

Beyond Prompts: The Rest of the Reusable Asset Inventory

Prompts are the visible part of this problem and the easiest to explain, which is why they get the attention. They are not the largest part. Once an organization moves past chat-style use into retrieval systems and agents, the number of reusable artifacts that shape regulated output multiplies, and none of them live in a system of record today.

Retrieval configurations

A retrieval-augmented system finds relevant source material and puts it in front of the model. Every parameter of that process shapes what the model sees, and therefore what it produces. The chunking strategy, the chunk size and overlap, the embedding model and its version, the number of results retrieved, the reranking step, the similarity threshold, and the filters applied for access control all constitute a configuration.

That configuration is at least as consequential as the prompt. Research on how models use long contexts found that performance is highest when relevant information appears at the beginning or end of the input and degrades noticeably when the model has to find it in the middle.4 A change to how many documents you retrieve or how you order them can therefore change output quality without anyone touching the prompt or the model. The evaluation methods for retrieval systems have matured to the point where component-level measurement of faithfulness and relevance is practical.8 What has not matured is anyone treating the retrieval configuration as a controlled artifact with a version and an owner.

There is a second dimension here that is specific to regulated environments. The content behind the retrieval index changes. SOPs get revised, specifications get superseded, guidance documents get withdrawn. A retrieval configuration that indexes a document library without respecting effective dates and superseded status will confidently surface withdrawn content. The index refresh policy is part of the configuration and belongs in the same record.

Evaluation datasets

The evaluation set is the reference standard for every claim you make about a prompt or a system. In any other part of a quality organization, a reference standard is itself controlled. In AI programs it is usually a spreadsheet on someone’s drive.

An evaluation dataset needs a version, because comparing this quarter’s results against last quarter’s is meaningless if the set changed in between. It needs provenance, because a set assembled from real records carries data classification and privacy obligations that follow it wherever it is stored. It needs a stated coverage rationale, because a set that only contains easy cases will produce results that look excellent and mean nothing. And it needs protection from contamination, because an evaluation set that has been used to iteratively tune a prompt is no longer measuring generalization.

A practical split worth adopting early. Maintain two evaluation sets per regulated asset: a development set that authors can see and iterate against, and a smaller held-back set that only the reviewer runs. The held-back set is what goes in the evidence record. This costs very little to set up at the start and is close to impossible to retrofit once a prompt has been tuned against everything you have.

Output schemas

Where a generative system produces structured output, the schema defining that structure is a reusable asset with real consequences. It determines what fields exist, what is required, what is optional, and what the downstream system will accept. A schema change ripples into every consumer of that output.

Schemas are the one item on this list that development teams sometimes already version properly, because they live in code. The gap is usually not versioning. It is that the schema version is not linked to the prompt version or captured with the output, so a record produced under an older schema cannot be interpreted against the current one without archaeology.

Agent tool definitions

This is the newest item and the one with the sharpest edge. In agent architectures, the model is given a set of tools it can call, each with a name, a description, and a parameter specification. The description is not documentation. It is the instruction the model uses to decide when to call that tool and with what arguments. Interoperability standards for connecting models to tools and data sources make these definitions portable across systems, which increases both their usefulness and their reach.18

A tool definition is therefore a prompt with executable consequences. Rewriting a tool description from “retrieves batch records” to “retrieves and updates batch records” changes system behavior in a way no code review would necessarily catch, because no code changed. And because tools can write as well as read, the blast radius is different from a prompt that only produces text.

The security dimension compounds this. Established taxonomies of adversarial attacks against AI systems describe how instructions embedded in retrieved or supplied content can redirect model behavior, and how the boundary between instruction and data is weaker in these systems than in conventional software.11 Secure development guidance for generative AI systems adds practices around provenance and integrity of the components that make up an AI system.10 Both of these point at the same conclusion: the definitions that grant an agent capability need the same change control as the code that implements them, and arguably tighter review, because they are easier to change and harder to test.

Reusable asset What changing it silently affects Where it usually lives today Minimum control to add first
Prompt Structure, tone, and completeness of every downstream document Chat history, wiki page, saved assistant instructions Owner, intended use, model version, change history
Retrieval configuration What source material the model sees, and whether superseded content surfaces Application config file or platform console settings Versioned config record plus index refresh and effective-date policy
Evaluation dataset Every performance claim made about the system Spreadsheet on an individual’s drive Version, provenance, coverage rationale, held-back split
Output schema What downstream systems accept and how historical records are interpreted Application code repository Link the schema version to the prompt version and to the stored output
Agent tool definition When the agent acts, on what, and with what authority Agent configuration, often edited without code review Treat as controlled configuration with review by someone other than the author

Where These Live and How People Actually Find Them

Everything above is wasted effort if the library is somewhere nobody goes. This is not a minor implementation detail. It is the single most common reason prompt governance programs produce a register that describes a fictional organization while the real work happens somewhere else.

The failure pattern is predictable. The library gets built in the document management system because that is where controlled things live. Finding a prompt requires knowing its document number or navigating a folder structure designed for SOPs. Copying the prompt out requires opening a PDF. Nobody uses it. Within two quarters, the actual working prompts are back in team channels and the register describes assets that have been superseded three times.

Separate the record of control from the point of use

The resolution is to stop treating these as the same thing. There are two distinct needs and they have different homes.

The record of control is the approved metadata: owner, intended use, out-of-scope uses, qualified model version, evaluation evidence, approval, and change history. This belongs wherever your organization keeps controlled records, because that is where an inspector will look and where retention is managed.

The point of use is where a person or a system actually retrieves the prompt at the moment of work. That needs to be inside the tool people are already using: the assistant’s own shared prompt space, the platform’s registry, or a searchable interface that sits one click from where the work happens. Purpose-built prompt registries now support versioning, aliasing, and linking prompts to evaluation runs directly.9 Where such tooling exists in your stack, use it as the point of use and let the controlled record reference it, rather than duplicating content into a document system that nobody will consult.

The connection between the two is an identifier. The controlled record names the asset and its version. The point-of-use copy carries the same identifier. Any output produced for Tier 2 or Tier 3 use captures that identifier alongside the record, which is what makes the retrospective question answerable.

Capture the identifier at the moment of output, not afterward. The single highest-value habit to establish is that regulated AI-assisted output carries the prompt asset identifier and version, and the model version, in its metadata or in the record itself. This is a small technical change with an outsized effect. It converts “we think this was produced with version 3” into a fact, and it is the difference between an explainable process and an unexplainable one.

Findability is the actual requirement

People do not bypass libraries out of defiance. They bypass them because searching failed and rewriting was faster. If a regulatory writer cannot find an approved prompt for a task in under a minute, they will write their own, and that new prompt is now an unregistered asset that will be shared with two colleagues by Friday.

Practical measures that matter more than they sound:

  • Name assets by task, not by system or department. “Deviation investigation summary draft” is findable. “QA-AI-004” is not, and neither is “Manufacturing Excellence Prompt Set v2.”
  • Support search by the words people actually use. Tag with synonyms and near-miss terms. Someone looking for a “CAPA effectiveness check” should find the asset named “CAPA effectiveness review summary.”
  • Show the out-of-scope list in the search result, not three clicks deeper. The moment of highest attention is the moment of selection.
  • Make contribution possible in minutes. If proposing a Tier 1 asset takes twenty minutes of form filling, nothing gets contributed and the library stays thin, which reinforces the belief that searching it is pointless.
  • Retire aggressively. A library full of stale assets is worse than a small one, because a failed search in a large library teaches people the library is unreliable.

Shadow prompts are the same failure mode as shadow AI

The parallel is exact and worth making explicitly to leadership, because it maps onto a risk they already understand.

Shadow AI happens when the sanctioned tool is worse than the unsanctioned one, so people use the unsanctioned one. Shadow prompts happen when the sanctioned library is harder to use than a private note, so people keep private notes. In both cases the organization loses visibility, not because anyone acted badly, but because the governed path was more effort than the ungoverned one.

The response is the same in both cases too. Make the governed path genuinely better: faster to search, pre-qualified so the user does not have to think about model versions, and carrying the out-of-scope guidance that saves someone from a mistake. Governance that competes on usefulness gets adopted. Governance that competes on obligation gets routed around.

The generative AI risk management guidance published by NIST is useful here because it frames a set of these concerns as configuration and value chain issues rather than purely model issues, including the difficulty of tracking third-party components and their changes through a system.20 An organization that cannot enumerate its prompts and retrieval configurations cannot make any credible statement about its generative AI risk posture, whatever its policy documents say.

Running the Library Without Killing the Usefulness

The design above only works if someone runs it, and the running of it is where good intentions usually stop. A few decisions determine whether this becomes a durable capability or a project that produced a spreadsheet.

Who owns it

Prompt asset governance sits awkwardly between quality, IT, and the functions doing the work. Putting it entirely in quality produces control without usefulness. Putting it entirely in IT produces a tool without authority. Putting it entirely in the business functions produces four incompatible libraries.

The arrangement that tends to hold is a small central function that owns the register, the tiering criteria, the evaluation approach, and the model dependency map, with asset ownership distributed to the functions that use them. Central owns the system and the standard. Functions own the content. Quality reviews Tier 2 and Tier 3 assets against the standard rather than authoring them.

What review actually looks like

Review of a prompt asset is not review of prose. A reviewer working through a Tier 2 asset should be checking a short, specific list:

  1. Is the intended use stated concretely enough that a user can tell whether their case matches?
  2. Does the out-of-scope list name the plausible adjacent misuses, not just the obvious ones?
  3. Does the prompt include an explicit path for the model to decline or flag insufficient input, rather than always producing output?
  4. Is the model version recorded, and does the evaluation evidence correspond to that version?
  5. Does the evaluation set include at least one case where the correct behavior is not to produce a confident answer?
  6. Does the output carry the information a downstream reviewer needs to check it, including any source references the prompt was told to include?

Six questions. A competent reviewer works through this in twenty minutes for most assets. That is a sustainable review burden. A three-hour validation-style review of a paragraph of English is not, and organizations that try it stop after the fourth asset.

Periodic review that is actually periodic

Set review intervals by tier, not uniformly. Tier 1 assets can be reviewed annually or on a use-triggered basis. Tier 2 assets warrant an annual review plus a review on any model version change. Tier 3 assets warrant more frequent review and continuous monitoring of production output against the qualified baseline.

The review itself should include a question most periodic reviews omit: is this asset still used? Usage data from the point-of-use system answers it. Assets with no use in twelve months should be retired rather than reapproved, because carrying them dilutes the library and adds review load for no benefit.

A reasonable first ninety days. Inventory what exists by asking each function to list the prompts they share with colleagues, without judgment and without asking anyone to justify them. Assign tiers to what comes back. Register the Tier 2 and Tier 3 assets with full metadata, and accept that this will be a small number, probably between ten and forty in a mid-sized organization. Build the model dependency map for those. Leave Tier 0 and Tier 1 alone except to give them a searchable home. That is a realistic first pass, and it produces something defensible without a program that consumes a year.

What to avoid

Three patterns reliably produce a program that fails.

Registering everything. A register with four hundred entries, most of them personal-use prompts nobody has touched in months, is not evidence of control. It is a maintenance obligation that will be abandoned, and the abandonment is worse than never having started, because now there is a stale controlled record contradicting reality.

Approving prompts without evaluation evidence. A signature on a prompt with no attached evidence is a governance artifact that provides no assurance. It looks like control and functions as ceremony. Given the demonstrated sensitivity of model behavior to phrasing and format, a reviewer reading a prompt genuinely cannot tell how it will perform.23

Treating the library as a documentation exercise. If the deliverable is a document describing the prompt governance framework, the program has already failed. The deliverable is a working register, a searchable point of use, a dependency map that fires when models change, and a habit of capturing the asset identifier with the output. The document describing all of this should be short and should come last.

Conclusion

The organizations that get this right are not the ones with the most thorough policy. They are the ones that recognized early that a shared prompt is an instruction that shapes regulated output, that an instruction shaping regulated output has always been something pharma controls, and that the only genuinely new problem is the coupling between the instruction and a model version that someone else operates and will eventually retire. Everything else is the application of change control principles that quality organizations have used for decades, scaled down to fit an artifact that takes ten seconds to modify.

The design decision that matters most is the tiering. Control that is proportionate to what the output touches will be followed. Control that treats every prompt as a controlled document will be routed around, and the routing around is invisible, which makes it worse than having no program at all. Start with the small set of assets that genuinely feed regulated records, give them real metadata and real evaluation evidence, connect them to the model versions they depend on, and put them somewhere people can find in under a minute. Leave everything else alone and say so out loud.

Sakara Digital works with pharma and biotech organizations building this kind of practical AI governance, including the parts that sit between quality and IT and tend to get left out of both. If you are working out how to bring prompts, retrieval configurations, and agent tool definitions under proportionate control without slowing down the people using them, we are happy to have that conversation.

For Further Reading