In This Article
- Executive Summary
- How Token Pricing Works and Why Early Spend Is Hard to Predict
- What Drives Token Consumption
- The Document-Heavy Line Item That Surprises Pharma and Biotech
- Pricing Structures Differ More Than the Headline Rate
- How to Budget: Start From the Task, Not the Token
- How to Monitor and Cap Spend Across Teams
- Levers That Lower Spend Without Lowering Quality
- Token Spend in GxP Settings: Validation, Change Control, and Hard Caps
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Most enterprise AI tools that run on large language models are billed by the token, a small unit of text that the model reads or writes. That pricing model makes spend easy to start and hard to forecast. The first months of a rollout rarely look like steady state, because consumption depends on how people use the tool, how the application is built, and which model version is behind it. Enterprise AI token cost is less a matter of the headline price per million tokens and more a matter of how many tokens each task consumes.
For pharma and biotech, the item that surprises finance teams is document-heavy work. Regulatory documents are long. Using published vendor pricing and documentation as listed in September 2026, our illustrative arithmetic shows a single pass over a 300-page document can cost 120 to 230 times more than a short chat exchange, and an agent that re-reads that document across many steps multiplies it again.
This article explains how token pricing works, what drives consumption, and how pricing structures differ across vendors. It then sets out a practical approach: budget per task rather than per token, separate spend by team with limits and alerts, use the levers that lower spend without lowering quality, and treat token spend as part of validation and change control in GxP settings.
How Token Pricing Works and Why Early Spend Is Hard to Predict
Every major model provider prices its application programming interface (API) access the same basic way: a price per million input tokens (what you send to the model) and a higher price per million output tokens (what the model writes back). A token is a piece of a word. Anthropic’s pricing documentation gives a rough rule of thumb that one token is about four characters, or about 0.75 words, of English text, and notes that the exact count varies by language and content type1.
That sounds simple, and at the level of a single request it is. The difficulty is that almost nobody in a business thinks in tokens. A finance lead thinks in dollars per month. A regulatory affairs lead thinks in documents. An IT lead thinks in users and licenses. The FinOps Foundation, the industry body for cloud financial management, makes this point directly in its June 2026 paper on token economics: finance teams and many senior engineering leaders have no intuition for what a token costs or how many tokens a typical request consumes2.
Input, Output, and Reasoning Tokens
Three categories of tokens show up on the bill, and they are not priced the same way.
- Input tokens are everything sent to the model: the system instructions, the user’s question, any attached documents, the conversation so far, and the definitions of any tools the model is allowed to call.
- Output tokens are what the model writes. They are priced higher than input tokens, typically five to six times higher on the price lists we reviewed.
- Reasoning or thinking tokens are the intermediate work some models do before they answer. OpenAI’s documentation states that while reasoning tokens are not visible through the API, they “still occupy space in the model’s context window and are billed as output tokens.”3 Anthropic’s documentation says the same of its thinking tokens4, and Google’s price list labels its output price as “including thinking tokens”5.
Reasoning tokens are a common source of surprise. A short visible answer can carry a large hidden bill. OpenAI notes that its reasoning models may generate anywhere from a few hundred to tens of thousands of reasoning tokens depending on the problem, and recommends reserving at least 25,000 tokens for reasoning and outputs when teams start experimenting3.
What the Price Lists Say Today
Prices change often, sometimes several times a year, so every figure below is dated. The table shows a sample of mid-tier and small models from three vendors, as listed on each vendor’s own pricing page in September 2026. It is not a recommendation of any model. It is here so readers can see the shape of the numbers.
| Model (as listed in September 2026) | Input, per million tokens | Output, per million tokens | Notes |
|---|---|---|---|
| Anthropic Claude Sonnet 5.51 | $2.00 | $10.00 | Cache reads $0.20; batch rates are half the standard rates |
| Anthropic Claude Haiku 4.51 | $1.00 | $5.00 | Smaller, faster model |
| OpenAI gpt-6-sol6 | $2.00 | $10.00 | Short-context rate; the long-context rate is listed at $4.00 input and $15.00 output |
| OpenAI gpt-5.4-mini6 | $0.75 | $4.50 | Smaller model |
| Google Gemini 3.1 Pro Preview5 | $2.00 (prompts up to 200K tokens); $4.00 above | $12.00; $18.00 above 200K | Output price includes thinking tokens |
| Google Gemini 3.8 Flash5 | $0.75 through Dec 31, 2026; $1.50 from Jan 1, 2027 | $3.75 through Dec 31, 2026; $7.50 from Jan 1, 2027 | Listed price doubles at the start of 2027 |
Two things stand out. First, the headline rates for comparable tiers are close to each other. Choosing a vendor on the price per million tokens alone rarely changes the budget much. Second, the rules around the headline rate differ a lot: long-context surcharges, scheduled price changes, cache pricing, and batch discounts. We return to those in a later section, because they matter more than the headline.
Why Month One Tells You Little About Month Six
Leaders often ask for a forecast after a pilot’s first month. That forecast is usually wrong, for reasons that have nothing to do with poor planning.
- Usage patterns are still forming. Early users try short questions. Once they trust the tool, they attach whole documents and run longer sessions.
- The application is still changing. Engineers add retrieval, tools, and agent steps, and each one adds tokens to every request.
- Model versions change. Anthropic’s pricing page notes that its Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text1. A model upgrade at the same listed price can still raise the bill.
- Prices change on a schedule. Google’s listed rates for Gemini 3.8 Flash double on January 1, 20275. A 2027 budget built from September 2026 rates would be half the real figure.
- Mistakes are expensive and fast. The FinOps Foundation paper describes how an agentic workflow that loops unexpectedly, a system prompt that was doubled in length by accident, or a feature that drives far more usage than expected can produce invoices that exceed monthly budgets in a single day2.
What Drives Token Consumption
When spend runs ahead of plan, the cause is almost always a change in how many tokens each task uses, not a change in price. Six drivers account for most of it. Each one is easy to see once you know to look for it.
Context Length
Everything in the request counts: instructions, attached files, retrieved passages, and prior turns. A long system prompt is paid for on every single call.
Conversation Accumulation
Each new turn resends the full history. The twentieth question in a session is billed for the nineteen that came before it.
Long Document Processing
A PDF page is typically 1,500 to 3,000 tokens of text before any image processing. Documents with hundreds of pages dominate the bill.
Retries and Regeneration
Failed format checks, truncated outputs, and users pressing “regenerate” all resend the full input. The retry pays the whole price again.
Agentic Loops
An agent calls the model many times per task: to plan, to call a tool, to read the result, to check its work. Each call carries the growing context.
Hidden Overhead
Tool definitions, tool-use system prompts, reasoning tokens, and paid add-ons such as web search appear in the bill but not on the screen.
Context Length and Conversation Accumulation
Anthropic’s documentation on context windows describes the mechanism plainly. As a conversation moves through turns, each user message and each model response accumulates in the context window, and previous turns are preserved completely. Each turn’s input contains all previous conversation history plus the current message, and the response becomes part of the input for the next turn4. The same page lists what counts toward the window: the system prompt, every message including tool results, images, and documents, and the tool definitions4.
The effect on spend is that the price of a conversation grows faster than the number of turns. A ten-turn session where each turn adds 2,000 tokens does not cost ten times a single turn. The last turn alone carries about 20,000 tokens of history, and the total input across the session is the sum of all ten growing requests. Long sessions are where chat tools become expensive.
Long Document Processing
Documents are the largest single driver in regulated industries, and we give them their own section below. The key fact is this one. Anthropic’s PDF documentation states that each page typically uses 1,500 to 3,000 tokens depending on content density, and that because each page is also converted into an image, image-based token charges apply on top8. So the text alone of a 100-page document is roughly 150,000 to 300,000 tokens, and the page images add more.
Retries and Regeneration
A retry is not a small follow-up. It is a full new request. If an application asks a model to extract data from a 200-page document into a fixed format, and the output fails a format check, the retry sends the 200 pages again. OpenAI’s documentation describes a related case for reasoning models: if the model reaches its output limit before producing visible text, the response comes back incomplete, and “you could incur costs for input and reasoning tokens without receiving a visible response.”3 Paying for a failed attempt and then paying again for the retry is one of the most common patterns we see in early usage data.
In chat tools the same thing happens through people. A user who does not like an answer presses “regenerate,” or pastes the same long document into a new chat to start fresh. Neither shows up as an error. Both show up on the invoice.
Agentic Loops
An agent is an application where the model decides what to do next, calls tools, reads the results, and repeats until the task is done. That loop is the point of an agent, and it is also what drives its token use. Anthropic’s engineering team, writing about its own multi-agent research system, reported that in its data agents typically use about four times more tokens than chat interactions, and multi-agent systems use about 15 times more tokens than chats7. The same post notes that token usage by itself explained 80% of the variance in performance on the browsing evaluation they studied7. In other words, agents often perform better partly because they spend more.
The post draws the practical conclusion: multi-agent systems make economic sense for tasks where the value of the task is high enough to pay for the increased performance7. That is the right test for any agent in a regulated company. The FinOps Foundation paper adds the risk side: an agent that calls itself, or tool chains that recurse without end, can consume millions of tokens in minutes2.
Hidden Overhead
Some tokens never appear on screen. Anthropic’s pricing page lists a tool-use system prompt that the API adds whenever tools are provided, in the range of roughly 290 to 800 tokens for recent models depending on the model and settings, plus the tokens for the tool names, descriptions, and schemas themselves1. Browser and computer-use toolsets add several thousand input tokens per request. Paid add-ons carry their own meters: Anthropic lists web search at $10 per 1,000 searches plus the tokens for the results1, and Google lists grounding with Google Search for Gemini 3.x models at $14 per 1,000 requests after 5,000 free requests per month5. None of these are large on a single call. Across an agent that makes dozens of calls per task and thousands of tasks per month, they add up.
The Document-Heavy Line Item That Surprises Pharma and Biotech
Most public discussion of AI pricing is built around chat and customer support, where a request is a few hundred to a few thousand tokens. Life sciences work is different. The documents that matter most are long, and many of the highest-value use cases involve reading them in full: summarizing a clinical study report, checking a submission section against its source data, drafting a response to a health authority question, reviewing a deviation investigation against the relevant procedures, or screening literature for safety signals.
How Long Regulatory Documents Are
Clinical study reports (CSRs) are the clearest example. A CSR is the full report of a clinical trial prepared for regulators, including the protocol, the statistical analysis plan, tables, and patient listings. A 2013 review in BMJ Open by Peter Doshi and Tom Jefferson assembled 78 CSRs covering 90 randomized controlled trials of 14 medicines, dated 1991 to 2011, totaling 144,610 pages9. By our arithmetic that is an average of about 1,850 pages per report. The review reported median lengths for individual sections, including 337 pages of attached tables, 62 pages of trial protocol, and 447 pages of individual efficacy listings9. The authors describe their sample as non-random and say they cannot tell whether it is representative9, so treat the average as an illustration of scale, not a benchmark.
Apply the per-page figure from Anthropic’s PDF documentation (1,500 to 3,000 text tokens per page8) to an 1,850-page report and the text alone comes to roughly 2.8 million to 5.6 million tokens. That is larger than the 1-million-token context window of Anthropic’s current models4, and Anthropic’s PDF support caps a single request at 600 pages8. A full CSR cannot be read in one call. It has to be split, searched, or summarized in stages, and every one of those design choices changes the token count.
An Illustrative Comparison
The table below uses Claude Sonnet 5.5 rates as listed in September 2026 ($2.00 per million input tokens, $10.00 per million output tokens, $2.50 per million for a five-minute cache write, and $0.20 per million for a cache read)1. The token counts are our own assumptions for illustration, not measurements from any company. The point is the ratio between rows, which holds at any vendor with similar pricing.
| Task (illustrative) | Assumed tokens | Price per task |
|---|---|---|
| Short chat question and answer | 1,500 input, 500 output | About $0.008 |
| One pass over a 300-page document, 4,000-token summary | 450,000 to 900,000 input (text only), 4,000 output | About $0.94 to $1.84 |
| Same document, sent through the Batch API at half price | Same as above | About $0.47 to $0.92 |
| Agent that reads the document across 12 model calls, no caching | 12 × 450,000 input, 12 × 2,000 output | About $11.04 |
| Same agent with prompt caching (one cache write, 11 cache reads) | Same volume, cached | About $2.36 |
By this arithmetic, one pass over a 300-page document costs roughly 120 to 230 times the short chat exchange. An agent that re-reads the same document on every step, without caching, costs more than ten times the single pass, even before counting the growth of its own history. With caching, the same agent comes back down to between two and three single passes. The agent rows use the low end of the page estimate and ignore history growth, so real figures would be higher.
Now multiply by volume. At 2,000 documents a month, the single-pass design comes to roughly $1,880 to $3,680 a month on these assumptions. The uncached agent design comes to about $22,000. The cached agent comes to about $4,700. Same model, same documents, same listed price, and a spread of more than ten times depending on how the application is built.
Why this line item surprises people. Budgets for enterprise AI are often set from a pilot where people asked short questions. The first document-heavy workflow to go live, such as CSR summarization, submission quality checks, or literature screening, can use more tokens in a week than the whole pilot used in a quarter. Nothing is broken. The work simply involves reading far more text. Ask for a token estimate per document before any document-heavy use case goes live, and budget it as its own line.
Where Document-Heavy Use Cases Show Up
In our experience, these are the pharma and biotech workflows where document volume drives the bill:
- Regulatory writing and review: drafting or checking sections of a submission against source reports, where the source material runs to hundreds or thousands of pages.
- Clinical documentation: CSR summaries, protocol amendment impact reviews, and consistency checks across the protocol, statistical analysis plan, and report.
- Quality and manufacturing: deviation and CAPA (corrective and preventive action) investigations that pull in batch records, procedures, and past investigations.
- Pharmacovigilance and medical information: literature screening and case narrative work at high volume.
- Due diligence: reviewing data rooms for licensing or acquisition, where hundreds of documents are read once under time pressure.
None of these is a reason to avoid AI. Several are among the strongest use cases in the industry. They are a reason to design for document volume from the start, which is where most of the savings are. For more on the writing use cases specifically, see our earlier article on generative AI for regulatory writing.
Pricing Structures Differ More Than the Headline Rate
When procurement compares AI vendors, the comparison usually starts with the price per million tokens. For long-document work, the rules around that price often matter more. Five differences deserve attention.
Long-Context Surcharges
Some vendors charge a higher rate once a request passes a size threshold. Google lists Gemini 3.1 Pro Preview at $2.00 per million input tokens for prompts up to 200,000 tokens and $4.00 above that, with output rising from $12.00 to $18.005. OpenAI’s price list shows separate short-context and long-context columns, with gpt-6-sol at $2.00 input and $10.00 output for short context and $4.00 input and $15.00 output for long context6. Anthropic, by contrast, states that its Claude 4.6 and later models include the full 1-million-token context window at standard pricing, so a 900,000-token request is billed at the same per-token rate as a 9,000-token request1. For a use case built around long regulatory documents, these rules can matter more than a small difference in the base rate.
Tokenizers
A token is not a fixed unit across vendors or even across model versions. Each model family splits text in its own way. As noted above, Anthropic states that its newer tokenizer produces approximately 30% more tokens for the same text1. Two models with the same listed rate can therefore charge different amounts for the same document. The only reliable comparison is to run a sample of your own documents through each candidate and read the token counts from the responses.
Scheduled and Promotional Prices
Price lists now carry dates. Google lists Gemini 3.8 Flash at one rate through December 31, 2026 and double that rate from January 1, 20275. OpenAI’s pricing page notes promotional pricing for one of its models “available at least through November 21, 2026.”6 Anthropic’s page records that introductory pricing for Claude Sonnet 5 became the standard price and a scheduled increase will not occur1. Prices move in both directions. Any budget that crosses a calendar year should check each vendor’s page for dated changes, and any contract should say which price applies.
Data Residency and Regional Premiums
Many life sciences companies need data to stay in a given country or region. That often carries a premium. Anthropic lists a 1.1 times multiplier on all token categories for US-only inference on Claude 4.6 and later models, and notes that regional endpoints on Amazon Bedrock and Google Cloud carry a 10% premium over global endpoints1. OpenAI’s pricing page states that FedRAMP endpoints and eligible data residency endpoints are charged a 10% uplift6. If your security or privacy review requires regional processing, add the premium to the budget from day one.
Seats, Credits, and Reserved Capacity
Not every enterprise AI tool is billed per token. Three other models are common, and many companies use all three at once.
| Billing model | How it works | Where the surprise comes from |
|---|---|---|
| Per-token API | Pay for input and output tokens used | Consumption per task grows with documents, agents, and retries |
| Per-seat license with included usage | Fixed monthly fee per user; some features are included | Agents and add-ons built on top may draw from a separate consumption meter |
| Prepaid or pay-as-you-go credits | Credits consumed per action, varying by task complexity | Unused credits may expire; exceeding capacity can lead to enforcement |
| Reserved capacity | Pay per hour for dedicated throughput, whether used or not | Idle capacity is still billed; too little capacity leads to throttled requests |
Microsoft’s Copilot Studio documentation, last updated in August 2026, is a good example of the credit model. It describes Copilot Credits as the common currency for agents, says the number of credits counted for each response depends on the complexity of the task, states that unused credits do not carry over to the next month, and warns that if usage exceeds purchased capacity, technical enforcement applies and can result in service denial10. The same page notes that for users with a Microsoft 365 Copilot license, certain agent answers inside Microsoft 365 apps are zero-rated and do not draw on the credit pool10. The practical lesson is that a seat license does not mean a fixed bill once agents are involved.
Reserved capacity works the other way. Microsoft’s documentation on provisioned throughput in Microsoft Foundry explains that provisioned deployments are billed at an hourly rate per provisioned throughput unit (PTU), “regardless of the number of tokens consumed,” and that 1-month or 1-year reservations give a discounted rate11. The same page recommends provisioned throughput for predictable traffic and standard pay-per-token deployments for development, testing, low volume, or highly variable traffic11. That guidance fits the budgeting approach below: pay per token while you learn the pattern, and consider reserved capacity only once the pattern is stable.
How to Budget: Start From the Task, Not the Token
The most useful change a leadership team can make is to stop budgeting AI in tokens and start budgeting it in tasks. “We expect to process 2,000 deviation investigations a year at about $X each” is a number a CFO, a quality head, and an engineer can all discuss. “We expect 400 million tokens” is not. The FinOps Foundation paper makes the same recommendation in its own terms: unit cost metrics such as cost per query, per user, or per outcome make AI spend legible to business stakeholders in a way raw token counts cannot2.
We suggest a six-step approach for each use case.
Define the Task Unit
Name the unit of work the business cares about: one CSR summary, one submission section check, one investigation draft, one literature screen. Estimate the monthly volume with the business owner, not the engineering team.
Measure Tokens Per Task on Real Documents
Run a representative sample of real documents through the actual application and read the input, output, cached, and reasoning token counts from each response. Vendors provide token counting tools for estimates before sending, but measured usage from real runs is better. Include the short, typical, and longest documents, because the long ones drive the bill.
Price It With Dated Rates
Multiply by the vendor’s listed rates and write down the date you read them, for example “Sonnet 5.5 rates as listed in September 2026.” Include any long-context, residency, cache, or batch adjustments. Check each vendor’s page for scheduled changes inside the budget period.
Add Allowances for Retries, Testing, and Growth
Retries, regeneration, and development runs are real spend. So is validation testing in GxP use cases, and so is re-testing after a model version change. Measure the retry rate during the pilot rather than assuming it is zero, and plan for usage to grow as people trust the tool.
Express the Budget as a Range
Give leadership a low, expected, and high figure per month, with the assumptions behind each. A range is more honest than a single number in the first six months, and it tells finance how much variance to hold in reserve.
Re-Forecast Monthly Until the Pattern Is Stable
Compare actual cost per task with the estimate each month. When three months in a row fall inside the range, the use case is ready for a normal annual budget line, and possibly for reserved capacity if volume is high and steady.
A Worked Budget Line
Suppose a regulatory team wants AI support for quality checks on submission sections. The business owner expects about 150 sections a month. Suppose measurement on twenty real sections shows an average of 250,000 input tokens and 6,000 output tokens per check, a retry rate of about one in ten, and no need for real-time answers. At Sonnet 5.5 batch rates as listed in September 2026 ($1.00 input and $5.00 output per million tokens1), one check is about $0.25 for input and $0.03 for output, or about $0.28. With retries it is about $0.31. At 150 checks, the monthly figure is under $50. That is small. The same check run as a 15-step agent without caching, at standard rates, would be many times larger. The budget conversation should be about design, not vendor.
Budget per task, report per team, review per quarter. The unit that finance approves is the task. The unit that operations watches is the team or use case. The unit that leadership reviews is the quarter, with the price list dates noted. Keeping those three levels separate stops the most common argument in AI budgeting: a monthly invoice that nobody can explain.
How to Monitor and Cap Spend Across Teams
Budgets only work if spend can be seen and, when needed, stopped. The good news is that the major vendors now provide the basic controls. The less good news is that they do not go as far as a finance team expects, so some of the work falls to you.
Separate Spend by Team and Use Case
The first control is structural: give each team or use case its own container for API keys, usage, and limits. On Anthropic’s platform that container is a workspace. The documentation describes using workspaces to separate projects, environments, or teams while keeping centralized billing, and lists development, staging, and production as a common split12. OpenAI uses projects for the same purpose13. On Amazon Bedrock, application inference profiles let you attach tags to a model endpoint and track costs through AWS cost allocation tags14.
This separation is what makes chargeback or showback possible. The FinOps Foundation paper notes that an invoice from a model provider typically shows aggregate token consumption broken down at most by API key or project, and that there is “no native concept of business unit, cost center, application, or workload.”2 If you want to know what the pharmacovigilance team spent, you have to have given the pharmacovigilance team its own workspace, project, or tagged profile before the spend happened.
Know the Difference Between an Alert and a Cap
Vendors offer two kinds of spend control, and they behave very differently.
- Alerts notify someone when spend crosses a threshold. Traffic keeps flowing.
- Hard limits stop traffic. OpenAI’s documentation says that when tracked spend reaches a hard limit, affected API requests return a 429 error, and warns that “Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount.”13
Anthropic’s workspaces support both monthly spend limits and alerts at chosen thresholds, as well as rate limits on requests, input tokens, and output tokens per minute12. Workspace limits can be set lower than the organization’s limits but not higher, and cannot be set on the default workspace12. That last detail matters: if teams are working in the default workspace, you cannot cap them individually.
The FinOps Foundation paper points out a gap between these controls and how finance thinks: native rate limits are expressed in tokens per minute, not dollars per month2. A rate limit protects against a runaway loop. It does not keep a team inside its quarterly budget. You need both.
| Control | What it protects against | Where to set it |
|---|---|---|
| Monthly spend alert at 50%, 80%, 100% of budget | Slow overspend that nobody notices until the invoice | Vendor console per workspace or project |
| Hard monthly spend limit | Large overspend in non-critical work (development, experiments) | Vendor console; use with care in production |
| Tokens-per-minute rate limit | Runaway loops and sudden spikes | Vendor console per workspace or project |
| Maximum output tokens per request | Unexpectedly long or looping outputs | Application code |
| Maximum steps per agent task | Agents that never finish | Application or agent framework |
| Per-task cost check in logs | Design changes that raise cost per task | Your own monitoring from usage data |
Watch Usage Daily, Not Monthly
Monthly invoices arrive too late to act on. Vendors now publish usage data that can be pulled much sooner. Anthropic’s Usage and Cost API returns token counts grouped by model, workspace, API key, and other dimensions in one-minute, one-hour, or one-day buckets, and cost in US dollars by day; the documentation says data typically appears within five minutes of a request completing15. It lists budget monitoring and cost attribution by workspace for chargebacks among the common uses15.
Pull that data into whatever your finance and IT teams already use, and look at three numbers per use case every week: total spend against budget, cost per task against estimate, and the cache hit rate. The FinOps Foundation paper adds a useful distinction for alerting: alerts at the account level tell you spend happened, while alerts at the workload level tell you which application or team caused it, and both layers are needed2. For runaway loops, it recommends monitoring tokens per minute at the application level, not just daily spend2.
Assign an Owner
Controls without an owner do not get used. Name one person per use case who receives the alerts, explains variances, and approves changes that affect cost per task. In most companies that should be the business owner of the use case, with support from IT. It should not be the vendor account manager.
Levers That Lower Spend Without Lowering Quality
Once you can see cost per task, you can lower it. The levers below are all documented by the vendors themselves, and most do not reduce output quality when applied with care.
Prompt Caching
Caching stores the processed form of a repeated prompt prefix, such as a long system prompt or a document that several requests will read. On Anthropic’s platform, a cache read costs 10% of the standard input price for most models, and a five-minute cache write costs 1.25 times the input price, so caching pays off after one cache read1. OpenAI enables caching by default for supported models, describes the cached-input rate as “discounted up to 90%,” and advises placing stable instructions and shared reference material first16. For document-heavy agents, caching is often the single biggest saving, as the illustrative table above shows.
Caching has conditions. Caches expire (five minutes or one hour on Anthropic’s platform1), and the cached part must be identical from one request to the next. An application that inserts a timestamp at the top of every prompt will never get a cache hit. This is a design question for the engineering team, and it is worth asking about in every architecture review.
Batch Processing
Much document work does not need an answer in seconds. Overnight processing is fine for literature screening, back-file summarization, or bulk quality checks. Anthropic’s Message Batches API charges 50% of standard prices, with most batches finishing in less than an hour and results available within 24 hours17. Google and OpenAI list the same 50% batch discount56. Anthropic notes that batch and caching discounts can be combined1.
Model Routing
Not every step needs the most capable model. Classifying a document, extracting a date, or checking a format can often run on a smaller model at a fraction of the price, while the drafting or reasoning step uses a larger one. Anthropic’s own pricing guidance suggests using smaller models for simple tasks, mid-tier models for most production work, and the largest models for the most complex reasoning1. Test quality at each step before routing, and record the routing decision, because in a GxP context it becomes part of the system’s design.
Send Less Text
The cheapest token is the one you never send. Four practices help:
- Retrieve, do not attach. For questions about a large document set, a retrieval step that finds the relevant passages is usually far cheaper than attaching whole documents.
- Extract text where images add nothing. Because PDF pages are also processed as images8, sending plain text for text-only pages avoids paying for page images the model does not need.
- Keep conversation history short. Anthropic documents server-side compaction and context editing, which summarize or clear earlier parts of a long conversation or agent session4. Starting a fresh session for a new topic does the same thing by hand.
- Trim tool lists. Only give an agent the tools it needs for the task. Every tool definition is paid for on every call.
Set Limits in the Application
Set a maximum output length for each request type, a maximum number of steps for each agent task, and a maximum retry count for each failure type. These limits protect the budget and also make the system more predictable, which matters for validation.
Where to start. For most document-heavy use cases, the order of impact is: design (retrieve rather than attach, cap agent steps), then caching, then batch for anything not time-sensitive, then model routing. Price negotiation with the vendor comes last. Anthropic, for example, notes that volume discounts are negotiated case by case1, and a well-designed application gives you a better starting position for that conversation.
For companies weighing whether some workloads should move off per-token APIs entirely, our article on small language models on-prem covers when that makes sense. It trades a variable token bill for fixed infrastructure and operating effort, which is a different kind of commitment, not a free one.
Token Spend in GxP Settings: Validation, Change Control, and Hard Caps
Most writing on AI spend ignores regulated work. In pharma and biotech, three issues connect token spend to quality and compliance.
Validation and Re-Testing Consume Tokens
Where an AI tool supports a GxP process, it is validated for its intended use, typically following a risk-based approach such as GAMP 5 (the ISPE guide for computerized system validation). That validation involves running test cases, often many times over, and keeping the evidence. Every test run is a model call and is billed like any other. When the model version changes, the tests run again. None of this is large compared with production volume for most use cases, but it belongs in the budget, and it should not come out of a team’s production allowance where a hard cap could block it.
Model and Tokenizer Changes Belong in Change Control
A model version change is already a change control event in a validated system, because behavior can change. It can also change cost. Anthropic’s note that its newer tokenizer produces approximately 30% more tokens for the same text1 is a clear example: the same prompt and the same document, on a newer model at the same listed price, can produce a larger bill. We recommend adding a cost-per-task comparison to the impact assessment for any model change, alongside the quality and behavior comparison. It takes very little extra effort when the test set already exists.
A Hard Cap Can Stop a Validated Process
Hard spend limits are a good control for experiments and development. In production GxP workflows they need more thought. If a hard limit is reached, the vendor returns an error and requests stop13. In a credit-based product, exceeding capacity can lead to service denial10. If an AI step is part of a validated process, a stop in the middle of a run is a process interruption that may need to be handled as a deviation, with an impact assessment on any affected records.
The safer pattern for GxP production work is to use alerts with enough headroom for someone to act, a documented procedure for what happens when an AI step is unavailable (including a manual fallback), and hard caps reserved for non-GxP workspaces. The business continuity plan for the process should cover this case as it would any other system outage.
A short checklist for GxP AI budgets:
- Validation and re-test runs have their own budget line, separate from production.
- Production GxP workspaces use alerts, not hard caps, with a named owner and a documented fallback.
- Model change impact assessments include a cost-per-task comparison.
- Price list dates are recorded with each budget and each contract.
- Agent step limits and output limits are part of the validated configuration, not left to defaults.
Two related governance points are covered in more depth elsewhere on our site. The agentic designs that drive the largest token counts are discussed in our article on agentic AI in pharma. The commercial side of committing to one vendor’s pricing, tokenizer, and roadmap is covered in our piece on AI vendor lock-in.
Conclusion
Enterprise AI token cost is predictable once you stop looking at it as a price and start looking at it as a volume. The listed rates for comparable models are close together, and they change often enough that any figure needs a date. What separates a manageable bill from a surprising one is the number of tokens each task uses, and that is set by design choices: how much text is sent, how often it is resent, how many steps an agent takes, and whether caching and batch processing are used. In pharma and biotech, the long documents at the center of regulatory, clinical, and quality work make those choices matter more than in most industries. A single design decision can move the monthly figure by ten times or more.
The companies that manage this well budget per task, separate spend by team from the first day, watch cost per task weekly during the first months, and treat model changes as both a quality event and a cost event. None of that requires new tools. It requires the same discipline life sciences already applies to other systems. Sakara Digital works with pharma and biotech organizations planning and governing enterprise AI, including the budgeting and controls that support it. If you are sizing a document-heavy use case or trying to explain an AI invoice that grew faster than expected, and want an independent perspective on where to start, we are happy to have that conversation.
For Further Reading
For Further Reading
- Small Language Models On-Prem: When Pharma Should Not Use a Frontier API
- The Hidden Cost of AI Vendor Lock-In in Regulated Life Sciences
- Agentic AI in Pharma: Moving Beyond Chatbots to Autonomous Workflows
- Generative AI for Regulatory Writing: Opportunities and Guardrails
- Measuring ROI from AI Investments in Life Sciences
References & Sources
- Anthropic. “Pricing.” Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/about-claude/pricing
- FinOps Foundation. “Tokenomics: Managing AI Value in SaaS Model Token Costs.” FinOps for AI Working Group, last updated June 3, 2026. https://www.finops.org/wg/token-economics-saas/
- OpenAI. “Reasoning Models.” OpenAI API Documentation, accessed September 2026. https://developers.openai.com/api/docs/guides/reasoning
- Anthropic. “Context Windows.” Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/build-with-claude/context-windows
- Google. “Gemini Developer API Pricing.” Google AI for Developers, accessed September 2026. https://ai.google.dev/gemini-api/docs/pricing
- OpenAI. “Pricing.” OpenAI API Documentation, accessed September 2026. https://developers.openai.com/api/docs/pricing
- Anthropic. “How We Built Our Multi-Agent Research System.” Anthropic Engineering, June 13, 2025. https://www.anthropic.com/engineering/multi-agent-research-system
- Anthropic. “PDF Support.” Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/build-with-claude/pdf-support
- Doshi P, Jefferson T. “Clinical Study Reports of Randomised Controlled Trials: An Exploratory Review of Previously Confidential Industry Reports.” BMJ Open, 2013;3(2):e002496. https://pmc.ncbi.nlm.nih.gov/articles/PMC3586134/
- Microsoft. “Standard Harness Licensing: Microsoft Copilot Studio.” Microsoft Learn, updated August 3, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/billing-licensing
- Microsoft. “Provisioned Throughput for Foundry Models.” Microsoft Learn, updated July 15, 2026. https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput
- Anthropic. “Workspaces.” Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/manage-claude/workspaces
- OpenAI. “Spend Limits.” OpenAI API Documentation, accessed September 2026. https://developers.openai.com/api/docs/guides/spend-limits
- Amazon Web Services. “Set Up a Model Invocation Resource Using Inference Profiles.” Amazon Bedrock User Guide, accessed September 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/inference-profiles.html
- Anthropic. “Usage and Cost API.” Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/manage-claude/usage-cost-api
- OpenAI. “Prompt Caching.” OpenAI API Documentation, accessed September 2026. https://developers.openai.com/api/docs/guides/prompt-caching
- Anthropic. “Batch Processing.” Claude Platform Docs, accessed September 2026. https://platform.claude.com/docs/en/build-with-claude/batch-processing








Your perspective matters—join the conversation.