How Token Pricing Works and Why Early Spend Is Hard to Predict

Every major model provider prices its application programming interface (API) access the same basic way: a price per million input tokens (what you send to the model) and a higher price per million output tokens (what the model writes back). A token is a piece of a word. Anthropic’s pricing documentation gives a rough rule of thumb that one token is about four characters, or about 0.75 words, of English text, and notes that the exact count varies by language and content type1.

That sounds simple, and at the level of a single request it is. The difficulty is that almost nobody in a business thinks in tokens. A finance lead thinks in dollars per month. A regulatory affairs lead thinks in documents. An IT lead thinks in users and licenses. The FinOps Foundation, the industry body for cloud financial management, makes this point directly in its June 2026 paper on token economics: finance teams and many senior engineering leaders have no intuition for what a token costs or how many tokens a typical request consumes2.

Input, Output, and Reasoning Tokens

Three categories of tokens show up on the bill, and they are not priced the same way.

  • Input tokens are everything sent to the model: the system instructions, the user’s question, any attached documents, the conversation so far, and the definitions of any tools the model is allowed to call.
  • Output tokens are what the model writes. They are priced higher than input tokens, typically five to six times higher on the price lists we reviewed.
  • Reasoning or thinking tokens are the intermediate work some models do before they answer. OpenAI’s documentation states that while reasoning tokens are not visible through the API, they “still occupy space in the model’s context window and are billed as output tokens.”3 Anthropic’s documentation says the same of its thinking tokens4, and Google’s price list labels its output price as “including thinking tokens”5.

Reasoning tokens are a common source of surprise. A short visible answer can carry a large hidden bill. OpenAI notes that its reasoning models may generate anywhere from a few hundred to tens of thousands of reasoning tokens depending on the problem, and recommends reserving at least 25,000 tokens for reasoning and outputs when teams start experimenting3.

What the Price Lists Say Today

Prices change often, sometimes several times a year, so every figure below is dated. The table shows a sample of mid-tier and small models from three vendors, as listed on each vendor’s own pricing page in September 2026. It is not a recommendation of any model. It is here so readers can see the shape of the numbers.

Model (as listed in September 2026)Input, per million tokensOutput, per million tokensNotes
Anthropic Claude Sonnet 5.51$2.00$10.00Cache reads $0.20; batch rates are half the standard rates
Anthropic Claude Haiku 4.51$1.00$5.00Smaller, faster model
OpenAI gpt-6-sol6$2.00$10.00Short-context rate; the long-context rate is listed at $4.00 input and $15.00 output
OpenAI gpt-5.4-mini6$0.75$4.50Smaller model
Google Gemini 3.1 Pro Preview5$2.00 (prompts up to 200K tokens); $4.00 above$12.00; $18.00 above 200KOutput price includes thinking tokens
Google Gemini 3.8 Flash5$0.75 through Dec 31, 2026; $1.50 from Jan 1, 2027$3.75 through Dec 31, 2026; $7.50 from Jan 1, 2027Listed price doubles at the start of 2027

Two things stand out. First, the headline rates for comparable tiers are close to each other. Choosing a vendor on the price per million tokens alone rarely changes the budget much. Second, the rules around the headline rate differ a lot: long-context surcharges, scheduled price changes, cache pricing, and batch discounts. We return to those in a later section, because they matter more than the headline.

Why Month One Tells You Little About Month Six

Leaders often ask for a forecast after a pilot’s first month. That forecast is usually wrong, for reasons that have nothing to do with poor planning.

  • Usage patterns are still forming. Early users try short questions. Once they trust the tool, they attach whole documents and run longer sessions.
  • The application is still changing. Engineers add retrieval, tools, and agent steps, and each one adds tokens to every request.
  • Model versions change. Anthropic’s pricing page notes that its Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text1. A model upgrade at the same listed price can still raise the bill.
  • Prices change on a schedule. Google’s listed rates for Gemini 3.8 Flash double on January 1, 20275. A 2027 budget built from September 2026 rates would be half the real figure.
  • Mistakes are expensive and fast. The FinOps Foundation paper describes how an agentic workflow that loops unexpectedly, a system prompt that was doubled in length by accident, or a feature that drives far more usage than expected can produce invoices that exceed monthly budgets in a single day2.
~4×More tokens used by single agents than chat interactions, in Anthropic’s own data7
~15×More tokens used by multi-agent systems than chats, in the same Anthropic data7
~30%More tokens for the same text with Anthropic’s newer tokenizer (Claude 4.7 and later)1

What Drives Token Consumption

When spend runs ahead of plan, the cause is almost always a change in how many tokens each task uses, not a change in price. Six drivers account for most of it. Each one is easy to see once you know to look for it.

Driver 1

Context Length

Everything in the request counts: instructions, attached files, retrieved passages, and prior turns. A long system prompt is paid for on every single call.

Driver 2

Conversation Accumulation

Each new turn resends the full history. The twentieth question in a session is billed for the nineteen that came before it.

Driver 3

Long Document Processing

A PDF page is typically 1,500 to 3,000 tokens of text before any image processing. Documents with hundreds of pages dominate the bill.

Driver 4

Retries and Regeneration

Failed format checks, truncated outputs, and users pressing “regenerate” all resend the full input. The retry pays the whole price again.

Driver 5

Agentic Loops

An agent calls the model many times per task: to plan, to call a tool, to read the result, to check its work. Each call carries the growing context.

Driver 6

Hidden Overhead

Tool definitions, tool-use system prompts, reasoning tokens, and paid add-ons such as web search appear in the bill but not on the screen.

Context Length and Conversation Accumulation

Anthropic’s documentation on context windows describes the mechanism plainly. As a conversation moves through turns, each user message and each model response accumulates in the context window, and previous turns are preserved completely. Each turn’s input contains all previous conversation history plus the current message, and the response becomes part of the input for the next turn4. The same page lists what counts toward the window: the system prompt, every message including tool results, images, and documents, and the tool definitions4.

The effect on spend is that the price of a conversation grows faster than the number of turns. A ten-turn session where each turn adds 2,000 tokens does not cost ten times a single turn. The last turn alone carries about 20,000 tokens of history, and the total input across the session is the sum of all ten growing requests. Long sessions are where chat tools become expensive.

Long Document Processing

Documents are the largest single driver in regulated industries, and we give them their own section below. The key fact is this one. Anthropic’s PDF documentation states that each page typically uses 1,500 to 3,000 tokens depending on content density, and that because each page is also converted into an image, image-based token charges apply on top8. So the text alone of a 100-page document is roughly 150,000 to 300,000 tokens, and the page images add more.

Retries and Regeneration

A retry is not a small follow-up. It is a full new request. If an application asks a model to extract data from a 200-page document into a fixed format, and the output fails a format check, the retry sends the 200 pages again. OpenAI’s documentation describes a related case for reasoning models: if the model reaches its output limit before producing visible text, the response comes back incomplete, and “you could incur costs for input and reasoning tokens without receiving a visible response.”3 Paying for a failed attempt and then paying again for the retry is one of the most common patterns we see in early usage data.

In chat tools the same thing happens through people. A user who does not like an answer presses “regenerate,” or pastes the same long document into a new chat to start fresh. Neither shows up as an error. Both show up on the invoice.

Agentic Loops

An agent is an application where the model decides what to do next, calls tools, reads the results, and repeats until the task is done. That loop is the point of an agent, and it is also what drives its token use. Anthropic’s engineering team, writing about its own multi-agent research system, reported that in its data agents typically use about four times more tokens than chat interactions, and multi-agent systems use about 15 times more tokens than chats7. The same post notes that token usage by itself explained 80% of the variance in performance on the browsing evaluation they studied7. In other words, agents often perform better partly because they spend more.

The post draws the practical conclusion: multi-agent systems make economic sense for tasks where the value of the task is high enough to pay for the increased performance7. That is the right test for any agent in a regulated company. The FinOps Foundation paper adds the risk side: an agent that calls itself, or tool chains that recurse without end, can consume millions of tokens in minutes2.

Hidden Overhead

Some tokens never appear on screen. Anthropic’s pricing page lists a tool-use system prompt that the API adds whenever tools are provided, in the range of roughly 290 to 800 tokens for recent models depending on the model and settings, plus the tokens for the tool names, descriptions, and schemas themselves1. Browser and computer-use toolsets add several thousand input tokens per request. Paid add-ons carry their own meters: Anthropic lists web search at $10 per 1,000 searches plus the tokens for the results1, and Google lists grounding with Google Search for Gemini 3.x models at $14 per 1,000 requests after 5,000 free requests per month5. None of these are large on a single call. Across an agent that makes dozens of calls per task and thousands of tasks per month, they add up.

The Document-Heavy Line Item That Surprises Pharma and Biotech

Most public discussion of AI pricing is built around chat and customer support, where a request is a few hundred to a few thousand tokens. Life sciences work is different. The documents that matter most are long, and many of the highest-value use cases involve reading them in full: summarizing a clinical study report, checking a submission section against its source data, drafting a response to a health authority question, reviewing a deviation investigation against the relevant procedures, or screening literature for safety signals.

How Long Regulatory Documents Are

Clinical study reports (CSRs) are the clearest example. A CSR is the full report of a clinical trial prepared for regulators, including the protocol, the statistical analysis plan, tables, and patient listings. A 2013 review in BMJ Open by Peter Doshi and Tom Jefferson assembled 78 CSRs covering 90 randomized controlled trials of 14 medicines, dated 1991 to 2011, totaling 144,610 pages9. By our arithmetic that is an average of about 1,850 pages per report. The review reported median lengths for individual sections, including 337 pages of attached tables, 62 pages of trial protocol, and 447 pages of individual efficacy listings9. The authors describe their sample as non-random and say they cannot tell whether it is representative9, so treat the average as an illustration of scale, not a benchmark.

Apply the per-page figure from Anthropic’s PDF documentation (1,500 to 3,000 text tokens per page8) to an 1,850-page report and the text alone comes to roughly 2.8 million to 5.6 million tokens. That is larger than the 1-million-token context window of Anthropic’s current models4, and Anthropic’s PDF support caps a single request at 600 pages8. A full CSR cannot be read in one call. It has to be split, searched, or summarized in stages, and every one of those design choices changes the token count.

An Illustrative Comparison

The table below uses Claude Sonnet 5.5 rates as listed in September 2026 ($2.00 per million input tokens, $10.00 per million output tokens, $2.50 per million for a five-minute cache write, and $0.20 per million for a cache read)1. The token counts are our own assumptions for illustration, not measurements from any company. The point is the ratio between rows, which holds at any vendor with similar pricing.

Task (illustrative)Assumed tokensPrice per task
Short chat question and answer1,500 input, 500 outputAbout $0.008
One pass over a 300-page document, 4,000-token summary450,000 to 900,000 input (text only), 4,000 outputAbout $0.94 to $1.84
Same document, sent through the Batch API at half priceSame as aboveAbout $0.47 to $0.92
Agent that reads the document across 12 model calls, no caching12 × 450,000 input, 12 × 2,000 outputAbout $11.04
Same agent with prompt caching (one cache write, 11 cache reads)Same volume, cachedAbout $2.36

By this arithmetic, one pass over a 300-page document costs roughly 120 to 230 times the short chat exchange. An agent that re-reads the same document on every step, without caching, costs more than ten times the single pass, even before counting the growth of its own history. With caching, the same agent comes back down to between two and three single passes. The agent rows use the low end of the page estimate and ignore history growth, so real figures would be higher.

Now multiply by volume. At 2,000 documents a month, the single-pass design comes to roughly $1,880 to $3,680 a month on these assumptions. The uncached agent design comes to about $22,000. The cached agent comes to about $4,700. Same model, same documents, same listed price, and a spread of more than ten times depending on how the application is built.

Why this line item surprises people. Budgets for enterprise AI are often set from a pilot where people asked short questions. The first document-heavy workflow to go live, such as CSR summarization, submission quality checks, or literature screening, can use more tokens in a week than the whole pilot used in a quarter. Nothing is broken. The work simply involves reading far more text. Ask for a token estimate per document before any document-heavy use case goes live, and budget it as its own line.

Where Document-Heavy Use Cases Show Up

In our experience, these are the pharma and biotech workflows where document volume drives the bill:

  • Regulatory writing and review: drafting or checking sections of a submission against source reports, where the source material runs to hundreds or thousands of pages.
  • Clinical documentation: CSR summaries, protocol amendment impact reviews, and consistency checks across the protocol, statistical analysis plan, and report.
  • Quality and manufacturing: deviation and CAPA (corrective and preventive action) investigations that pull in batch records, procedures, and past investigations.
  • Pharmacovigilance and medical information: literature screening and case narrative work at high volume.
  • Due diligence: reviewing data rooms for licensing or acquisition, where hundreds of documents are read once under time pressure.

None of these is a reason to avoid AI. Several are among the strongest use cases in the industry. They are a reason to design for document volume from the start, which is where most of the savings are. For more on the writing use cases specifically, see our earlier article on generative AI for regulatory writing.

Pricing Structures Differ More Than the Headline Rate

When procurement compares AI vendors, the comparison usually starts with the price per million tokens. For long-document work, the rules around that price often matter more. Five differences deserve attention.

Long-Context Surcharges

Some vendors charge a higher rate once a request passes a size threshold. Google lists Gemini 3.1 Pro Preview at $2.00 per million input tokens for prompts up to 200,000 tokens and $4.00 above that, with output rising from $12.00 to $18.005. OpenAI’s price list shows separate short-context and long-context columns, with gpt-6-sol at $2.00 input and $10.00 output for short context and $4.00 input and $15.00 output for long context6. Anthropic, by contrast, states that its Claude 4.6 and later models include the full 1-million-token context window at standard pricing, so a 900,000-token request is billed at the same per-token rate as a 9,000-token request1. For a use case built around long regulatory documents, these rules can matter more than a small difference in the base rate.

Tokenizers

A token is not a fixed unit across vendors or even across model versions. Each model family splits text in its own way. As noted above, Anthropic states that its newer tokenizer produces approximately 30% more tokens for the same text1. Two models with the same listed rate can therefore charge different amounts for the same document. The only reliable comparison is to run a sample of your own documents through each candidate and read the token counts from the responses.

Scheduled and Promotional Prices

Price lists now carry dates. Google lists Gemini 3.8 Flash at one rate through December 31, 2026 and double that rate from January 1, 20275. OpenAI’s pricing page notes promotional pricing for one of its models “available at least through November 21, 2026.”6 Anthropic’s page records that introductory pricing for Claude Sonnet 5 became the standard price and a scheduled increase will not occur1. Prices move in both directions. Any budget that crosses a calendar year should check each vendor’s page for dated changes, and any contract should say which price applies.

Data Residency and Regional Premiums

Many life sciences companies need data to stay in a given country or region. That often carries a premium. Anthropic lists a 1.1 times multiplier on all token categories for US-only inference on Claude 4.6 and later models, and notes that regional endpoints on Amazon Bedrock and Google Cloud carry a 10% premium over global endpoints1. OpenAI’s pricing page states that FedRAMP endpoints and eligible data residency endpoints are charged a 10% uplift6. If your security or privacy review requires regional processing, add the premium to the budget from day one.

Seats, Credits, and Reserved Capacity

Not every enterprise AI tool is billed per token. Three other models are common, and many companies use all three at once.

Billing modelHow it worksWhere the surprise comes from
Per-token APIPay for input and output tokens usedConsumption per task grows with documents, agents, and retries
Per-seat license with included usageFixed monthly fee per user; some features are includedAgents and add-ons built on top may draw from a separate consumption meter
Prepaid or pay-as-you-go creditsCredits consumed per action, varying by task complexityUnused credits may expire; exceeding capacity can lead to enforcement
Reserved capacityPay per hour for dedicated throughput, whether used or notIdle capacity is still billed; too little capacity leads to throttled requests

Microsoft’s Copilot Studio documentation, last updated in August 2026, is a good example of the credit model. It describes Copilot Credits as the common currency for agents, says the number of credits counted for each response depends on the complexity of the task, states that unused credits do not carry over to the next month, and warns that if usage exceeds purchased capacity, technical enforcement applies and can result in service denial10. The same page notes that for users with a Microsoft 365 Copilot license, certain agent answers inside Microsoft 365 apps are zero-rated and do not draw on the credit pool10. The practical lesson is that a seat license does not mean a fixed bill once agents are involved.

Reserved capacity works the other way. Microsoft’s documentation on provisioned throughput in Microsoft Foundry explains that provisioned deployments are billed at an hourly rate per provisioned throughput unit (PTU), “regardless of the number of tokens consumed,” and that 1-month or 1-year reservations give a discounted rate11. The same page recommends provisioned throughput for predictable traffic and standard pay-per-token deployments for development, testing, low volume, or highly variable traffic11. That guidance fits the budgeting approach below: pay per token while you learn the pattern, and consider reserved capacity only once the pattern is stable.

How to Budget: Start From the Task, Not the Token

The most useful change a leadership team can make is to stop budgeting AI in tokens and start budgeting it in tasks. “We expect to process 2,000 deviation investigations a year at about $X each” is a number a CFO, a quality head, and an engineer can all discuss. “We expect 400 million tokens” is not. The FinOps Foundation paper makes the same recommendation in its own terms: unit cost metrics such as cost per query, per user, or per outcome make AI spend legible to business stakeholders in a way raw token counts cannot2.

We suggest a six-step approach for each use case.

1

Define the Task Unit

Name the unit of work the business cares about: one CSR summary, one submission section check, one investigation draft, one literature screen. Estimate the monthly volume with the business owner, not the engineering team.

2

Measure Tokens Per Task on Real Documents

Run a representative sample of real documents through the actual application and read the input, output, cached, and reasoning token counts from each response. Vendors provide token counting tools for estimates before sending, but measured usage from real runs is better. Include the short, typical, and longest documents, because the long ones drive the bill.

3

Price It With Dated Rates

Multiply by the vendor’s listed rates and write down the date you read them, for example “Sonnet 5.5 rates as listed in September 2026.” Include any long-context, residency, cache, or batch adjustments. Check each vendor’s page for scheduled changes inside the budget period.

4

Add Allowances for Retries, Testing, and Growth

Retries, regeneration, and development runs are real spend. So is validation testing in GxP use cases, and so is re-testing after a model version change. Measure the retry rate during the pilot rather than assuming it is zero, and plan for usage to grow as people trust the tool.

5

Express the Budget as a Range

Give leadership a low, expected, and high figure per month, with the assumptions behind each. A range is more honest than a single number in the first six months, and it tells finance how much variance to hold in reserve.

6

Re-Forecast Monthly Until the Pattern Is Stable

Compare actual cost per task with the estimate each month. When three months in a row fall inside the range, the use case is ready for a normal annual budget line, and possibly for reserved capacity if volume is high and steady.

A Worked Budget Line

Suppose a regulatory team wants AI support for quality checks on submission sections. The business owner expects about 150 sections a month. Suppose measurement on twenty real sections shows an average of 250,000 input tokens and 6,000 output tokens per check, a retry rate of about one in ten, and no need for real-time answers. At Sonnet 5.5 batch rates as listed in September 2026 ($1.00 input and $5.00 output per million tokens1), one check is about $0.25 for input and $0.03 for output, or about $0.28. With retries it is about $0.31. At 150 checks, the monthly figure is under $50. That is small. The same check run as a 15-step agent without caching, at standard rates, would be many times larger. The budget conversation should be about design, not vendor.

Budget per task, report per team, review per quarter. The unit that finance approves is the task. The unit that operations watches is the team or use case. The unit that leadership reviews is the quarter, with the price list dates noted. Keeping those three levels separate stops the most common argument in AI budgeting: a monthly invoice that nobody can explain.

How to Monitor and Cap Spend Across Teams

Budgets only work if spend can be seen and, when needed, stopped. The good news is that the major vendors now provide the basic controls. The less good news is that they do not go as far as a finance team expects, so some of the work falls to you.

Separate Spend by Team and Use Case

The first control is structural: give each team or use case its own container for API keys, usage, and limits. On Anthropic’s platform that container is a workspace. The documentation describes using workspaces to separate projects, environments, or teams while keeping centralized billing, and lists development, staging, and production as a common split12. OpenAI uses projects for the same purpose13. On Amazon Bedrock, application inference profiles let you attach tags to a model endpoint and track costs through AWS cost allocation tags14.

This separation is what makes chargeback or showback possible. The FinOps Foundation paper notes that an invoice from a model provider typically shows aggregate token consumption broken down at most by API key or project, and that there is “no native concept of business unit, cost center, application, or workload.”2 If you want to know what the pharmacovigilance team spent, you have to have given the pharmacovigilance team its own workspace, project, or tagged profile before the spend happened.

Know the Difference Between an Alert and a Cap

Vendors offer two kinds of spend control, and they behave very differently.

  • Alerts notify someone when spend crosses a threshold. Traffic keeps flowing.
  • Hard limits stop traffic. OpenAI’s documentation says that when tracked spend reaches a hard limit, affected API requests return a 429 error, and warns that “Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount.”13

Anthropic’s workspaces support both monthly spend limits and alerts at chosen thresholds, as well as rate limits on requests, input tokens, and output tokens per minute12. Workspace limits can be set lower than the organization’s limits but not higher, and cannot be set on the default workspace12. That last detail matters: if teams are working in the default workspace, you cannot cap them individually.

The FinOps Foundation paper points out a gap between these controls and how finance thinks: native rate limits are expressed in tokens per minute, not dollars per month2. A rate limit protects against a runaway loop. It does not keep a team inside its quarterly budget. You need both.

ControlWhat it protects againstWhere to set it
Monthly spend alert at 50%, 80%, 100% of budgetSlow overspend that nobody notices until the invoiceVendor console per workspace or project
Hard monthly spend limitLarge overspend in non-critical work (development, experiments)Vendor console; use with care in production
Tokens-per-minute rate limitRunaway loops and sudden spikesVendor console per workspace or project
Maximum output tokens per requestUnexpectedly long or looping outputsApplication code
Maximum steps per agent taskAgents that never finishApplication or agent framework
Per-task cost check in logsDesign changes that raise cost per taskYour own monitoring from usage data

Watch Usage Daily, Not Monthly

Monthly invoices arrive too late to act on. Vendors now publish usage data that can be pulled much sooner. Anthropic’s Usage and Cost API returns token counts grouped by model, workspace, API key, and other dimensions in one-minute, one-hour, or one-day buckets, and cost in US dollars by day; the documentation says data typically appears within five minutes of a request completing15. It lists budget monitoring and cost attribution by workspace for chargebacks among the common uses15.

Pull that data into whatever your finance and IT teams already use, and look at three numbers per use case every week: total spend against budget, cost per task against estimate, and the cache hit rate. The FinOps Foundation paper adds a useful distinction for alerting: alerts at the account level tell you spend happened, while alerts at the workload level tell you which application or team caused it, and both layers are needed2. For runaway loops, it recommends monitoring tokens per minute at the application level, not just daily spend2.

Assign an Owner

Controls without an owner do not get used. Name one person per use case who receives the alerts, explains variances, and approves changes that affect cost per task. In most companies that should be the business owner of the use case, with support from IT. It should not be the vendor account manager.

Levers That Lower Spend Without Lowering Quality

Once you can see cost per task, you can lower it. The levers below are all documented by the vendors themselves, and most do not reduce output quality when applied with care.

Prompt Caching

Caching stores the processed form of a repeated prompt prefix, such as a long system prompt or a document that several requests will read. On Anthropic’s platform, a cache read costs 10% of the standard input price for most models, and a five-minute cache write costs 1.25 times the input price, so caching pays off after one cache read1. OpenAI enables caching by default for supported models, describes the cached-input rate as “discounted up to 90%,” and advises placing stable instructions and shared reference material first16. For document-heavy agents, caching is often the single biggest saving, as the illustrative table above shows.

Caching has conditions. Caches expire (five minutes or one hour on Anthropic’s platform1), and the cached part must be identical from one request to the next. An application that inserts a timestamp at the top of every prompt will never get a cache hit. This is a design question for the engineering team, and it is worth asking about in every architecture review.

Batch Processing

Much document work does not need an answer in seconds. Overnight processing is fine for literature screening, back-file summarization, or bulk quality checks. Anthropic’s Message Batches API charges 50% of standard prices, with most batches finishing in less than an hour and results available within 24 hours17. Google and OpenAI list the same 50% batch discount56. Anthropic notes that batch and caching discounts can be combined1.

Model Routing

Not every step needs the most capable model. Classifying a document, extracting a date, or checking a format can often run on a smaller model at a fraction of the price, while the drafting or reasoning step uses a larger one. Anthropic’s own pricing guidance suggests using smaller models for simple tasks, mid-tier models for most production work, and the largest models for the most complex reasoning1. Test quality at each step before routing, and record the routing decision, because in a GxP context it becomes part of the system’s design.

Send Less Text

The cheapest token is the one you never send. Four practices help:

  • Retrieve, do not attach. For questions about a large document set, a retrieval step that finds the relevant passages is usually far cheaper than attaching whole documents.
  • Extract text where images add nothing. Because PDF pages are also processed as images8, sending plain text for text-only pages avoids paying for page images the model does not need.
  • Keep conversation history short. Anthropic documents server-side compaction and context editing, which summarize or clear earlier parts of a long conversation or agent session4. Starting a fresh session for a new topic does the same thing by hand.
  • Trim tool lists. Only give an agent the tools it needs for the task. Every tool definition is paid for on every call.

Set Limits in the Application

Set a maximum output length for each request type, a maximum number of steps for each agent task, and a maximum retry count for each failure type. These limits protect the budget and also make the system more predictable, which matters for validation.

Where to start. For most document-heavy use cases, the order of impact is: design (retrieve rather than attach, cap agent steps), then caching, then batch for anything not time-sensitive, then model routing. Price negotiation with the vendor comes last. Anthropic, for example, notes that volume discounts are negotiated case by case1, and a well-designed application gives you a better starting position for that conversation.

For companies weighing whether some workloads should move off per-token APIs entirely, our article on small language models on-prem covers when that makes sense. It trades a variable token bill for fixed infrastructure and operating effort, which is a different kind of commitment, not a free one.

Token Spend in GxP Settings: Validation, Change Control, and Hard Caps

Most writing on AI spend ignores regulated work. In pharma and biotech, three issues connect token spend to quality and compliance.

Validation and Re-Testing Consume Tokens

Where an AI tool supports a GxP process, it is validated for its intended use, typically following a risk-based approach such as GAMP 5 (the ISPE guide for computerized system validation). That validation involves running test cases, often many times over, and keeping the evidence. Every test run is a model call and is billed like any other. When the model version changes, the tests run again. None of this is large compared with production volume for most use cases, but it belongs in the budget, and it should not come out of a team’s production allowance where a hard cap could block it.

Model and Tokenizer Changes Belong in Change Control

A model version change is already a change control event in a validated system, because behavior can change. It can also change cost. Anthropic’s note that its newer tokenizer produces approximately 30% more tokens for the same text1 is a clear example: the same prompt and the same document, on a newer model at the same listed price, can produce a larger bill. We recommend adding a cost-per-task comparison to the impact assessment for any model change, alongside the quality and behavior comparison. It takes very little extra effort when the test set already exists.

A Hard Cap Can Stop a Validated Process

Hard spend limits are a good control for experiments and development. In production GxP workflows they need more thought. If a hard limit is reached, the vendor returns an error and requests stop13. In a credit-based product, exceeding capacity can lead to service denial10. If an AI step is part of a validated process, a stop in the middle of a run is a process interruption that may need to be handled as a deviation, with an impact assessment on any affected records.

The safer pattern for GxP production work is to use alerts with enough headroom for someone to act, a documented procedure for what happens when an AI step is unavailable (including a manual fallback), and hard caps reserved for non-GxP workspaces. The business continuity plan for the process should cover this case as it would any other system outage.

A short checklist for GxP AI budgets:

  • Validation and re-test runs have their own budget line, separate from production.
  • Production GxP workspaces use alerts, not hard caps, with a named owner and a documented fallback.
  • Model change impact assessments include a cost-per-task comparison.
  • Price list dates are recorded with each budget and each contract.
  • Agent step limits and output limits are part of the validated configuration, not left to defaults.

Two related governance points are covered in more depth elsewhere on our site. The agentic designs that drive the largest token counts are discussed in our article on agentic AI in pharma. The commercial side of committing to one vendor’s pricing, tokenizer, and roadmap is covered in our piece on AI vendor lock-in.

Conclusion

Enterprise AI token cost is predictable once you stop looking at it as a price and start looking at it as a volume. The listed rates for comparable models are close together, and they change often enough that any figure needs a date. What separates a manageable bill from a surprising one is the number of tokens each task uses, and that is set by design choices: how much text is sent, how often it is resent, how many steps an agent takes, and whether caching and batch processing are used. In pharma and biotech, the long documents at the center of regulatory, clinical, and quality work make those choices matter more than in most industries. A single design decision can move the monthly figure by ten times or more.

The companies that manage this well budget per task, separate spend by team from the first day, watch cost per task weekly during the first months, and treat model changes as both a quality event and a cost event. None of that requires new tools. It requires the same discipline life sciences already applies to other systems. Sakara Digital works with pharma and biotech organizations planning and governing enterprise AI, including the budgeting and controls that support it. If you are sizing a document-heavy use case or trying to explain an AI invoice that grew faster than expected, and want an independent perspective on where to start, we are happy to have that conversation.

For Further Reading