Why the Build Question Is Back

For most of the last twenty years, pharma and biotech companies bought their quality, document, training, and laboratory systems. Building one took a development team, many months, and a validation project to match. Few companies outside large pharma had the people for that, and most of those that did still chose to buy.

AI coding tools changed the first part of that. A quality specialist can now describe a deviation tracker in plain English and have a working version by the end of the week. A clinical operations lead can build an agent that drafts site correspondence. The results often look good, they solve a real problem, and the first version is cheap compared with a vendor subscription.

The numbers show how fast this is moving. In McKinsey’s 2026 State of AI survey, nearly a third of respondents (32 percent) said their organizations had decided against buying one or more software products or features because they could be built in-house with agentic coding tools. Among the small group McKinsey calls AI high performers, nearly half said the same, compared with 31 percent of other respondents.1

The people doing the building have changed too. Opsin Labs, an AI security company, analyzed enterprise production environments from March 2025 to June 2026. It reported that 67 percent of AI agents were built by employees without an engineering background, that companies now average about one agent per employee, live or in draft, and that 60 percent of agents provisioned beyond default settings were given allow-all access rather than access limited to what their tasks required.2

32% of organizations decided against buying software they could build with AI coding tools
(McKinsey, 2026)1
67% of enterprise AI agents were built by employees without an engineering background
(Opsin Labs, 2026)2
60% of agents given more than default permissions received allow-all access
(Opsin Labs, 2026)2

None of this is a problem in itself. Much of what people build with AI is personal productivity: a script that cleans up a spreadsheet, a prompt that drafts meeting notes, an agent that sorts an inbox. The problem starts when a tool built in a week begins to hold GxP records, support a GxP decision, or replace a validated system. At that point the company is no longer comparing a subscription fee with a few weeks of effort. It is deciding whether to become its own software supplier, with every obligation that role carries.

The rest of this article lists those obligations in the terms a budget owner uses: staff hours, contractor fees, and recurring spend.

Scope of this article. Everything below applies to software that supports GxP activities: systems that create, change, or hold records regulators can inspect, or that support decisions about product quality, patient safety, or data integrity. Personal productivity tools and non-GxP business tools have a much shorter list of hidden costs. Our article on citizen developers in regulated industries sets out a four-tier model for deciding which category a tool belongs in.

What You Were Buying Without Noticing

A software subscription looks like a fee for code. A commercial GxP product includes much more than that, but the buyer rarely lists it out, because it comes with the license and nobody has to plan for it.

Veeva’s annual report gives a sense of the scale. In its fiscal year ending January 31, 2026, Veeva spent $767 million on research and development.3 Veeva describes its industry focus as giving it an in-depth view of “the needs and best practices of life sciences companies,” which it says lets it “quickly adapt to regulatory changes.”3

Most vendors in this market make some version of that claim, and buyers should test it during selection. The point here is simpler. When you buy a commercial product, you share the cost of that knowledge with every other customer. When you build a custom tool, you pay for all of it yourself.

Design

Industry Practices Built In

Workflows, required fields, signature rules, and status logic shaped by many customers and many inspections.

Roadmap

New Features Every Year

Improvements you did not request and often did not know you needed, funded from the vendor’s development budget.

Evidence

Testing You Can Rely On

Supplier test records and documentation that FDA’s CSA guidance and GAMP 5 allow you to use to reduce your own testing.

Compliance

Regulatory Updates

Product changes for new or revised regulations, made once and delivered to every customer.

Operations

Hosting and Security

Servers or cloud services, backups, disaster recovery, monitoring, and security patching.

Support

Support at Every Level

A help desk, product specialists, and developers who fix defects in the code itself.

People

Training and Help Materials

User guides, release notes, help pages, and training courses that are kept current with each release.

Inspection

A Supplier to Point To

A vendor with its own quality system, audit reports, and a customer base whose use of the product has already been inspected.

When you build, each of these becomes your responsibility. For a low-risk tool, some can be scaled down. For a GxP system, very few can be skipped.

The Full List: 20 Hidden Costs of Building GxP Software

Each item below is a budget line. Most are staff hours from quality, IT, the business, or contractors. Some are recurring spend on hosting, security tools, or AI usage. The items are grouped by when they appear in the life of the system: before the first line of code, building and proving it works, running it in production, keeping it current, people and knowledge, and the end of the system’s life. Only one of them, the build itself, is smaller because of AI.

Before the First Line of Code

1. Deciding how the process should work. When you buy a quality or document system, the vendor has already decided how a deviation moves from open to closed, which fields must be filled in before a record can be approved, who can sign what, and what counts as complete. You can change many of those choices through configuration, but you start from a design that already works for other companies. When you build, someone has to make every one of those decisions and write them down. In GxP, that written version is the user requirements specification, and it has to be reviewed, approved, and kept current for the life of the system. AI can help draft it, but it cannot make the decisions, and the people who own the process have to spend time making them. A common pattern is to let the AI tool make design choices during the build, then discover that nobody can explain why the system behaves the way it does. In a regulated company, that usually means writing the requirements again from the beginning.

2. Finding the industry practices a vendor would have built in. This is the item most build estimates leave out. A mature GxP product reflects years of lessons about what inspectors expect and where processes fail. Examples include linking a deviation to its CAPA and to the CAPA effectiveness check, requiring a reason for every change to a controlled record, preventing the author of a document from also approving it, and flagging training that will expire before a new procedure takes effect. A team building its own tool has to find these practices, decide which ones apply, and design them in. That means research time from people who know the regulations, review of guidance documents and inspection findings, and often a consultant. When this step is skipped, the gaps tend to appear in an audit rather than in testing.

3. Risk assessment and validation planning. Under ISPE’s GAMP 5, software written for a single company is Category 5, custom software, which calls for the most life-cycle activity: design specifications, code review, and testing at a depth a configured commercial product does not need.4 FDA’s Computer Software Assurance (CSA) guidance supports a risk-based approach that concentrates testing where the risk is.5 It was written for medical device production and quality system software, but pharma and biotech quality teams widely apply its principles. It does not remove the need to assess risk, choose the testing approach, and write the plan. For a commercial product, much of that planning starts from the vendor’s documentation. For a custom tool, it starts from nothing.

Building It and Proving It Works

4. Building it, and reviewing what the AI wrote. This is the line AI has made smaller, and the savings are meaningful. But AI-written code needs more review than its speed suggests. Veracode tested code from more than 100 large language models and found that 45 percent of the code samples failed security tests and introduced vulnerabilities from the OWASP Top 10 list. Newer and larger models wrote more working code, but no more secure code.6 GitClear’s analysis of 211 million changed lines of code found that copied and pasted code rose from 8.3 percent of changed lines in 2021 to 12.3 percent in 2024, while code that was reorganized to improve its structure fell from 25 percent to under 10 percent. Both trends make code harder to maintain.7 Developers see the same problem. In Stack Overflow’s 2025 Developer Survey of more than 49,000 developers, 46 percent said they do not trust the accuracy of AI tools’ output, and 45 percent named the time spent debugging AI-generated code as a main frustration.8 Someone qualified has to read, test, and fix what the AI produces. That time belongs in the build estimate.

5. Testing and documentation without vendor evidence. When you buy, a large share of the testing evidence already exists. FDA’s CSA guidance says manufacturers can rely on validation activities performed by others, including developers, suppliers, and cloud service providers. It also says that for some supporting software, the vendor’s evaluation and validation records, or records of the software’s installation and configuration, may be enough on their own, so that additional scripted or unscripted testing is unnecessary.5 GAMP 5 makes the same point about using supplier documentation and testing where the supplier has been assessed.4 When you build, there is no supplier evidence. Your team writes and runs every test, from unit tests through user acceptance, builds the trace from each requirement to its tests, and writes the design and configuration documents a supplier would normally hold. In the first year this is often the largest single line, and part of it repeats with every significant change.

6. Building Part 11 controls and proving they work. Commercial GxP systems come with audit trails, electronic signatures, access controls, and record export already built and tested. A custom tool has to build all of them. Among other things, 21 CFR Part 11 requires secure, computer-generated, time-stamped audit trails that record when operators create, change, or delete records, without hiding the earlier entries. It requires limiting system access to authorized people, authority checks on who can sign or change records, and the ability to produce accurate and complete copies of records for inspection.9 Each control needs a design, code, testing, and evidence. These are also among the first controls an inspector asks to see working.

The audit trail is not a log file. Many tools built quickly record activity in a technical log that a developer can edit, delete, or overwrite. That does not meet Part 11. A compliant audit trail has to be generated by the system, protected from change, time-stamped, kept for as long as the records it describes, and available for agency review and copying.9 Building this correctly is one of the more expensive parts of a custom GxP tool, and it is easy to underestimate.

7. Security testing and data protection. A system that holds GxP or personal data needs vulnerability scanning, penetration testing, monitoring of the open-source libraries it depends on, secure configuration of its hosting, and regular access reviews. The draft revision of EU GMP Annex 11, released for consultation in July 2025, gives security its own section and expects regulated users to keep up with new security threats and improve their protections in a timely way. It also adds dedicated sections on supplier and service management, periodic review, backup, and archiving.10 A vendor spreads its security program across its whole customer base. A company that builds pays for scanning tools, outside testing, and the staff time to act on the results.

Running It in Production

8. Hosting, backups, and recovery. Someone has to run the servers or cloud services, keep separate environments for development, testing, and production, take backups, prove the backups can be restored, and maintain a disaster recovery plan that has been tried. These are monthly cloud bills plus staff time for as long as the system is in use. For a commercial SaaS product, most of this is included in the subscription.

9. Production support, where you are Level 3. Support for a business system usually has three levels. Level 1 answers user questions and resets access. Level 2 diagnoses problems with configuration, data, and integrations. Level 3 fixes defects in the code itself. When you buy, the vendor usually provides Level 3 and often Level 2. When you build, all three levels are yours. Someone who knows the code has to be reachable when the system fails during a batch release, a submission deadline, or an inspection, and more than one person needs that knowledge so there is cover for vacations and departures. In GxP, a production failure can also lead to a deviation, an investigation, and a CAPA against the system, each with its own staff hours.

10. Monitoring and incident management. A production system needs monitoring for errors, performance problems, and unusual activity, and a process for logging incidents, assessing their GxP impact, and tracking fixes to closure. For a tool that uses an AI model, monitoring also has to cover the quality of the model’s output over time. That is a newer task for most quality teams, and it needs defined measures, sampling, and review.

Keeping It Current

11. Ordinary maintenance and technical debt. Software ages even when nobody changes it. Browsers, operating systems, libraries, and connected systems change underneath it, and any of those changes can break something. Technical debt is the accumulated effort needed to fix shortcuts and outdated design. In a McKinsey survey, CIOs estimated technical debt at 20 to 40 percent of the value of their entire technology estate before depreciation, and reported that 10 to 20 percent of the technology budget meant for new products was diverted to resolving it.11 Code written quickly, by AI or by people, tends to add to that debt unless someone spends time cleaning it up.

12. Paying for new features yourself. Commercial products improve every year without a separate charge for each improvement: new features, better reports, new integrations, and, increasingly, AI capabilities. Many of these are things a customer did not know it needed until it saw them in a release. A custom tool does not improve unless you pay for the improvement. Over five years, the gap between a commercial product that receives several releases a year and a custom tool that receives none tends to grow. Closing it means another round of design, building, testing, and validation, paid for by one company instead of shared across a customer base.

13. Keeping up with regulatory change. Regulations and guidance keep changing. FDA finalized its CSA guidance in September 2025 and reissued it in February 2026.5 The European Commission’s July 2025 consultation covered a revised Annex 11, a revised Chapter 4 on documentation, and a new Annex 22 on artificial intelligence, all of which will change expectations for computerized systems once they are final.10 A vendor reads each change once, updates the product once, and releases the update to every customer, often with documentation to support the customer’s own assessment. The builder of a custom tool has to track every relevant change, assess its impact, update the system, and revalidate it, alone.

14. AI model retirement. If the tool calls an AI model, that model is a supplier component with its own end date. Anthropic, for example, publishes a model deprecation schedule and commits to at least 60 days’ notice before retiring a publicly released model. The schedule shows how often this happens: the Claude Sonnet 4 and Claude Opus 4 versions dated May 14, 2025 were retired on June 15, 2026, and Claude Sonnet 4.5 is scheduled for retirement on November 30, 2026.12 Every forced model change is a change to a validated system. It needs an impact assessment, regression testing against the tool’s intended use, and change control, on the model provider’s schedule rather than yours. Vendors that build AI into their products handle much of this for their customers. A company that builds is responsible for all of it. Our article on AI model changes and change control gives a decision tree for these events.

15. Change control and periodic review. In a validated system, every change goes through change control: an impact assessment, approval, testing in proportion to the risk, and documentation. AI tools make changes fast and easy, which tends to increase the number of changes, and each one carries the same review and records. Periodic review adds a scheduled check that the system is still in a validated state, that user access is still appropriate, and that the change history is complete. Our article on periodic review for AI systems covers what that review needs to include.

People and Knowledge

16. Training, rollout, and help materials. A commercial product comes with user guides, release notes, help pages, and often training courses. A custom tool comes with none of them. Someone has to write the user guide, the help pages, and the training material from the beginning, and update them each time the tool changes in a way users can see. Rollout is easy to underestimate because it does not feel like part of building. Yet a tool that people do not understand, or use in different ways, is a common source of deviations.

17. Qualified people, and more than one of them. Part 11 requires a determination that people who develop, maintain, or use electronic record and electronic signature systems have the education, training, and experience for their assigned tasks.9 For a custom tool, that includes the builder, and the qualification has to be documented. The larger risk is that the knowledge belongs to one person. Many custom tools depend on whoever built them. When that person changes roles or leaves, the company is left with a validated system nobody can safely change. Gartner’s September 2026 prediction about AI built by vendor engineers describes the same problem from another direction. Gartner expects that by 2028, 70 percent of enterprises will abandon agentic AI built through vendor forward-deployed engineering, because of rising costs and because they cannot evolve it on their own.13 If organizations struggle to sustain AI that a vendor’s engineers built for them, sustaining AI that a busy employee built in a few weeks will not be easier.

18. Inspection and due diligence readiness. When an inspector asks about a commercial system, the company can point to its supplier assessment, the supplier’s quality system, and the supplier’s audit reports. For a custom tool, the company is the supplier. Its own development records, design documents, code review evidence, and test results have to answer every question without a supplier to refer to, and the people who built the system may need to explain it in person. The same questions come up in due diligence when a company is acquired or signs a major partnership. Our article on what a biotech under 200 people should not build in-house describes how custom systems slow down that process.

The End of the System’s Life

19. Records retention, migration, and decommissioning. GxP records outlast the systems that create them. Part 11 requires protecting records so they can be retrieved accurately and readily throughout the records retention period.9 For clinical trials in the EU, the Clinical Trials Regulation requires the sponsor and the investigator to archive the content of the trial master file for at least 25 years after the trial ends, on media that keep the content complete and legible for that whole period.14 When a custom tool is retired, its records have to be migrated or archived in a readable form, with their audit trails, and someone has to show that nothing was lost. Vendors usually provide export tools and documented data formats. A builder has to create them. Our article on decommissioning a GxP AI system covers what has to be kept.

20. The work that does not get done. The people who build and maintain internal tools are usually the same people a small or mid-size company needs for its science, its submissions, and its quality system. Every hour they spend acting as an in-house software team is a paid hour not spent on that work. This line rarely appears in a build estimate, and it is often the largest.

What AI Makes Cheaper, and What It Does Not

AI coding tools reduce some of these lines, leave most of them unchanged, and add to a few. Separating the three is the most useful thing a budget owner can do before approving a build.

Smaller With AI

What Gets Cheaper

Writing code. Drafting requirements, test scripts, and user guides for people to review. Producing a working prototype to test an idea. Translating help materials. Finding likely defects during code review.

Unchanged or Larger

What Does Not

Process decisions and industry practices research. Testing with no supplier evidence. Part 11 controls. Security review of generated code. Level 3 support. Regulatory change. Model retirements and the change control they trigger. Decades of record retention. Dependence on the one person who understands the system.

There is also evidence that faster code creation does not automatically mean faster or safer delivery. The 2024 Accelerate State of DevOps report from Google Cloud’s DORA research program found that AI adoption improved individual productivity, flow, and job satisfaction. It also estimated that every 25 percent increase in AI adoption was associated with a 1.5 percent reduction in delivery throughput and a 7.2 percent reduction in delivery stability. The researchers had expected the opposite.15 In a regulated company, lower delivery stability shows up as failed changes, deviations, and rework.

Agentic AI projects in particular have a high failure rate even before GxP requirements are added. In June 2025, Gartner predicted that over 40 percent of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear business value, or inadequate risk controls.16 The cancellation rate for custom GxP tools built without a full budget is unlikely to be lower.

45% of AI-generated code samples failed security tests
(Veracode, 2025)6
7.2% estimated drop in delivery stability for every 25% increase in AI adoption
(DORA, 2024)15
25 yrs minimum archiving period for an EU clinical trial master file
(EU Clinical Trials Regulation)14

A Five-Year Budget Worksheet

The worksheet below puts the 20 lines side by side for a build option and a buy option, plus the lines that apply only to buying. Fill it in with staff hours and dollar figures for your own situation. Two rules make the comparison complete. First, count five years, not one, because most hidden costs are recurring. Second, convert staff time into dollars at a fully loaded rate, so that hours from quality, IT, and the business appear next to the subscription fee instead of disappearing into existing salaries.

#Budget LineIf You BuildIf You Buy
1Process design and requirementsEvery decision and the full requirements specification, written by your teamConfiguration choices and a shorter requirements set based on the vendor’s design
2Industry practices researchYour team finds, selects, and designs them inLargely built into the product
3Risk assessment and validation planningStarts from nothing; GAMP Category 5Starts from vendor documentation; usually GAMP Category 4 (configured product)
4Building and code reviewSmaller with AI; review and fixes still requiredConfiguration time instead
5Testing and documentationAll of it, with no supplier evidenceReduced by relying on assessed supplier testing under a risk-based approach
6Part 11 controlsDesigned, built, and proven by youBuilt and tested by the vendor; you verify your configuration
7Security testingScanning, penetration testing, library monitoring, all yoursVendor’s program; you review its reports
8Hosting, backups, recoveryCloud bills plus staff timeUsually included in a SaaS subscription
9Production supportLevels 1, 2, and 3Level 1 internally; vendor usually provides Levels 2 and 3
10Monitoring and incidentsYour tools and your on-call timeShared: vendor monitors the platform, you monitor your use
11Maintenance and technical debtEvery year, for the life of the systemIncluded in the subscription
12New featuresFunded and validated by you aloneIncluded in releases; you assess and test what you adopt
13Regulatory changeYou track, assess, update, and revalidateVendor updates the product; you assess the release
14AI model retirementEach retirement is your change controlMostly handled by the vendor
15Change control and periodic reviewEvery change, with all evidence produced by youYour configuration changes plus assessment of vendor releases
16Training and help materialsWritten and maintained by youProvided and updated by the vendor
17Qualified people and backup coverageAt least two people who understand the codeAdministrators trained on the product
18Inspection and diligence readinessYour development records serve as the supplier recordsSupplier audit reports plus your supplier assessment
19Retention and decommissioningYou build export and archivingVendor export tools; you verify the result
20Time taken from core workOften the largest lineMostly selection and administration time
+Subscription and renewal increasesNot applicableAnnual fees, sometimes with automatic increases at renewal
+Implementation and configurationCovered in lines 1 to 7Vendor or partner fees plus internal time
+Supplier qualification and oversightApplies to your AI model provider and hosting providerApplies to the vendor, including audits

How to use the worksheet. Ask the person proposing the build to fill in the build column and the person who manages the vendor relationship to fill in the buy column, then compare them together.

Buying Has Hidden Costs Too

A fair comparison prices the buy side completely. Several costs of buying are easy to underestimate, and leaving them out makes the build option look worse than it is.

  • Price increases at renewal. Some contracts include automatic annual increases. Veeva’s annual report, for example, says certain of its contracts raise the price at each renewal by the lower of 4 percent or the U.S. Consumer Price Index.3 Over five years, that compounds, and it should be in the worksheet.
  • Implementation and validation of your configuration. Configuration, data migration, integrations, and validation of the configured system remain your responsibility, even when you rely on supplier evidence.
  • Supplier qualification and oversight. The draft revision of Annex 11 states plainly that relying on a vendor’s qualification of a system does not change the requirements, and that the regulated user remains fully responsible.10 That means a supplier assessment, often an audit, and ongoing monitoring. Our article on vendor qualification in a cloud-first world covers how to do this efficiently.
  • Lock-in and exit. Moving data and processes out of a commercial system at the end of a contract takes planning, time, and money. Our article on the hidden cost of AI vendor lock-in covers this side in detail.
  • A roadmap you do not control. A vendor builds what most of its customers need. A feature that matters a great deal to you may never arrive.
  • Usage-based pricing for AI features. When a vendor charges for AI features by usage, the annual bill is harder to predict. Ask for usage reporting by user and team before you sign.

For most GxP systems, these costs do not change the answer. They do belong in the worksheet, so that the comparison is complete in both directions and the decision can be defended later.

When Building Is the Right Call

Building is sometimes the right decision. The good cases share a pattern: nobody sells what you need, the tool reflects knowledge that is specific to your company, and you are prepared to own it for its whole life.

  • Analysis code that reflects your own science. No vendor sells your company’s specific analysis methods.
  • Integrations nobody sells. A connection between two systems that no vendor supports can be worth building, with an owner, a specification, and monitoring.
  • Tools outside GxP scope. For a non-GxP tool, the hidden-cost list is much shorter, and AI-assisted building is often a good choice.
  • Prototypes built to learn what you need. This is one of the best uses of AI coding tools in a regulated company. Build a working prototype in a week, use it to discover and write sharper requirements, then buy a product that meets them. The prototype is retired, not validated.
  • Configuration on a bought platform. Configuring a commercial product, GAMP Category 4, often gives most of the benefit of a custom build with far less to own.

Six Questions to Ask Before Anyone Starts Building

1

Will It Touch GxP Records or Decisions?

If it will create, change, or hold GxP records, or support a GxP decision, the full list of 20 applies.

2

Does a Commercial Product Already Do Most of This?

If so, what would building give you that configuration cannot? Write the answer down. If it is only price, finish the worksheet first.

3

Who Owns It for Its Whole Life?

Name the system owner, the person who provides Level 3 support, and the second person who understands the code.

4

Have You Priced All 20 Lines for Five Years?

Include staff time converted to dollars, recurring cloud and AI usage, and outside testing.

5

What Happens When Something Underneath It Changes?

Plan for the AI model being retired, a library becoming unsupported, and a regulation being revised.

6

How Will You Retire It?

Decide now how its records will stay complete and readable for the full retention period after the tool is gone.

Conclusion

AI made the first version of software cheap, and that is good news. More ideas get tested, more problems get a quick answer, and more people outside IT can contribute. But in a regulated company, the first version was never where most of the money went. The money goes into everything a vendor normally provides: the industry practices designed into the product, the testing evidence, the operations and support, the regulatory updates, the improvements nobody asked for, and the records that have to stay readable for decades.

So the useful question is not whether you can build it. With today’s tools, you almost always can. The question is whether you want to become the supplier, the Level 3 support team, the regulatory tracking function, and the product development team for that system for as long as it exists. Sometimes the answer is yes. For most GxP systems, a complete five-year budget points to buying, configuring what you need, and using AI to write sharper requirements, test faster, and help people get more out of the validated systems they already have.

At Sakara Digital, we help pharma and biotech teams make this decision based on real numbers: completing the five-year worksheet, assessing tools that were built without a formal decision, and deciding which ones to keep, replace, or retire. If you are deciding whether to build a GxP tool now, we are happy to talk it through.

For Further Reading