In This Article
- Executive Summary
- Why Both Of The Usual Answers Are Wrong
- Step One: Find Every Spreadsheet, Including The Ones You Are Not Supposed To Know About
- Step Two: Classify By What The File Actually Does
- What Validating A Spreadsheet Actually Means, And Where It Stops
- Step Three: Decide Per Tier, Not Per File
- The Migration Test: Which Spreadsheets Are Worth Moving At All
- The Uncontrolled Spreadsheet In An Inspection
- What One Quarter Actually Buys You
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Every quality control laboratory runs on spreadsheets nobody planned. Calculation workbooks that turn instrument output into a reportable result. Stability trackers that decide when a pull point is due. Sample logs, trend charts, reagent inventories, training matrices, out-of-specification logs. They accumulated one urgent problem at a time, and the collection now falls somewhere between the validated systems and the paper records, belonging to neither.
The two responses leaders usually choose both fail. Validating all of them turns a hundred small files into a hundred small validation packages that nobody will maintain. Replacing all of them with a new laboratory information management system turns a control problem into a two-year capital program, and the spreadsheets keep running the whole time. The workable answer is tiering. Inventory what exists, classify each file by what it actually does to a GxP record, and then apply four different decisions: validate and keep, move the logic into a system you already own, rebuild as a small controlled application, or delete.
This article gives the inventory method (including how to find files on personal drives), a three-way classification that holds up under questioning, an honest account of what spreadsheet validation can and cannot deliver under EU GMP Annex 11 and 21 CFR Part 11, the criteria that decide whether a given file is worth migrating, what the published research on spreadsheet error rates supports and what it does not, what United States Food and Drug Administration warning letters have actually said about spreadsheet use, and a quarter-length plan that produces a defensible position rather than a finished migration.
Why Both Of The Usual Answers Are Wrong
Ask a quality control director how many spreadsheets support GMP decisions in the laboratory and you will usually get a range rather than a number. The range is wide because nobody built the collection on purpose. A method transfer needed a conversion. A stability program needed a pull schedule. A regulator asked for trending on environmental data and the fastest way to produce a chart was a workbook. Each one was a good local decision. Together they form an infrastructure layer that no one owns, that changes without change control, and that produces numbers which end up in batch records and regulatory filings.
When the problem finally gets escalated, two answers get proposed in the same meeting.
Answer one: validate them all
This sounds rigorous and is usually a way of doing nothing slowly. A spreadsheet validation package that is worth having includes a requirements statement, a specification of every formula, testing against known inputs including boundary and error conditions, a protection scheme, a controlled location, a defined change procedure, and periodic review. That is real work. Multiply it by the number of files a mid-size site actually has and the program consumes a quality unit for a year while delivering nothing a patient or an inspector would notice for most of the files.
Worse, it produces a maintenance obligation in perpetuity. A validated spreadsheet is only validated while the controls hold. The moment an analyst copies it to a desktop to make a small adjustment, the package describes a file that no longer exists. Organizations that validate everything tend to end up with a shelf of documentation and a laboratory that operates on uncontrolled copies, which is a worse position than having neither, because the documentation now creates an expectation the practice does not meet.
Answer two: replace everything with a LIMS
The other proposal is to buy or extend a laboratory information management system and move all of it there. This is right for some of the content and wrong as a total answer, for three reasons.
First, timing. A LIMS implementation or major extension is measured in quarters at best and years at worst. During that entire period the spreadsheets keep running, uncontrolled, and the control problem is deferred rather than solved. Second, fit. A meaningful share of laboratory spreadsheets are not laboratory data at all. A reagent inventory, a training matrix, an instrument calibration due list, and an equipment log are administrative trackers. Forcing them into a LIMS is expensive configuration for low value, and the configuration itself becomes a validation burden. Third, arithmetic. Some calculations belong in the chromatography data system, not the LIMS, because that is where the raw data and the integration parameters already live and where recalculation is already covered by the system’s own controls.
The reframe: the goal is not to eliminate spreadsheets. It is to make sure that no spreadsheet is doing work that requires controls it cannot provide. Some files pass that test today. Some can be made to pass it with modest effort. Some cannot pass it at any effort and need to move. And a large number should simply stop existing.
Step One: Find Every Spreadsheet, Including The Ones You Are Not Supposed To Know About
You cannot tier what you have not found. The inventory is the part teams want to skip and the part that determines whether the rest of the work is credible.
Note the relationship to the wider computerized system inventory required under Annex 11 clause 4.3, which calls for an up-to-date listing of all relevant systems and their GMP functionality.6 The spreadsheet inventory is a feeder into that listing, not a replacement for it. Build it as a working document that resolves into inventory entries for the files that survive tiering, and into disposal records for the ones that do not. Treating it as a separate permanent register creates a second system of record, which is the failure mode this whole exercise is meant to avoid.
Where they actually live
Four locations, in ascending order of difficulty.
Managed shared drives and document systems. The easy population. A file crawl over the laboratory’s shared folders, filtered to spreadsheet extensions and sorted by last-modified date, produces a candidate list in an afternoon. Include macro-enabled formats, because those are the highest-risk files in the estate and the ones most likely to be treated as ordinary documents.
Attachments inside controlled systems. Spreadsheets attached to deviation records, change controls, protocols, and validation reports. These are often overlooked because the containing record is controlled and the assumption is that the attachment inherits that control. It does not inherit formula protection or a recalculation history. If an attached workbook computed a number that supported a conclusion in an investigation, that workbook is in scope.
Instrument workstations. Local drives on chromatography, spectroscopy, dissolution, and particle counting workstations. Analysts export to a local file, calculate, then print or transcribe. The local file is frequently not backed up, not retained, and not visible to anyone outside the bench. This population overlaps with the question of where raw data ends and processed data begins in chromatography systems, which is a separate discussion in its own right, but for inventory purposes the rule is simple: if a file on that workstation contributed to a reported result, it counts.
Personal drives and mailboxes. The population everyone knows about and nobody lists. This is where the honest part of the exercise happens.
How to get the personal-drive files without starting a witch hunt
The instinct is to run a discovery sweep across user home directories. It works technically and fails organizationally. People hide what they think will be taken away, and a laboratory that hides files from quality is in a much worse position than one that keeps files in the open.
The approach that works is an amnesty framed around usefulness rather than compliance. Announce a defined window. State plainly that the purpose is to find out which calculations people are doing by hand because no system supports them, that nothing will be deleted without the owner being consulted, and that the outcome for most files will be either a controlled version of the same thing or a small application that does the job better. Then ask each analyst and supervisor a short set of questions.
- What do you calculate outside a validated system, and why?
- Which file do you open at the start of a shift that is not in a system?
- What would you have to do by hand if your laptop failed today?
- Which number that leaves this laboratory passed through a spreadsheet on the way out?
The last question is the one that finds the material risk. It reframes the request from an audit of behavior into a question about data flow, and people answer it accurately because it does not sound accusatory.
Pair the amnesty with a technical sweep of managed locations so that the interview data can be cross-checked. Where the sweep finds a file that nobody declared, treat it as an inventory gap to close rather than a disciplinary matter, at least on the first pass. The first inventory is for learning what is there. Enforcement belongs to the operating model that follows it.
What to record about each file
Resist the urge to build a rich metadata model. Most of the fields will be wrong within a month and the extra effort delays the classification step, which is where the value is. Ten fields are enough.
| Field | Why it matters |
|---|---|
| File name and storage location | Identifies the file and reveals whether it is already in a controlled location. |
| Named owner (a person, not a department) | Every later decision needs someone to sign it. Departments do not sign. |
| Business purpose in one sentence | Forces the owner to state what the file is for. Files that cannot be described in one sentence are usually several files pretending to be one. |
| Inputs and where they come from | Manual entry, instrument export, or a system report. Determines whether Annex 11 clause 6 accuracy checks apply to manual entry of critical data.6 |
| Outputs and where they go | The single most important field. A number that reaches a batch record, a certificate of analysis, an investigation, or a submission changes the classification immediately. |
| Does it calculate, or only display? | Separates calculation engines from trackers and dashboards. |
| Does it contain macros or scripts? | Custom code changes the validation category and the effort profile substantially. |
| Is it the record, or a copy of a record held elsewhere? | Determines retention, backup, and archiving obligations. |
| Frequency of use | Distinguishes daily production tools from files opened twice a year. |
| Last modified date and modifier | Files untouched for two years are usually deletion candidates hiding in the list. |
Collect these in a structured form rather than free text, and have the owner complete it rather than a central team guessing. The exercise of writing the business purpose and the output destination is itself the beginning of the classification, and owners are the only people who reliably know where the numbers go.
Expect the count to be higher than the estimate and the risk to be lower. In practice the file count usually exceeds what management expected, while the number of files that actually touch a GxP record is a small fraction of the total. That combination is good news: it means the problem is large in volume and small in risk concentration, which is exactly the shape that tiering handles well.
Step Two: Classify By What The File Actually Does
Classification schemes fail when they ask about importance, because everyone’s file is important. They work when they ask a factual question with a checkable answer. Use three questions in order, and stop at the first yes.
The file is the record
The spreadsheet holds original GMP data that exists nowhere else in a controlled form. A sample receipt log that is the only record of when material arrived. A stability pull log that is the authoritative schedule. If the file is lost, the record is lost. Highest obligation: retention, backup, archiving, and the ability to detect change all apply directly.
The file performs a calculation whose result reaches a record
Inputs come from somewhere controlled, the workbook applies formulas, and the output is transcribed or attached to a certificate of analysis, a batch record, an investigation, or a submission. The file is not the record but it determines the record. Formula correctness and formula protection are the whole game here.
The file tracks or displays, and decides nothing on its own
Reagent inventory, training matrix, instrument due list, a chart summarizing data that is authoritative elsewhere. Useful, sometimes operationally critical, but no GMP record depends on the file’s arithmetic. Control obligation is real but light.
The file does nothing anyone can name
Superseded versions, personal working copies of a controlled file, one-off analyses from a project that closed, duplicates with a year in the name. This tier is always larger than expected and it is the fastest source of progress in the first month.
Getting the tier decision right
Two rules keep the classification honest.
Follow the output, not the intent. A file described as “just for trending” that produces the trend chart pasted into an annual product quality review is influencing a GMP conclusion. It may still be Tier 3 if the underlying data and the conclusion are both documented elsewhere and the chart is illustrative. It moves to Tier 2 if the chart is the only place the trend was ever evaluated. Ask where the evaluation is recorded, not what the file is called.
Split files that span tiers. Many workbooks contain a calculation tab, a log tab, and a chart tab. Classifying the whole file by its highest tier means applying Tier 1 controls to a chart, which is wasteful. Classifying it by its lowest tier is a compliance failure. The right answer is usually to split the file, which is also the first step toward moving the calculation into a system and leaving the tracker as a tracker.
Layering risk on top of tier
Tier says what the file does. Risk says how much it matters if it is wrong. The combination determines effort, and it is a straightforward application of the quality risk management principle that runs through Annex 11 clause 1 and is carried forward in the draft revision of the annex.67 Score two dimensions coarsely, on three points each.
| Dimension | Low | Medium | High |
|---|---|---|---|
| Consequence of an undetected error | Internal inconvenience only. Nothing leaves the laboratory. | Rework, a deviation, a delayed release. Detected before product moves. | A wrong result on a certificate of analysis, a wrong release decision, a wrong number in a filing, a missed out-of-specification result. |
| Likelihood an error goes undetected | Output is independently checked against another system that would disagree. | Second-person review of the printed output, no independent recalculation. | Single analyst, no second check, no downstream system that would notice. |
A Tier 2 calculation with high consequence and high likelihood of an error going undetected is the priority set. In most laboratories that set is between five and twenty files. That is a tractable number, and identifying it in the first six weeks is worth more than perfecting the classification of the remaining hundreds.
What Validating A Spreadsheet Actually Means, And Where It Stops
Some spreadsheets legitimately stay. Before deciding which, leaders need an accurate picture of what a validated spreadsheet does and does not give them, because the gap between the two is where inspection findings live.
The regulatory frame
A spreadsheet used in a GMP activity is a computerized system. Annex 11 opens with the statement that it applies to all forms of computerized systems used as part of GMP regulated activities, and that the application should be validated while the IT infrastructure should be qualified.6 There is no size threshold and no exemption for desktop tools. In the United States, 21 CFR 211.68(b) requires appropriate controls over computer or related systems to assure that changes in master production and control records or other records are instituted only by authorized personnel, along with backup of records.9
The FDA’s data integrity guidance makes the connection to spreadsheets explicit in its discussion of static and dynamic records. Describing what makes a record dynamic, the guidance notes that the format may allow the user to modify formulas or entries in a spreadsheet used to compute test results or other information such as calculated yield.5 That single sentence carries the operative point: a spreadsheet is a dynamic record, and dynamic records carry the full set of expectations around retention, review, and the ability to see what changed.
Where 21 CFR Part 11 applies, the electronic records and electronic signatures requirements apply to the spreadsheet as they would to any other electronic record system, including the controls in subpart B for closed systems.8 The draft revision of Annex 11, published for consultation jointly by the European Commission and the Pharmaceutical Inspection Co-operation Scheme, has no final text and no implementation date, so nothing in it is in force. It is still worth reading for direction of travel. Its qualification and validation section states that increased focus should be on testing a system’s handling of key functional requirements and functionality designed to ensure data integrity, and it lists calculations explicitly among the areas named.7 For a spreadsheet estate, that is a direct signal about where inspector attention is heading.
What a spreadsheet validation package contains
Under GAMP 5 second edition, software categorization is a way of scaling effort to risk and complexity rather than a fixed label.10 A spreadsheet is a useful illustration of why the categories are a continuum rather than boxes. The underlying application is commercial software in wide use. What the laboratory built on top of it, formulas, links, lookups, and in some cases macros, is the part that carries the GMP function and the part that needs the evidence. The practical consequence is that validation effort should scale with the complexity of what was built, not with the name of the product it was built in.
A defensible package for a Tier 2 calculation workbook contains, at minimum:
- A requirements statement. What the workbook must calculate, from which inputs, to what precision, with which rounding conventions, and what it must reject.
- A formula specification. Every calculation written out in a form a reviewer can check without opening the file. This is the document that survives the file.
- Test evidence against known inputs. Normal cases, boundary cases, and error cases. Test what happens with a blank cell, a text entry in a numeric field, a value outside the calibration range, and a division by zero. Annex 11 clause 4.7 calls for evidence of appropriate test methods and test scenarios covering parameter limits, data limits, and error handling.6
- A protection scheme. Which cells accept input, which are locked, which sheets are hidden, and who holds the password. The password is a controlled item, not a shared convenience.
- A controlled location and version identity. A single master copy in a location where users cannot overwrite it, with a version number visible on the printed output.
- A change procedure. How a formula change is requested, assessed, tested, approved, and released, and how the previous version is retired. This is change control, not a note in a file name.
- A periodic review. Confirmation that the file in use is still the validated file, that the protection is intact, and that the calculation still matches the method it supports. Annex 11 clause 11 requires periodic evaluation to confirm systems remain in a valid state.6
The limits, stated plainly
This is the part that gets glossed over in vendor material and it is the part senior leaders need to understand before they authorize keeping a population of validated spreadsheets.
Cell protection is a control, not a barrier. Sheet and workbook protection in common spreadsheet applications is designed to prevent accidental modification. It is not a security boundary, and passwords on protected sheets can be removed with widely available tools. It reduces the probability of an unintentional change and it does very little against a determined intentional one. Any risk assessment that treats protection as preventing deliberate alteration is overstating the control.
There is no native audit trail of the kind Annex 11 clause 9 contemplates. That clause asks for a system-generated record of GMP-relevant changes and deletions, with the reason documented for changes and deletions of GMP-relevant data, available in an intelligible form and regularly reviewed.6 A spreadsheet application does not generate that record. Change tracking features and version history in cloud document platforms provide something related but not equivalent: they typically capture that a file changed and who saved it, not that a specific formula was altered, why, and with what effect on a previously reported value. Where a platform’s version history is used as the compensating control, that use has to be specified, tested, and reviewed like any other control, and its gaps have to be documented rather than assumed away.
Attribution is weak. Annex 11 clause 12.4 expects management systems for data and documents to record the identity of operators entering, changing, confirming, or deleting data, including date and time.6 A shared network file opened under a shared session does not do this. Where the file is the record (Tier 1), this limitation is usually decisive: the file cannot meet the requirement and the data needs to move.
Copies defeat everything. The validated master exists in one place. The moment a user saves a copy to work on, the controls do not travel with it. Every control in the list above depends on a behavioral assumption that people will use the master. That assumption is the weakest link in the entire approach, and it is the reason spreadsheet validation should be reserved for the files that genuinely warrant it rather than applied broadly.
What the error research actually supports
The literature on spreadsheet errors is often cited loosely in vendor material, usually as a single dramatic percentage. The underlying picture is more useful and more defensible than the headline.
Panko’s synthesis of field audits and laboratory experiments concluded that spreadsheet errors are both common and non-trivial, that the field audits conducted from 1997 onward found errors in at least 86% of the spreadsheets examined, and that across roughly a thousand experimental subjects building spreadsheets from word problems, 51% of the resulting files contained errors even though most were only 25 to 50 cells.11
Powell, Baker and Lawson reviewed the same literature critically and were more cautious. They noted that the widely repeated figure of a 5% cell error rate rests on five studies covering only 43 spreadsheets in total, some of them informally conducted.12 When they built an explicit auditing protocol and applied it to 50 diverse operational spreadsheets, they found errors in 0.9% to 1.8% of formula cells depending on how an error was defined, and they found that the rate varied widely from one spreadsheet to another.13
Use the conservative number, because the conservative number is more than sufficient. At roughly one error per hundred formula cells, a workbook with 300 formulas is likely to contain several. A laboratory calculation that produces a single reportable result from a chain of twenty formulas has a meaningful probability of being wrong somewhere, and the wrongness will not announce itself. This is the actual argument for controlling calculation spreadsheets, and it does not require the dramatic version.
The life-sciences example that makes the point concretely is gene name corruption. A 2016 study found that a fifth of papers with supplementary gene lists in spreadsheet format contained gene names that had been converted to dates or floating point numbers without anyone noticing by default application settings.14 A follow-up covering 2014 to 2020 found the problem had not improved, identifying errors in 30.9% of 11,117 articles examined.15 Nobody chose to corrupt that data. The application did it on open, and the researchers did not notice. A quality control laboratory pasting lot numbers, batch identifiers, or sample codes into a workbook is exposed to the same class of automatic transformation.
For the wider pattern, the European Spreadsheet Risks Interest Group has collated public reports of spreadsheet failures for many years, and the collection continues to grow at a steady rate across sectors.16 The value of that record is not any single incident. It is the demonstration that these failures are ordinary rather than exceptional.
Step Three: Decide Per Tier, Not Per File
Once the inventory is classified, the decisions become policy rather than a hundred separate arguments. Four dispositions cover the estate.
Validate and keep
For Tier 2 calculations that are stable, well understood, low in volume, and not supported by any system already owned. A method-specific conversion used by one laboratory a few times a month, with a formula set that has not changed in three years, is a reasonable candidate. Apply the full package described above, put it under change control, and put it on the periodic review schedule. Accept that you are keeping it, not tolerating it, and resource it accordingly.
Move the calculation into a system you already own
For calculations that duplicate functionality already present in the LIMS or the chromatography data system. Most result calculations belong in the chromatography data system, because that is where the raw data, the integration parameters, and the recalculation controls are already held, and moving the calculation there removes an export step and a transcription step at the same time. Specification comparisons, result rounding, and reportable value determination usually belong in the LIMS. This disposition often requires configuration rather than development, and configuration inside an already validated system is a smaller undertaking than a new validation.
Rebuild as a small controlled application or a workflow in an existing platform
For Tier 1 files where the spreadsheet is the record and cannot meet attribution and change-visibility requirements, and for Tier 3 trackers that are operationally important enough to deserve real support. A stability pull schedule, a reagent inventory with expiry alerting, a training matrix, or an out-of-specification log are all natural fits for a workflow in a quality management system, an electronic laboratory notebook, or a low-code platform that already carries identity management and a change history. The rebuild is usually days to weeks rather than months, and the validation effort is proportionate because the platform underneath is already qualified.
Delete
For Tier 0 and for anything that duplicates a controlled record without adding anything. Deletion needs a documented decision, a retention check against the record retention schedule, and an owner sign-off, but it does not need a project. This is the disposition that produces visible progress fastest and the one teams are most reluctant to use, because deleting a file feels riskier than keeping it. In an inspection the opposite is true.
A default disposition by tier and risk
| Tier | Low risk | Medium risk | High risk |
|---|---|---|---|
| Tier 1: the file is the record | Rebuild as a workflow, or move into an existing controlled log | Rebuild as a workflow | Rebuild as a workflow. Do not validate in place. Attribution cannot be met. |
| Tier 2: calculation feeding a record | Validate and keep, or fold into an existing system when convenient | Validate and keep with full change control, or move to the LIMS or chromatography data system | Move to the LIMS or chromatography data system. Validate and keep only if no system supports it, and treat that as a temporary state with a review date. |
| Tier 3: tracker or display | Leave in place under a light standard: controlled location, named owner, no GMP claim | Leave in place under the light standard, or rebuild if it is operationally critical | Rebuild as a workflow. A tracker that is high risk is usually a Tier 1 record in disguise. |
| Tier 0: no identified purpose | Delete | Delete | Delete after confirming it is genuinely Tier 0 |
The light standard for Tier 3
Most of the estate ends up here, so the standard applied to it determines whether the program is sustainable. Four requirements, no validation package.
- Controlled location. The file lives on a managed drive or in a document platform with backup. Not on a desktop, not in a mailbox.
- Named owner. One person, recorded, who is responsible for the file’s continued fitness.
- An explicit statement of what it is not. A visible note on the file stating that it is a working tool and not a GMP record, and naming where the authoritative record is held. This one line prevents the most common drift, which is a tracker gradually becoming the thing people rely on because it is easier to read than the system.
- Annual confirmation. At periodic review of the laboratory’s system inventory, the owner confirms the file still serves its purpose or it is deleted. This is what stops the estate regrowing.
The Migration Test: Which Spreadsheets Are Worth Moving At All
A migration decision made file by file becomes a negotiation. A migration decision made against stated criteria becomes a review. Five questions, answered before any configuration work is scoped.
1. Does a system you already own do this?
If the LIMS or the chromatography data system already performs the calculation and the laboratory does not use that function, the problem is adoption, not capability. This is common and it is usually the cheapest win available, because the functionality is validated already and the work is training, configuration review, and a procedure change. Before scoping any development, ask the system owner to demonstrate the function. Teams are frequently surprised.
2. Will the calculation still exist in two years?
Some workbooks support methods that are being replaced, products that are being transferred, or processes that are being redesigned. Migrating a calculation into a system three months before the method retires is effort spent on something that will be decommissioned. Check the method lifecycle and the product roadmap before committing. A spreadsheet with a short remaining life is a candidate for interim controls and a documented end date, not migration.
3. Is the logic stable enough to specify?
A workbook that changes every month because the underlying approach is still being worked out is not ready to be configured into a validated system. Moving unstable logic into a change-controlled environment converts a monthly edit into a monthly change control, and the laboratory will work around it. Stabilize the logic first, in a controlled spreadsheet with a defined review cycle, and migrate when it stops moving.
4. Does moving it remove a transcription step?
The strongest migration cases are the ones that eliminate a manual handoff rather than just relocating a calculation. If the current process is export from the instrument, calculate in a workbook, then type the result into the LIMS, moving the calculation removes two error opportunities and a manual accuracy check obligation under Annex 11 clause 6.6 If the current process is a self-contained calculation from manually entered values that will still be manually entered afterward, the gain is control rather than error reduction, which is a weaker but still legitimate case.
5. What happens to the historical data?
This question is routinely deferred and routinely becomes the reason a migration stalls. Historical results in a Tier 1 workbook are GMP records subject to the retention schedule. Options are to migrate the history into the target system, to archive the workbook as a static record with a documented readability check, or to retain it in place under maintained controls until the retention period expires. All three are defensible. Choosing none of them and hoping the question does not come up is not. Annex 11 clause 17 expects archived data to be checked for accessibility, readability, and integrity, and expects the ability to retrieve it to be tested when relevant system changes are made.6
The disqualifying answer
If the answer to question one is yes and the answer to question four is yes, migrate. If the answer to question three is no, do not migrate yet. If nobody can answer question five, the migration is not scoped, whatever the project plan says.
The Uncontrolled Spreadsheet In An Inspection
The regulatory argument for this work does not rest on interpretation. Inspectors have written about spreadsheets in warning letters, in specific language, for years. The pattern in those letters is worth understanding, because it shows what triggers a finding and what the finding then does to the rest of the inspection.
The calculation nobody verified
In a September 2021 warning letter to Missouri Analytical Laboratories, the FDA wrote that the firm’s analysts used individualized non-validated spreadsheets to calculate assay, impurity, content uniformity, and dissolution test results for a variety of drug products.1 The word doing the work in that sentence is “individualized.” Each analyst had a version. There was no master, so there was no way to demonstrate that any two results had been produced the same way.
An August 2021 letter to Adamson Analytical Laboratories cited a failure to validate electronic worksheets used by laboratory personnel for microbial challenge efficacy testing, and stated that microbial worksheets reviewed were found to use unvalidated cell formulas resulting in erroneous data generation such as negative log reductions and percent reductions.2 This is the outcome the error literature predicts. The formulas were wrong, the outputs were physically impossible, and nobody caught it, because there was no test evidence and no independent recalculation. The letter also noted the firm had been cited for a similar violation in a previous warning letter, which is the pattern that turns a laboratory control finding into a quality system finding.
The workbook that was the record
A February 2020 letter to Chemland Co., Ltd. stated that the firm stored its master batch records as unlocked Excel files which were open to alteration, duplication, and deletion by unauthorized personnel, and that analysts used open Excel files for documenting sample preparation information and final calculations.3 This is Tier 1 exposure in its clearest form. The file was the record, the file had no protection, and the file’s contents could not be shown to be what they had been.
The files that disappeared overnight
A January 2025 letter to Global Calcium Pvt. Limited described investigators observing Microsoft Excel documents on a desktop computer in a production office on the first morning of the inspection, including files relating to cleaning validation samples and production details, and then returning on the second day to review those files and finding that all of them had been deleted and could not be recovered.4 Whatever the intent, the position is unrecoverable. Files that exist outside any controlled location can be removed without trace, and once an investigator has seen them the absence of a trace is itself the finding.
The escalation pattern. Read across these letters and the sequence is consistent. An investigator finds a spreadsheet performing a GMP function. There is no validation record, so the results it produced cannot be relied on. That calls into question every result the file touched, which extends the scope of the observation from one file to a body of data. Because the file was not in the computerized system inventory, the finding also becomes evidence that the inventory is incomplete, which moves the issue from the laboratory into the quality system. One workbook becomes a data reliability question and a governance question in the same conversation.
What this means for how you present the estate
A laboratory that can produce a current spreadsheet inventory, a stated classification, a validation package for the files it kept, and disposal records for the files it removed is in a fundamentally different position from one that cannot, even if both have spreadsheets in operation on the day of the inspection. The first is managing a known population under a documented risk-based approach. The second is discovering its own estate in front of an investigator.
This is why the quarter-length version of the work has value even though it does not finish the migration. Completing the inventory and the classification, and being able to show the disposition decisions with dates against them, converts an open-ended exposure into a managed program with a plan. That is a defensible answer. “We are replacing all of this with a new LIMS in 2028” is not, because it says nothing about the eighteen months in between.
Practitioner commentary in the analytical literature has made the same point for years. McDowall’s treatment of spreadsheet use alongside chromatography data systems examines, through case studies, how post-run calculation in spreadsheets creates regulatory exposure and when spreadsheets should and should not be used in a regulated laboratory.17 The consistent conclusion is that the calculation should happen where the data already is.
What One Quarter Actually Buys You
Here is what is genuinely achievable in thirteen weeks with a part-time working group, roughly one full-time equivalent split across a quality unit representative, a laboratory supervisor, and someone from the computerized systems or validation function. This is not a migration plan. It is the plan that makes a migration scopeable and defensible in the meantime.
| Weeks | Activity | Deliverable at the end |
|---|---|---|
| 1–2 | Charter the work. Agree the tier definitions and the risk scoring with the quality unit before any files are examined, so the criteria cannot be argued backward from a result. Announce the amnesty window and the four questions. | A one-page classification standard, approved. An announced amnesty with a closing date. |
| 3–5 | Run the technical sweep of managed locations and instrument workstations. Collect owner submissions from the amnesty. Cross-check the two lists. | A raw inventory with the ten fields populated by named owners. Expect gaps; record them as gaps rather than filling them by guessing. |
| 6–7 | Classify. Work through the inventory tier by tier with the owners in the room. Split multi-purpose workbooks. Score risk on the two dimensions. | A classified inventory. The priority set of Tier 2 high-risk calculation files, named. |
| 8 | Delete Tier 0. Documented decision, retention check, owner sign-off, removal. Do this before anything else so the estate stops growing and the team sees progress. | Disposal records. A materially smaller inventory. In most laboratories this step removes a substantial share of the file count. |
| 9–11 | Validate the priority set. Full packages on the five to twenty highest-risk Tier 2 calculations: requirements, formula specification, testing including boundary and error cases, protection, controlled location, change procedure. | The files that most directly determine reported results are under control, with evidence. |
| 12 | Apply the light standard to Tier 3. Move files to controlled locations, assign owners, add the “not a GMP record” statement, register the annual confirmation. | The bulk of the estate is in known locations with known owners and no ambiguity about status. |
| 13 | Write the disposition plan for everything remaining. For each Tier 1 file and each Tier 2 file not validated in the quarter, state the target disposition, the target system, the owner, and the date. Fold the surviving entries into the computerized system inventory. | A dated migration plan tied to real system capability, and a spreadsheet population that is fully accounted for. |
What will not be finished, and why that is acceptable
At the end of the quarter, the migrations have not happened. Tier 1 rebuilds and Tier 2 moves into the LIMS or the chromatography data system take longer than thirteen weeks and depend on system roadmaps outside the laboratory’s control. Configuration work in an existing validated system carries its own change control and testing obligations, and it competes with everything else in the system owner’s queue.
What has happened is more valuable than a partial migration would have been. The estate is known. The highest-risk calculations are under control with evidence. The unnecessary files are gone. Everything else has a named owner, a stated disposition, and a date. The laboratory can answer the inspector’s question, and the organization can now scope the migration against a real list instead of an estimate.
The sequencing point that matters most. Delete before you validate. Teams instinctively start with the interesting files and leave cleanup for later, which means the validation effort is sized against an estate that includes hundreds of files nobody needed. Removing Tier 0 first shrinks the problem, and it shrinks it in the direction of the files that were never going to be defensible anyway.
What sustains it afterward
An estate that was cleaned once and then left alone regrows in about eighteen months, because the conditions that created it are still there. Three controls keep it from happening.
A stated rule about new spreadsheets. Not a ban, which will be ignored, but a rule with a low threshold: any new spreadsheet whose output will reach a GMP record must be registered before first use, and registration triggers the tier question. A five-minute form is enough. The point is that the file becomes visible on creation rather than at the next inventory.
A named home for the demand. Most Tier 1 and Tier 2 spreadsheets exist because someone needed something a system did not provide and had no route to request it. If that route still does not exist after the cleanup, the next workbook is already being built. Giving the laboratory a real intake path into the systems team, with a service level that beats building a workbook, addresses the cause rather than the symptom.
A place in periodic review. The annual confirmation for Tier 3 and the periodic review for validated Tier 2 files both belong in an existing review cycle rather than in a standalone process. A control that has its own separate schedule is a control that gets skipped in a busy quarter.
Conclusion
The spreadsheet population in a quality control laboratory is not a discipline problem and it is not a technology gap. It is the visible record of every occasion when the laboratory needed something the systems did not provide and solved it locally. Treating it as misconduct produces hidden files. Treating it as a system selection problem produces a multi-year program that leaves the exposure open the whole way through. Treating it as an estate to be inventoried, classified, and dispositioned produces a defensible position within one quarter and a real migration plan at the end of it.
The judgment that matters is proportion. A small number of calculation workbooks genuinely determine reported results and deserve full validation, full change control, and eventual migration into a system that can attribute changes properly. A larger number are records that a spreadsheet cannot adequately hold and need to move. The largest group by count are trackers that need a controlled location, a named owner, and an honest label, and nothing more. And a substantial share should simply be deleted, which is the disposition with the best ratio of risk reduction to effort and the one most consistently avoided.
Sakara Digital works with pharma and biotech organizations on exactly this kind of scoping problem: separating the part of a legacy estate that needs formal validation from the part that needs a smaller answer, and sequencing the work so that the highest-risk items are controlled first rather than last. If you are looking at a laboratory spreadsheet population and trying to decide whether it is a validation program, a systems program, or a cleanup, we are happy to have that conversation.
For Further Reading
For Further Reading
References & Sources
- U.S. Food and Drug Administration. “Missouri Analytical Laboratories Inc – 615319 – 09/30/2021.” Warning Letter, September 30, 2021. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/missouri-analytical-laboratories-inc-615319-09302021
- U.S. Food and Drug Administration. “Adamson Analytical Laboratories, Inc. – 614644 – 08/17/2021.” Warning Letter, August 17, 2021. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/adamson-analytical-laboratories-inc-614644-08172021
- U.S. Food and Drug Administration. “Chemland Co., Ltd. – 593158 – 02/11/2020.” Warning Letter, February 11, 2020. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/chemland-co-ltd-593158-02112020
- U.S. Food and Drug Administration. “Global Calcium Pvt. Limited – 692000 – 01/16/2025.” Warning Letter, January 16, 2025. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/global-calcium-pvt-limited-692000-01162025
- U.S. Food and Drug Administration. “Data Integrity and Compliance With Drug CGMP: Questions and Answers, Guidance for Industry.” December 2018. https://www.fda.gov/media/119267/download
- European Commission. “EudraLex Volume 4, Good Manufacturing Practice, Annex 11: Computerised Systems.” Revision 1, effective June 30, 2011. https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf
- European Commission and PIC/S. “Annex 11: Computerised Systems (draft for public consultation).” EudraLex Volume 4 consultation document, 2025. https://health.ec.europa.eu/document/download/40231f18-e564-4043-94de-c031f813d38b_en
- Electronic Code of Federal Regulations. “21 CFR Part 11: Electronic Records; Electronic Signatures.” https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- Electronic Code of Federal Regulations. “21 CFR 211.68: Automatic, mechanical, and electronic equipment.” https://www.ecfr.gov/current/title-21/chapter-I/subchapter-C/part-211/subpart-D/section-211.68
- International Society for Pharmaceutical Engineering. “GAMP 5 Guide, 2nd Edition: A Risk-Based Approach to Compliant GxP Computerized Systems.” 2022. https://ispe.org/publications/guidance-documents/gamp-5-guide-2nd-edition
- Panko, Raymond R. “Spreadsheet Errors: What We Know. What We Think We Can Do.” Proceedings of the Spreadsheet Risk Symposium, European Spreadsheet Risks Interest Group, Greenwich, July 2000. https://arxiv.org/abs/0802.3457
- Powell, Stephen G., Kenneth R. Baker, and Barry Lawson. “A critical review of the literature on spreadsheet errors.” Decision Support Systems 46(1), 2008, pages 128 to 138. https://mba.tuck.dartmouth.edu/spreadsheet/product_pubs_files/literature.pdf
- Powell, Stephen G., Kenneth R. Baker, and Barry Lawson. “Errors in Operational Spreadsheets.” Journal of Organizational and End User Computing 21(3), July to September 2009, pages 24 to 36. https://mba.tuck.dartmouth.edu/spreadsheet/product_pubs_files/errors.pdf
- Ziemann, Mark, Yotam Eren, and Assam El-Osta. “Gene name errors are widespread in the scientific literature.” Genome Biology 17, article 177, August 23, 2016. https://link.springer.com/article/10.1186/s13059-016-1044-7
- Abeysooriya, Mandhri, Megan Soria, Mary Sravya Kasu, and Mark Ziemann. “Gene name errors: Lessons not learned.” PLOS Computational Biology 17(7), July 30, 2021. https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008984
- European Spreadsheet Risks Interest Group. “Horror Stories: spreadsheet mistakes news stories.” Collated by Patrick O’Beirne and contributors. https://eusprig.org/research-info/horror-stories/
- McDowall, R.D. “Are Spreadsheets a Fast Track to Regulatory Non-Compliance?” LCGC North America 38(6), 2020, pages 346 to 354. https://www.chromatographyonline.com/view/are-spreadsheets-a-fast-track-to-regulatory-non-compliance








Your perspective matters—join the conversation.