Thirty-Year Retention Is an Engineering Problem, Not a Policy Statement

Ask a quality organization about records retention and you will usually be handed a retention schedule. It is a table. It lists record types down one side and periods across the other. It is approved, it is version controlled, and it is almost always correct as far as it goes. What it does not do is describe how a record written in 2026 will be readable in 2056.

That gap matters because the regulatory expectation is not simply that records exist. It is that they remain accessible, readable, and intact for the whole period. The MHRA defines an archive as a designated secure area or facility for the long-term retention of data and metadata, and states that archived data should permit recovery and readability of the data and its associated metadata throughout the retention period. The World Health Organization’s data integrity guideline makes the same point: records must remain legible and available for the full retention period, and the means of retrieval must be maintained.3 EudraLex Volume 10 guidance on computer systems in clinical trials requires that archived data be readable and that the ability to retrieve it be tested when relevant changes are made.4

Those are engineering requirements written in regulatory language. “Readable in 2056” is a statement about file formats, character encodings, rendering software, decryption keys, checksum algorithms, and the availability of people who understand what a given column header meant. None of that is settled by a retention schedule.

The mismatch between record life and system life

Consider the arithmetic. A validated laboratory information management system might run for twelve to fifteen years before it is replaced. A chromatography data system might last ten. An electronic trial master file platform is typically re-tendered on a five to seven year cycle. A cloud application’s underlying platform may change materially within three years, without the customer choosing anything.

Against that, the obligations run much longer. A nonclinical study record under Good Laboratory Practice must be kept at least five years after the results are submitted to the FDA in support of an application, and at least two years after that application is approved, with the longest applicable period governing.5 A clinical trial master file in the EU runs to 25 years minimum.1 Pharmacovigilance records can run indefinitely while a product remains authorized, then ten years beyond.2 For a product authorized for thirty years, the pharmacovigilance tail alone reaches forty.

25 years Minimum archiving period for the clinical trial master file after the end of a trial, under EU Regulation 536/2014, Article 58
10 years Minimum retention for pharmacovigilance documents after the marketing authorization ends, per EMA GVP Module I
700+ File formats with published preservation action plans in the US National Archives digital preservation framework

The last figure is the interesting one. A national archive with a statutory permanent-preservation mandate has found it necessary to publish risk assessments and preservation actions for more than seven hundred distinct file formats.6 That is the scale of the problem when someone treats format survival seriously. Most regulated companies have never counted the formats in their own archive.

What “designed for thirty years” actually means

A designed answer has four properties that a policy statement does not.

  • It names the formats. Not “electronic records” but the specific formats each record type will be preserved in, with a stated reason for each choice.
  • It names what is lost. Every preservation choice discards something. A design that does not say what it discards has not been thought through.
  • It has a scheduled test. Retrievability is verified on a defined frequency, with documented evidence, not assumed until someone asks.
  • It has an end. The design covers disposal as explicitly as it covers retention, including the evidence that disposal was authorized and complete.

None of this is exotic. The archival community has worked on it for three decades and produced usable standards. What is unusual in life sciences is treating those standards as applying to GxP records rather than to libraries and museums.

Determining the Retention Period You Actually Owe

Before anything can be designed, the period has to be known. This turns out to be harder than it looks, because several separate obligations attach to the same physical record, and they do not expire together.

A batch record for a commercial product carries a GMP obligation tied to the expiration date of the batch. If that batch supported a regulatory submission, a second obligation attaches. If the product is the subject of litigation or a legal hold, a third. If the batch was made under a contract that requires the manufacturer to retain evidence for a defined period after termination, a fourth. Tax and corporate record rules may add a fifth. Each of these has a different trigger event and a different clock.

The obligations that stack on a single record

Obligation typeTypical triggerWhat makes it hard to compute
GMP product lifecycle Batch expiration date, or distribution date for products without expiry dating Expiry can be extended by stability data after the record was created, moving the clock forward
Clinical trial records End of trial as defined in the protocol and the regulation The EU sets a 25-year floor; sponsor policy and other jurisdictions may set a longer one1
Nonclinical study records Submission date, approval date, or study completion, whichever applies Which trigger applies depends on whether and when the study was ever submitted5
Pharmacovigilance End of the marketing authorization The clock does not start while the product remains authorized anywhere in scope2
Contractual Contract termination or expiry Terms differ by counterparty and are rarely visible to the archive team
Legal hold Notice of anticipated or actual litigation Suspends disposal indefinitely and can be layered several deep on the same record
Privacy limitation Purpose for processing personal data ends Pushes in the opposite direction, requiring deletion rather than retention7

In practice the answer is nearly always the longest applicable period, with one important exception. Privacy law pulls the other way. Article 5(1)(e) of the GDPR requires that personal data be kept in a form permitting identification of data subjects for no longer than is necessary for the purposes for which it is processed.7 A record that contains identifiable personal data and is being kept forever because nobody computed an end date is not a conservative choice. It is a compliance failure in a different regime.

How to determine and document the period

The workable method is to compute the period once per record class rather than per record, and to record the reasoning rather than only the answer.

  1. Define the record class narrowly enough that the obligations are uniform. “Laboratory records” is too broad. “Release testing raw data for commercial product, EU and US markets” is a class where the obligations are the same for every member.
  2. List every obligation that attaches, with its legal or contractual source. Include the ones that shorten the period, not only the ones that lengthen it.
  3. Identify the trigger event for each and where that event is recorded. This is where most schedules break down. If the trigger is “batch expiration date” and the archive has no reliable feed of expiration dates, the schedule cannot be executed.
  4. Compute the controlling period and write down why. A one-paragraph rationale per class, reviewed by legal and quality, is worth more than a table of numbers nobody can defend.
  5. Record the review date. Obligations change. The EU clinical trial regulation moved the floor for many sponsors. A schedule with no review cadence quietly goes stale.

A test worth running. Pick one archived record at random and ask the team to state, without looking anything up afterward, the exact date on which it becomes eligible for disposal and the evidence that supports that date. If the answer requires a research project, the retention schedule is a document rather than a control.

The alternative that many organizations fall into is keeping everything forever. It looks safe and it is not. It defers the determination problem rather than solving it, it accumulates privacy exposure, and it makes the eventual clean-up harder every year. The OECD advisory document on GLP archives makes the point plainly in its own domain: an archive requires defined procedures for retention periods and for the disposal of material at the end of them, not simply indefinite storage.8

Format Obsolescence: The Failure Mode Nobody Budgets For

Here is the centerpiece. A record retained for thirty years will outlive the application that created it, its file format, its operating system, and quite possibly its vendor. The question is not where the bytes sit. Bits are cheap and durable if you keep copying them. The question is whether anyone can render and interpret those bits decades later.

The Library of Congress has spent twenty years cataloging what makes a digital format survivable, and the seven factors it identifies are the right checklist for a GxP archive as much as for a national collection: disclosure, adoption, transparency, self-documentation, external dependencies, impact of patents, and technical protection mechanisms.9

The seven factors, read as GxP risk

  • Disclosure. Is the specification published? A proprietary binary instrument format with no public specification is a record you can only read while the vendor exists and chooses to support you.
  • Adoption. How widely is the format used? A format used by one instrument line in one company has no ecosystem to keep it readable.
  • Transparency. Can the content be understood by direct inspection? A text-based format with readable structure can be recovered by a competent person with no original software. A compressed proprietary container cannot.
  • Self-documentation. Does the file carry its own metadata? A file that contains its units, its acquisition parameters, and its provenance is interpretable on its own. One that relies on a database elsewhere is not.
  • External dependencies. Does rendering require a specific runtime, license server, font, codec, or network call? Every dependency is a separate thing that has to survive thirty years.
  • Patents. Can a third party restrict the writing of a reader? Rare but real.
  • Technical protection. Is the file encrypted or DRM-wrapped? If so, the key management scheme now has the same retention period as the record, and key loss destroys the record as completely as deleting it.

The encryption trap. Encrypting archived records at rest is good security practice and it introduces a thirty-year key management obligation that almost nobody documents. If the key custodian process, the key escrow, and the algorithm agility plan are not part of the archive design, an organization has built a record that can be destroyed by an HR event. This applies with equal force to digital signatures whose validation depends on certificate chains that expire long before the record does.

Where obsolescence actually bites in regulated data

Three categories account for most of the real exposure.

Instrument raw data. Chromatography, mass spectrometry, spectroscopy, flow cytometry, and imaging systems mostly write proprietary binary formats. These are the records regulators care most about, because they are the original observations. They are also the least portable. The vendor’s own format may change between major software versions, and reading a twenty-year-old file may require an old version of the software that will not install on a supported operating system.

Structured data in retired applications. When a LIMS or an electronic batch record system is decommissioned, the data is usually extracted to files or to a relational archive. What survives is the table content. What often does not survive is the application logic that gave the content meaning: the lookup tables, the calculated fields, the workflow states, the display layouts that determined what a reviewer actually saw when they approved something.

Compound documents. Reports with embedded objects, spreadsheets with external links, documents with linked images, and anything that references a network path. These render perfectly today and become progressively hollow as the referenced things disappear.

The UK National Archives maintains PRONOM, a public technical registry of file formats and the software that reads them, precisely because knowing what a file is turns out to be a prerequisite for keeping it readable.10 A regulated organization that cannot produce an inventory of the formats in its own archive has not started the work.

Migration, Emulation, and Rendition: What Each One Preserves and Loses

There are three durable strategies, plus the option of doing nothing and hoping. Each of the three preserves something different and discards something different. Serious archives use more than one, deliberately, for different record classes.

STRATEGY 1

Migration to an open format

Convert the record to a published, widely adopted, self-documenting format at ingest or on a schedule. Preserves the content and makes it independent of the original vendor. Loses whatever the open format cannot express, and introduces a conversion step that must itself be verified.

STRATEGY 2

Emulation of the original environment

Keep the original files and run the original software on emulated hardware. Preserves behavior and interactivity that no conversion can capture. Loses simplicity: the emulator, the operating system image, and the application licenses all now have their own retention problem.

STRATEGY 3

Readable rendition alongside the native file

Keep the native file untouched and store a human-readable rendition next to it. Preserves guaranteed readability now and the option to recover more later. Loses storage efficiency and creates two things that must be kept in step and proved to correspond.

STRATEGY 4

Do nothing and hope

Keep bytes, trust the vendor, revisit when someone asks. Preserves budget. Loses the record, silently, at an unknown future date, with discovery typically occurring during an inspection or a legal request.

Migration, honestly assessed

Migration is the workhorse. Converting instrument data to an open exchange format, or documents to an archival PDF profile designed for long-term preservation, removes the dependency on a single vendor and puts the record into something with a large enough user base to stay readable.

What migration loses is real and should be stated rather than glossed over. Converting a chromatography data file to a text-based exchange format typically preserves the acquired signal and the acquisition parameters. It generally does not preserve the ability to reprocess: to change integration parameters and regenerate the result the way the original software would have. For a record where reprocessing capability is part of what makes the data meaningful, that is a material loss. The MHRA’s position that data must be retained in dynamic form where the dynamic nature is critical to its integrity or later verification is the relevant constraint here. A static extract of something that needs to stay dynamic is not a compliant archive.

Migration also introduces a verification obligation. Every conversion must be performed under change control with documented evidence that the converted record is a true and complete representation of the original. The Digital Preservation Coalition’s guidance on preservation action treats this as normal practice: preservation actions are planned, tested on samples, verified, and documented, not run as a bulk job and assumed to have worked.11

Emulation, honestly assessed

Emulation runs historical software on current hardware so that the original files can be opened by the original application. For records where behavior matters, this is the only strategy that preserves it. The Council on Library and Information Resources’ overview of emulation as a preservation method sets out both the appeal and the difficulty: emulation keeps the original object intact and reproduces the original experience, but it moves the obsolescence problem up a layer rather than eliminating it, because the emulator itself is software running in a changing environment.12

For a regulated archive, emulation carries three specific burdens. The software licenses must permit it, and perpetual licenses for retired products are not always obtainable. The emulated environment is a computerized system in its own right and, if it is used to produce a record for a regulatory purpose, someone will ask whether it is qualified. And the skills to operate a twenty-year-old application do not automatically persist in the organization.

Emulation is worth the effort for a small number of high-value record classes: the ones where a regulator or a court might need to see exactly what the original reviewer saw. It is not a general answer.

Rendition alongside the native file, honestly assessed

The third strategy is the pragmatic one and, in our experience, the most under-used. Keep the native file exactly as it was written, unmodified, with its checksum. Alongside it, store a rendition in a format guaranteed to be readable: a self-contained archival PDF of the report, a text export of the underlying table, an image of the chromatogram as displayed at the moment of approval.

What this preserves is a floor. Whatever happens to the native format, there is always something a human can read and an inspector can be shown. What it preserves in addition is optionality: the native file is still there, so if a reader is ever needed and obtainable, more can be recovered later.

What it loses is storage efficiency, which matters less every year, and it creates a correspondence obligation. The rendition and the native file must be demonstrably of the same record, generated at a known point, with the relationship recorded. If the rendition was produced years later from a possibly-degraded source, its evidential value is weaker.

A workable default. For most GxP record classes, the combination that holds up is: keep the native file with a checksum, generate a readable rendition at the moment of approval rather than at decommissioning, migrate to an open format where a credible one exists for that data type, and reserve emulation for the handful of classes where dynamic behavior is genuinely part of the record. Write down which classes got which treatment and why.

The Audit Trail and the Metadata Are Part of the Record

This is the most common real failure, and it is worth stating as plainly as possible. An archived result without its audit trail, without its context, and without the identity of who approved it is not a complete GxP record. It is a number. Many archiving approaches export the data and drop exactly that.

The pattern is easy to fall into. A decommissioning project extracts the results tables because those are the obvious thing to keep. The audit trail lives in a separate schema, sometimes in a proprietary internal structure, sometimes in a log file that the extract tool does not know about. The user directory that maps user identifiers to actual people is a different system entirely and is decommissioned on its own schedule. The electronic signature manifestations are rendered at display time from data the extract did not capture. Six months after go-live on the replacement system, all of that is gone, and nobody notices until an inspector asks who approved a specific result in 2019.

What a complete archived GxP record contains

ComponentWhy it is part of the recordCommon failure at archiving
The result or content itself The observation being preserved Rarely lost; this is what everyone remembers to keep
Acquisition and processing metadata Makes the result interpretable: units, method, instrument, parameters, calculation version Held in application configuration rather than in the data, so the extract misses it
The GxP audit trail Shows what changed, when, by whom, and why. Required for review and for demonstrating integrity Stored separately, in a different format, or truncated by a retention setting inside the source application
Identity resolution Maps user IDs in the audit trail to real, identifiable people Depends on a directory service that is decommissioned independently
Electronic signature manifestation Shows the printed name, date and time, and meaning of the signature Rendered dynamically at display time and never persisted as data
Relationships and context Links the result to its batch, study, subject, sample, specification, and deviation Expressed as foreign keys that lose meaning once the parent tables are not archived together
Integrity evidence Checksums and a chain of custody proving the record has not changed since archiving Not generated at all, so there is nothing to compare against later

The regulatory basis for treating all of this as one record is well established. The MHRA’s data integrity guidance defines a true copy as an exact verified copy of an original record that preserves the meaning of the data, including date formats, context, layout, electronic signatures and authorizations, and the full GxP audit trail. The OECD advisory document on GLP data integrity makes the same point for nonclinical work: metadata and audit trails are part of the raw data and must be retained with it.13 EudraLex Volume 10 Annex III requires that clinical trial data be archived in a way that keeps it available and readable together with the information needed to interpret it.4

The audit trail retention setting that quietly destroys records

Check this before your next decommissioning. Many applications have a configurable audit trail retention or purge setting, often defaulted by the vendor to a period far shorter than the record retention period. It is set once during implementation, usually for performance reasons, and forgotten. The result is a system where the data has a twenty-five-year obligation and the audit trail for that data was purged after three years. This is discoverable in minutes by looking at the configuration, and it is worth checking on every validated system in the estate rather than only the one being retired.

Practical ways to keep the record whole

Three approaches work, in rough order of preference.

Archive the record as a package, not as tables. The archival community solved this problem with the concept of an information package: the content plus everything needed to understand, use, and prove the provenance of that content, bundled together and treated as one unit. The OAIS reference model, published as an international standard, defines this structure and the preservation description information that travels with the content, including provenance, context, reference, fixity, and access rights.14 A GxP archive built this way holds a package per batch, per study, or per sample, rather than a set of related tables in a database that has to be reassembled by someone who understands the original schema.

Resolve identities at archiving time. Do not archive a user identifier and assume the directory will still be there. Substitute or supplement it with the printed name and role as they were at the time of the action, stored as data in the archive package.

Persist rendered signatures. If a signature manifestation is generated at display time, generate it once at archiving time and keep the output. The same applies to any calculated field that a reviewer relied on.

Designing the Archive: What Stays Active, What Moves, What the Archive Must Guarantee

With the period known, the formats chosen, and the completeness of the record defined, the architecture becomes tractable. The core decision is what belongs in an active system and what moves out.

The line between active and archived

The useful test is not age. It is whether the record is still being used to make decisions. A batch record for product still in distribution is active whether it is two years old or ten, because it will be consulted during an investigation or a complaint. A trial master file for a completed study whose product was never approved is archival on the day the study closes.

Three tiers, defined by use rather than by date, work in most organizations.

1

Active

Records in systems that are still creating, changing, or routinely reading them. Full application functionality, full audit trail capture, normal backup and disaster recovery. No special preservation treatment beyond making sure the exit path exists.

2

Near-line

Records no longer being changed but still consulted often enough that retrieval needs to be quick. Typically read-only within the originating system, or in a queryable archive that retains the application’s interpretation of the data. This tier is where most organizations should keep records during the years when investigations and inquiries are most likely.

3

Deep archive

Records that must be kept but are consulted rarely. Application-independent packages, preserved formats, fixity checking, and a documented retrieval procedure that does not depend on the original system existing. Retrieval measured in days rather than seconds is acceptable here if the procedure is proven.

Most failures happen at the transition from tier two to tier three, because that transition usually coincides with a decommissioning project that is being judged on how quickly the old system can be turned off.

What the archive must guarantee

Whatever technology holds the deep archive, it has to be able to demonstrate four things. These are worth writing into the requirements document for any archive platform, and into the quality agreement for any third party that provides archiving as a service.

GuaranteeWhat it means in practiceHow it is evidenced
Integrity The bits have not changed since ingest, and any change would be detected Cryptographic checksums computed at ingest, stored separately from the content, and re-verified on a schedule with results recorded
Immutability No one, including an administrator, can silently alter or delete an archived record Write-once storage or object lock, privileged access management, and an audit trail of the archive system itself that is separately protected
Retrievability A named person can locate and produce a specific record within a defined time, without the original system Search and retrieval tested against real requests on a defined frequency, with the elapsed time recorded
Provable stability You can show an inspector that the record produced today is the record that was archived Fixity history: the checksum at ingest, every verification since, and any preservation action performed, all held as part of the package

The fourth is the one most often missing. Producing a record for an inspector is not enough on its own. The question that follows is how you know it has not changed. An archive that keeps a fixity history can answer with evidence. An archive that cannot is relying on assertion.

On third-party archiving services. Outsourcing storage does not outsource the obligation. The quality agreement needs to address format migration decisions, notification before any change to storage technology, the fixity evidence the provider will supply, retrieval service levels tested rather than promised, and the exit provisions: what happens to your records, in what format, on what timeline, if the provider is acquired, changes strategy, or fails. The exit clause matters more than the service level, because a thirty-year obligation will outlive most vendor relationships.

Testing Retrievability on a Schedule, Not During an Inspection

Everything above is design. This section is the operating control that keeps the design honest, and it is the cheapest insurance available.

The regulatory expectation is explicit. Archived data should be checked for accessibility, readability, and integrity, and where relevant changes are made to the system, the ability to retrieve the data must be ensured and tested. WHO’s data integrity guideline requires that the means to retrieve data be maintained and periodically verified for the duration of the retention period.3 These are not aspirational statements. They describe a periodic test with a documented outcome.

What a retrievability test should look like

The test that produces useful information is a full end-to-end retrieval performed by someone who does not already know the answer, against a randomly selected record, timed and documented. Not a system health check. Not a report that the storage is available.

  1. Select randomly, across formats and vintages. A sample that includes the oldest records and the most exotic formats, not only the recent and convenient ones.
  2. Have the retrieval performed by the role that would actually do it. If the only person who can retrieve a 2011 chromatogram is one specialist who is planning to retire, that is the finding.
  3. Render the record, do not just fetch the file. The test passes when a human can read the content, the metadata, the audit trail, and the signature. Successfully downloading an unreadable file is a failure.
  4. Verify fixity as part of the test. Compare the current checksum against the one recorded at ingest.
  5. Record the elapsed time. Retrieval time is the number that tells you whether you can meet an inspection request, and it is the number that drifts as systems age.
  6. Raise findings through the normal quality process. A failed retrieval test is a deviation. Treating it as an IT ticket removes the visibility that makes the control worth having.

Frequency that works. Annual for each major record class, with an additional test triggered by any of the following: a storage technology change, an archive platform upgrade, a change of archiving provider, a migration or preservation action, or the decommissioning of any system that the archive depends on for interpretation. The triggered tests catch more problems than the calendar-driven ones, because change is what breaks retrievability.

The value of running this on a schedule is not only that it finds problems. It is that it produces a documented history of successful retrieval spanning years, which is a considerably stronger answer to an inspector than a policy stating that records are retrievable.

Defensible Disposal: The Half of the Problem Nobody Plans

Almost every retention program stops at the point where the period ends. The record sits there. Nobody has authority to delete it, nobody wants the responsibility, and the safest-looking option is to keep it. This is a mistake with real consequences.

Why over-retention is a liability

ISACA’s analysis of data over-retention sets out the exposure clearly: retaining data past its required period increases litigation and regulatory risk, expands the scope of what must be preserved and produced under legal hold, enlarges the damage from a breach, and creates direct conflict with privacy laws that require deletion when the processing purpose ends.15 Records kept beyond their period are still discoverable. They still have to be searched, reviewed, and produced. They still contain personal data that a data subject can ask about.

The counsel’s perspective is the same. Guidance on defensible disposition of data makes the point that organizations without a disposal program end up preserving everything by default, which raises the effort of every future legal matter and undermines the defensibility of their preservation decisions, because a company that never deletes anything cannot easily explain why a particular thing is missing.16

What makes disposal defensible

Defensible disposal is not the same as deletion. Deletion is a technical act. Defensible disposal is deletion plus the evidence that it was authorized, that it followed a consistently applied policy, that no hold was in force, and that what was destroyed was exactly what the authorization covered and nothing else. Four elements are required.

ELEMENT 1

A policy applied consistently

Disposal follows a written schedule that is applied the same way to every record in a class. Selective deletion, or a policy that is applied only when convenient, is worse than no policy because it invites the inference that the missing records were chosen.

ELEMENT 2

A hold check that actually blocks

Disposal is suspended for any record under legal hold, regulatory inquiry, or an active investigation. The check must be technical, not procedural. A record under hold should be incapable of being disposed of, not merely policy-protected.

ELEMENT 3

Documented authorization

A named, appropriately senior person authorizes each disposal event against a defined scope, with the retention rule and trigger date cited. Quality and legal review is part of the approval, not a courtesy notification afterward.

ELEMENT 4

A certificate of destruction

Evidence of what was destroyed, when, by what method, by whom, and under which authorization. The certificate is itself a record with its own retention period, and it should identify the records at class and identifier level so that a future question can be answered precisely.

The certificate of destruction is the element that most often does not exist, and it is the one that matters when a record is later requested. Being able to state that a specific record was destroyed on a specific date under a specific authorization, in accordance with a schedule that predated the request, is a complete and defensible answer. Being unable to explain why something is missing is not.

Disposal in cloud and backup systems. Deleting the record from the archive does not delete it from backups, replicas, disaster recovery copies, prior migration staging areas, or the retired system’s decommissioning extract sitting on a file share. A disposal program that only addresses the primary copy leaves the organization holding data it has certified as destroyed. The scope of a disposal authorization has to include every known copy, and the mapping of where copies exist has to be maintained as part of the archive design rather than reconstructed at disposal time.

Personal data and the shortening pressure

One structural tension deserves attention. GxP obligations push retention longer. Privacy law pushes it shorter for personal data specifically.7 The resolution is usually not to delete the GxP record but to reduce the personal data it contains, through pseudonymization or through separating identifying information into a component with its own shorter retention rule. That design decision has to be made at ingest, because retrofitting it across a thirty-year archive is close to impossible. Clinical archives in particular should be designed on the assumption that the identifying key and the study data will have different lifespans.

A Practical Sequence for the First Twelve Months

Organizations that have never approached retention as an engineering problem often ask where to start. The sequence below assumes an existing retention schedule and an existing archive of some kind, which is the usual starting position.

1

Months 1 to 2: inventory the formats, not the systems

Produce a list of every distinct file format and data structure in the archive, with its volume, its source system, whether its specification is public, and what software is currently required to read it. Most organizations have never done this and find the result uncomfortable. It is the input to every subsequent decision.

2

Months 2 to 3: run a retrieval test against the oldest records you hold

Pick ten records from the oldest vintages and different formats and try to render them completely, including audit trail and signature. Time it. Document the failures. This single exercise usually reframes the conversation with leadership more effectively than any assessment document.

3

Months 3 to 5: fix the retention determination

Rebuild the schedule by record class with documented rationale, trigger events, and the location where each trigger is recorded. Include legal hold and privacy obligations. Have legal and quality sign the rationale, not only the periods.

4

Months 4 to 7: define the preservation approach per record class

Decide migration, emulation, rendition, or a combination for each class, and write down what each choice preserves and what it discards. Start with the classes that carry the longest obligations and the most proprietary formats, because those are where the exposure concentrates.

5

Months 6 to 9: fix the completeness gap

Check the audit trail purge settings across the validated estate. Establish how identity, signature manifestation, and context will be captured at archiving time. Change the standard decommissioning procedure so that a project cannot close without demonstrating a complete record package.

6

Months 9 to 12: stand up disposal and the testing cadence

Build the disposal process with its hold check, authorization step, and certificate of destruction, and run it on one low-risk record class to prove it works. Put the periodic retrievability test into the quality calendar with named owners and a defined evidence package.

Twelve months is realistic for the design and the first controls. The preservation actions themselves, particularly format migration across a large instrument estate, run for years and should be planned as continuing operational work rather than as a project with an end date.

Conclusion

The reason long-term retention keeps surprising organizations is that it is filed under policy when it belongs under engineering. A retention schedule tells you how long. It does not tell you whether the record will still render, whether the audit trail came along with it, whether anyone can prove it has not changed, or what happens on the day the period ends. Those are design questions, and they have to be answered while the source system is still running and the people who understand the data are still available. After decommissioning, the options narrow sharply and the recovery effort rises.

The organizations that handle this well are not the ones with the largest archive budgets. They are the ones that made a small number of deliberate choices early: they inventoried their formats, they decided per record class what they were preserving and what they were willing to lose, they treated the audit trail and the metadata as part of the record rather than as an attachment, they tested retrieval on a schedule and treated failures as deviations, and they built a disposal process so that the archive has an end as well as a beginning. None of that is expensive relative to the exposure it removes. All of it is much harder to do retroactively.

Sakara Digital works with pharma and biotech organizations designing archives that have to outlast the systems that feed them. If you are facing a decommissioning, a retention schedule that has drifted from what your systems can actually deliver, or a growing archive that nobody has tried to read in years, we are happy to have that conversation.

For Further Reading