In This Article
- Executive Summary
- Why Spatial Data Breaks the Sequencing Playbook
- What Spatial Data Actually Weighs
- The Retention Decision Nobody Makes Deliberately
- The Analysis Bottleneck Is People, Not Processors
- Standards, Formats, and What Portability Actually Requires
- Architecture: Tiers, Cloud, and On-Premises
- Planning Capacity When the Assay Roadmap Is Uncertain
- When the Data Supports a Regulatory Submission
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Spatial transcriptomics and multiplexed imaging have moved from specialist method to standard part of the translational toolkit in pharma and biotech research. The science is genuinely good. The infrastructure conversation almost never happens before the first instrument is installed. By the time someone in IT gets asked about storage, the group has already run forty sections, filled a shared drive, and started emailing about slow file transfers.
The core problem is that spatial assays produce two different kinds of data with two completely different retention profiles, and most organizations treat them as one thing. There are large, mostly static image files that are expensive to keep and impossible to regenerate. There are much smaller processed outputs that are cheap to keep and can be rebuilt if the images survive. Deciding which of those you keep, for how long, and on what storage tier is the single decision that determines whether a spatial program becomes a manageable line item or an open-ended commitment. Almost nobody makes that decision on purpose.
This article gives the infrastructure view behind the science: what spatial data actually weighs and which experimental choices drive the number, a framework for the raw image retention decision, an honest account of why the real constraint is analyst time rather than compute, what format standards do and do not give you in terms of portability, and how to plan storage capacity when the assay roadmap for next year is genuinely unknown. It closes with the boundary you should watch for, which is the point where spatial data starts supporting a regulatory submission and the retention and provenance expectations change.
Why Spatial Data Breaks the Sequencing Playbook
Most life sciences organizations built their research data infrastructure around bulk sequencing, then extended it for single-cell work. That architecture has a recognizable shape. Instruments write FASTQ files to a landing area. A pipeline converts them into count matrices. The matrices are small enough to hand to a biologist on a laptop. The FASTQ files go to cheap archival storage because they are the only thing that cannot be regenerated, and everything downstream can be rebuilt from them by rerunning the pipeline. It is a clean model and it worked for over a decade.
Spatial assays break that model in three specific ways.
The raw data is images, and images behave differently
In sequencing, the archival object is a text file that compresses well and never needs to be looked at by a human. In spatial work, the archival object is a multi-channel, multi-cycle microscopy image of a tissue section, and it is looked at constantly. Pathologists want to see it. Analysts want to overlay segmentation results on it. Reviewers want to check whether a region call is defensible. That means the archive is not write-once, read-never. It is write-once, read-occasionally-but-unpredictably, and that changes which storage tier is appropriate.
The scale is also different in kind. A single benchmark specimen used in the ASHLAR image stitching work was a colon section of roughly 24 mm by 14 mm, imaged as 609 tiles per cycle across nine cycles of tissue-based cyclic immunofluorescence to build a 28-plex image. Each cycle produced roughly 7 GB, and the complete dataset came to 61 GB from a single tissue section.11 That is one section, one specimen, one moderate plex level. A study with sixty sections at that scale is close to 4 TB of image data before anyone has analyzed anything.
The processed output is not small
The second break is that the processed output stops being laptop-sized. Sequencing-based spatial platforms have pushed toward finer capture resolution, and the feature count rises quickly. Visium HD produces feature-barcode matrices at 2 µm native resolution as well as binned versions at 8 µm and 16 µm, and the smallest bins can reach roughly 10.5 million per dataset.3 10x Genomics own guidance on custom binning works through a rasterization example where representing the full data as a dense array of 3,250 by 3,250 by 19,059 values, one byte each, would require 187 GB.4 Nobody stores it that way in practice, which is precisely the point: sparse-aware tooling is not optional at this resolution, and the tools your analysts already know may not be sparse-aware.
Practical guidance on planning for spatial storage makes the same point from the other direction. A single high-definition section can hold over 516,000 spatial features across more than 1,600 genes, roughly 838 million values, which comes to something like 3.35 to 6.7 GB per section if held densely in memory. Aggregating from 2 µm to 16 µm bins reduces the feature count by roughly 69 times, which is why binning choices have such a large effect on what hardware an analyst needs.5
The pipeline is not settled
The third break is the one that matters most and gets discussed least. In sequencing, the value of keeping only raw FASTQ files rests on a stable, well-benchmarked pipeline. You can throw away derived outputs because you are confident you can rebuild them the same way, or better, later. Spatial analysis does not have that stability yet. Segmentation methods are actively changing, benchmarking frameworks are still being proposed, and the choice of method materially changes what cells you think you saw.12 That means “we can always reprocess” is a promise your infrastructure has to be able to keep, because you will probably need to.
What Spatial Data Actually Weighs
Vague statements about spatial data being “huge” are useless for planning. The useful version is a set of ranges tied to what the experiment actually did. Vendors publish enough to build that, and the published figures are more helpful than the industry conversation suggests.
Published output sizes
10x Genomics publishes example output directory sizes for Xenium runs, broken out by tissue area and assay type. For RNA-only runs, a core needle biopsy of roughly 0.01 cm² produces about 0.3 GB, a full coronal mouse brain section of roughly 1 cm² produces 29 to 31 GB, and a maximum-area section of 2.35 cm² produces 67 to 72 GB. Adding 12 to 27 protein markers to the same tissue areas raises those to 0.5 to 0.6 GB, 47 to 64 GB, and 110 to 149 GB. Public Xenium datasets in the wild span roughly 3.5 GB to 144 GB.1 At the run level, a single slide with the full imageable area selected lands between 7 and 60 GB, and a four-slide run between roughly 28 and 240 GB.5
Independent guidance on the field as a whole puts contemporary spatial transcriptomics datasets at several hundred gigabytes to a few terabytes each, depending on how much imaging is retained alongside the expression data.6 Multimodal work sits at the upper end, because a single section profiled by more than one method carries the image burden of each.
| What you changed | Effect on stored volume | Who usually decides |
|---|---|---|
| Tissue area imaged (0.01 cm² to 2.35 cm²) | Roughly 200-fold range on Xenium RNA-only outputs, from 0.3 GB to 72 GB1 | The scientist selecting the imageable region at run setup |
| Plex level and added modalities | Adding 12 to 27 protein markers roughly doubles output size at the same tissue area1 | Assay design, usually settled months earlier |
| Capture or bin resolution | Moving from 16 µm to 2 µm bins raises feature count by roughly 69 times5 | Pipeline defaults, frequently never revisited |
| Number of sections per study | Linear, and the one number leadership can actually forecast | Study design and program roadmap |
| Whether raw images are retained | Determines whether the archive is tens of gigabytes or tens of terabytes per sample1 | Nobody, in most organizations |
The number that surprises people
The figures above describe the analysis outputs, not the instrument’s internal sensor data. That distinction matters enormously and it is the source of most of the confusion about spatial storage. 10x Genomics states plainly that Xenium internal sensor data runs to roughly tens of terabytes per sample, cannot be reanalyzed after onboard processing has run, and is therefore not practically useful to store.1
Be precise about which “raw” you mean. When someone says a spatial run produces terabytes, ask which files. Instrument sensor data is measured in tens of terabytes per sample and is explicitly not worth keeping, because it cannot be reprocessed. The archival image and transcript outputs, which are the ones you genuinely have a choice about, are measured in tens of gigabytes per slide. Confusing the two produces storage budgets that are wrong by three orders of magnitude in either direction.
Building your own estimate
A sizing model that works looks like this. Start with sections per year multiplied by gigabytes per slide to get an annual archive. Multiply by retention years for a cumulative archive. Then multiply by the number of copies you keep, a multiplier for derived files, and a headroom factor to get provisioned capacity.5 Applied to realistic program sizes, a small program running roughly 40 sections a year accumulates something like 1.4 to 12 TB over five years, which doubles with a second copy. A medium program at 160 sections a year lands at 5.6 to 48 TB, and a large program at 320 sections a year at 11 to 96 TB.5
Those are not frightening numbers on their own. What makes storage budgets go wrong is that storage is billed by terabyte-month, not by terabyte. A program adding 9.6 TB a year accrues roughly 1,730 terabyte-months over five years, which is a very different figure from the 48 TB that appears on the capacity slide.5 If your finance model multiplies final capacity by a per-terabyte rate, it will understate the five-year figure substantially.
The Retention Decision Nobody Makes Deliberately
Here is the sharp question that a spatial program has to answer, usually within its first year: once processed outputs exist, do you keep the raw images?
In most organizations this question is never asked. It is answered by default, and the default is “keep everything,” because nobody wants to be the person who deleted the images from the study that later needed reprocessing. That default is understandable and it is also the most expensive possible answer. It commits the organization to an archive that grows forever, on storage that has to stay reasonably accessible because images get looked at, with no criteria for ever removing anything.
Why the answer is not obvious
The case for keeping raw images is real. Segmentation is the step most likely to be revisited, and the field is actively improving it. Recent work has shown that inaccurate segmentation misattributes large numbers of transcripts to neighboring cells, and newer probabilistic approaches produce better boundaries and better recovery of hard-to-segment immune cell populations.13 Nature Methods has described the field as needing an open benchmarking framework built on large-scale annotated data, reproducible infrastructure, and transcript-informed metrics rather than morphology alone.12 If your two-year-old study used the segmentation method that was standard at the time and the field has since moved on, reprocessing is the only way to bring that study up to current standards. Without the images, you cannot.
The case against keeping everything is equally real. Image archives are the largest component of spatial storage by a wide margin, they grow monotonically, and a large fraction of them will never be reopened. Vendors themselves draw a line here: 10x recommends archiving decoded transcripts, provided in Zarr and Parquet, together with high-resolution morphology images in OME-TIFF, on the grounds that all other outputs are derived from those and can be rebuilt.1 That is a retention policy stated as a vendor recommendation, and it is more specific than what most organizations have written down themselves.
A framework for deciding
The decision is not binary across the whole program. It should be made per study, at study close, against four criteria.
Reprocessing likelihood
Is the analysis method for this data type still moving? Segmentation-dependent imaging assays are far more likely to be reprocessed than sequencing-based assays with settled preprocessing. High likelihood argues for keeping images.
Sample recoverability
Could the tissue be sectioned and rerun? Cell line and animal model work often can be. Rare patient material, a closed clinical study, or a discontinued program cannot. Irreplaceable samples argue strongly for keeping images.
Decision weight
Did this data support a go or no-go decision, a publication, a patent filing, or a regulatory interaction? If a result may need to be defended or reconstructed, the underlying image is part of the record and retention is not really optional.
Reuse potential
Is this tissue type, indication, or panel one your organization will keep working on? Reference-quality images from a well-characterized cohort earn their keep. One-off exploratory sections in an abandoned area usually do not.
Two of four pointing toward retention is a reasonable threshold for keeping images on deep archive. Three or four means keeping them somewhere warmer, because you will probably reopen them. Zero or one means keeping only the processed outputs and the metadata needed to explain what was done, and writing down that you made that choice on purpose.
The point is not the specific threshold. It is that the decision gets made at a defined moment, by a named person, using criteria the organization agreed on, and recorded. A spatial program with a written retention policy and an imperfect threshold is in a far better position than one with no policy and a full disk.
What makes reprocessing possible
Keeping images is necessary but not sufficient. Reprocessing requires the image, the metadata that describes how it was acquired, and enough record of the original processing to know what you are comparing against. Vendors have made this easier: Xenium Ranger allows segmentation to be rerun or replaced with a user’s own segmentation results using the archived raw outputs.1 Community guidance on sharing spatial data goes further, recommending that segmentation masks be shared alongside clear documentation of which images were segmented and mappings that link segmented objects back to the other analytical data.6
If you keep the images but lose the acquisition metadata, the panel definition, or the record of which pipeline version produced the outputs you published, you have kept the expensive part and thrown away the part that made it useful. The metadata is small. Treat it as the highest-value object in the archive.
The Analysis Bottleneck Is People, Not Processors
Every conversation about spatial infrastructure eventually turns into a conversation about compute. It is the wrong conversation, and buying compute will not fix the problem it is meant to fix.
The constraint is skilled analyst time
When Cell Systems asked working researchers what the main bottleneck is in deriving biological understanding from spatial transcriptomic profiling, the answer was not hardware. It was the need for deep biological expertise paired with genuine computational skill, applied from experimental design through to interpretation, and the observation that the bottleneck has moved from generating data to analyzing it.25 That is a labor constraint, not an infrastructure constraint.
The labor market makes it worse. Bioinformatics and data science roles are among the hardest categories to fill across pharma and biotech, driven by demand from large-scale genomic datasets, real-world evidence platforms, and computational drug design, all competing for people who can work fluently in both biology and code.22 Spatial work raises the bar further, because it requires image processing expertise on top of standard computational biology. The person who can run a single-cell pipeline competently is not automatically the person who can judge whether a segmentation result is trustworthy.
A useful test for whether your constraint is compute or people. Ask how many completed runs are sitting unanalyzed, and how long the oldest one has been waiting. Then ask what the average utilization of your analysis cluster is over the same period. If the queue is long and the cluster is idle, more compute will change nothing. That pattern is the normal one in spatial programs.
There is no settled pipeline to hand someone
The second half of the bottleneck is that spatial analysis has no universally accepted pipeline in the way that single-cell RNA sequencing broadly does. Each platform brings its own protocols and its own analysis path, which makes it hard to obtain uniformly preprocessed data across a study that used more than one technology. Benchmarking work is underway across several fronts: eleven sequencing-based spatial methods have been systematically compared using shared reference tissues,14 and methods for identifying spatially variable genes have been benchmarked systematically,15 but the results establish that method choice changes the answer rather than establishing which method to use in every case.
This has a direct infrastructure implication that is easy to miss. In a field with a settled pipeline, you can standardize, automate, and hire people to operate a known process. In a field without one, every study involves methodological judgment, which means senior analyst time, which means you cannot scale by adding junior staff or more nodes. Recognizing this changes the investment case. The right spend is often a smaller amount on compute and a larger amount on retaining two or three people who can actually make those judgments.
What compute you do need
None of this means compute is irrelevant. It means the compute requirement is more specific and usually more modest than expected. The dominant requirement is memory rather than processor cores, because the limiting operations involve holding large feature matrices and image arrays in memory. Guidance on planning for spatial workloads recommends provisioning generously for memory rather than oversizing core counts, and notes that a well-specified server with generous memory is often sufficient, with shared high-performance computing needed mainly for segmentation work or large cohorts.5
Frameworks have also adapted. SpatialData provides lazy representation of larger-than-memory data, so that analysts can work with datasets that exceed the memory of the machine in front of them without a cluster.10 That capability shifts a real portion of the compute requirement from hardware to software, and it is free.
What to do instead of buying a cluster. Provision one or two high-memory analysis machines, standardize on tooling that handles sparse and larger-than-memory data, reserve shared high-performance computing for segmentation and cohort-scale reprocessing, and put the difference into analyst headcount or a named external partner. That allocation reflects where the constraint actually sits.
Standards, Formats, and What Portability Actually Requires
Choosing a spatial platform is partly choosing how portable your data will be. That is a strategic decision made by scientists on scientific grounds, which is correct, but the data consequence should be visible to whoever owns the infrastructure.
Where the standards are
The imaging side has a real standard and a credible successor. OME-TIFF is the established format with broad support across bioimaging software.6 OME-NGFF, the next-generation file format published by the Open Microscopy Environment, uses Zarr to store metadata as JSON and binary data as individually addressable chunk files, which makes very large multidimensional images workable over object storage rather than requiring a whole file to be pulled down before anything can be viewed.7 The OME-Zarr specification describes itself as a cloud-optimized complement to OME-TIFF and HDF5, developed with international community support, and is currently at version 0.5.89
The spatial omics side is converging on the same foundation. SpatialData builds directly on OME-NGFF and Zarr to provide a unified multiplatform format with alignment to common coordinate systems, organized around a small set of primitive elements.10 Vendor output has moved in the same direction: Xenium writes morphology images as OME-TIFF and writes transcripts, cell feature matrices, segmentation masks, and analysis results as Zarr and Parquet.2 That is meaningfully better than the situation a few years ago.
Expression data has its own well-established path. Count matrices belong in INSDC-compliant repositories such as GEO, ArrayExpress, EGA, or dbGaP, following the established reporting guidelines for high-throughput sequencing experiments.6 That part of the problem is largely solved.
What open formats do not give you
Here is where organizations get a false sense of security. Writing your images as OME-TIFF and your matrices as Parquet means the bytes are readable by open tools. It does not mean the data is portable in any sense that matters scientifically.
Portability requires four things, and the format only covers the first.
Readable bytes in a documented container
OME-TIFF, OME-Zarr, Parquet, and HDF5 all deliver this. Any modern platform will produce at least some of its outputs in one of these. Necessary, and the easiest part.
Metadata rich enough to interpret the bytes
Channel definitions, panel composition, pixel size, acquisition settings, tissue annotation, and processing history. Community guidance for spatial data sharing builds on imaging and single-cell metadata standards adapted for spatial context, and this is where most real datasets fall short.6
A coordinate system that survives the move
Spatial data is only spatial if the relationship between the image, the segmentation mask, and the expression values is preserved. Frameworks that align modalities to common coordinate systems exist precisely because this alignment is fragile and platform-specific.10
An analysis path that does not depend on the vendor’s software
The strongest form of lock-in is not the file format. It is that the only practical route from raw output to interpretable result runs through one vendor’s application. Open toolboxes for multiplexed image analysis and community analysis frameworks reduce this dependence, but only if your analysts actually use them.2324
The practical question to ask a vendor
Rather than asking whether a platform supports open formats, which will always get a yes, ask three concrete questions. Can I export the raw images, the segmentation masks, and the expression matrix in documented open formats with full acquisition metadata attached, without using a licensed application? If I replace my segmentation with my own, can I regenerate every downstream output from the archived files alone? And if I stop paying for the analysis software next year, what specifically can I no longer do with the data I already have?
The answers to those three questions describe your actual lock-in. They are also questions that a scientific evaluation team will not think to ask, which is a good reason for someone from the technology side to be in the platform selection conversation early rather than after the purchase order.
Architecture: Tiers, Cloud, and On-Premises
Once the retention policy exists, the architecture follows from it fairly directly. The mistake is doing it in the other order, buying storage first and letting the policy be whatever the storage happens to permit.
Match tiers to access patterns, not to file age
Most storage tiering policies move data based on how old it is. That is the wrong variable for spatial data, because a two-year-old image from an active indication may be opened weekly while a three-month-old image from a dropped program will never be opened again. Tier by expected access, and set that expectation at study close using the same criteria that drove the retention decision.
| Tier | What belongs here | Access expectation | Notes |
|---|---|---|---|
| Working | Active study images, matrices, and intermediate outputs | Daily, by analysts and scientists | Fast, expensive, and should be small. Time-box how long a study stays here. |
| Warm | Processed outputs, segmentation masks, metadata, images from active programs | Weekly to monthly | Where most reopened data lives. Processed outputs are small enough that keeping them warm indefinitely is usually reasonable. |
| Cold archive | Raw images from closed studies flagged for retention | Rarely, but must be recoverable | Deep archive tiers price at roughly a dollar per terabyte-month, with standard restore in up to 12 hours and bulk restore in up to 48 hours.20 |
| Public repository | Data supporting publications | External | Repository capacity is a real constraint and needs checking before submission, not after.6 |
| Deleted | Instrument sensor data, regenerable intermediates, images that failed the retention criteria | None | Requires a written decision and a record of it. This tier is the one that makes the model work. |
Read the fine print on deep archive before committing to it. Deep archive tiers typically carry a minimum storage duration, commonly 180 days, so data written and then deleted early is still billed for the full minimum. They also charge for retrieval by volume. Bulk retrieval is inexpensive per gigabyte but can take up to two days.20 That combination is fine for an archive you expect to open once every few years. It is a poor fit for images a pathologist wants to look at next week.
Where cloud makes sense and where it does not
For image data at this scale, the honest answer is that it depends on one variable more than any other: how often the data crosses the boundary between where it is stored and where it is analyzed.
Cloud is the better answer when analysis happens next to the data. Chunked, cloud-native formats like OME-Zarr exist precisely so that a viewer or an analysis job can read the region it needs from object storage without downloading the whole image.7 If your analysts work in cloud notebooks or a hosted workspace, storing images in object storage and computing beside them avoids moving the large objects at all. Consortium-scale efforts have gone this way for exactly this reason. The HuBMAP Data Portal, which as of June 2026 hosted 9,232 public datasets across 25 data types, 29 organ classes, and 498 donors, pairs its data holdings with collaborative workspaces that provide access to high-performance compute in place.17
Cloud is the worse answer when analysts work locally and pull images down repeatedly. Egress is billed by volume, with the first 100 GB per month free and subsequent tiers starting around nine cents per gigabyte before volume discounts.21 Ten analysts each pulling a few hundred gigabytes a month is a recurring bill that nobody forecast and that grows with adoption. It also gets worse as the program succeeds, which is the wrong shape for a research cost.
On-premises makes sense when you already have capable storage and a network fast enough that scientists do not think about it, when your analysis genuinely happens on local workstations, and when your volume growth is predictable enough to buy ahead. It stops making sense when growth is unpredictable, because the failure mode of on-premises is a hard capacity wall reached on a Friday afternoon.
A hybrid split that works
The pattern that holds up in practice separates the two data classes rather than choosing one location for everything. Keep processed outputs, metadata, segmentation masks, and anything an analyst touches regularly close to where the analysis happens, which for most groups now means cloud object storage adjacent to cloud compute. Put raw image archives on the cheapest durable tier available, wherever that is, and accept a slow restore, because by definition you have already decided this is data you rarely open. Keep one copy of anything genuinely irreplaceable in a second location under different administrative control.
The reason this works is that it puts the egress-sensitive data where it does not need to move and the volume-sensitive data where volume is cheap. It also makes the storage bill legible, because the large line item and the frequently accessed line item are no longer the same line item.
Planning Capacity When the Assay Roadmap Is Uncertain
The most common objection to any of this is that nobody knows what assays the research organization will run next year. That is true and it is not a reason to skip planning. It is a reason to plan in a form that tolerates being wrong.
Plan in bands, not point estimates
Ask the research leaders for three numbers rather than one: how many sections they expect to run next year if the current programs continue as planned, if one new program starts, and if a major initiative adopts spatial across the board. The published sizing ranges convert those directly into storage bands. A program at 40 sections a year, 160 sections a year, and 320 sections a year produces five-year cumulative archives in the ranges of 1.4 to 12 TB, 5.6 to 48 TB, and 11 to 96 TB respectively, before copies.5
Presenting a band rather than a number changes the conversation with finance in a useful way. It makes the width of the band the subject, and the width is driven by decisions the organization can actually control: tissue area, plex, resolution, and retention. That is a far more productive discussion than defending a single forecast that everyone knows is a guess.
Five practical steps
Measure what you already have
Before forecasting, get the actual gigabytes per section from your own last twenty runs, split into images and processed outputs. Vendor ranges are a starting point. Your own numbers reflect your tissue types, your panels, and your defaults, and they are usually different.
Write the retention policy before you buy capacity
Decide what gets kept, for how long, and who signs off. Without this, every capacity forecast is unbounded, because “keep everything forever” has no upper limit and therefore no meaningful budget.
Model in terabyte-months, not terabytes
Accumulating storage bills for every month it exists. A program adding 9.6 TB a year accrues roughly 1,730 terabyte-months across five years rather than the 48 TB that appears on a capacity chart.5 Use the accumulating figure in the budget.
Buy in increments that match your uncertainty
If the roadmap is unclear beyond twelve months, do not commit to three years of on-premises capacity. Elastic capacity is worth paying a premium for exactly when the forecast is weak, and worth abandoning once the volume becomes predictable.
Review quarterly against actuals
Compare forecast sections to actual sections and forecast gigabytes to actual gigabytes. Two quarters of data will tell you more about your growth rate than any vendor estimate, and will catch a resolution or plex change that quietly doubled your per-section volume.
Watch for the changes that do not look like infrastructure changes
The changes that break a capacity plan rarely arrive labeled as infrastructure decisions. A group switches to a higher-plex panel. A protocol moves to finer binning because a reviewer asked for more resolution. A study adds a second modality on the same sections. Sequencing depth guidance shifts, as it has for formalin-fixed paraffin-embedded Visium work where 100,000 to 120,000 reads per spot is now often the recommendation.16 Each of these is a scientific decision made for good scientific reasons, and each of them changes the storage curve.
The fix is not to give infrastructure a veto over assay design, which would be both unwelcome and wrong. It is to establish a lightweight notification: when a group changes plex, resolution, modality count, or tissue area as a standing default, someone tells the person who owns the storage forecast. That is a five-minute conversation that prevents a surprise a year later.
When the Data Supports a Regulatory Submission
Everything above describes research infrastructure, where you get to make your own retention decisions on scientific and financial grounds. There is a boundary, and crossing it changes the rules. It is worth knowing where the boundary sits even if your spatial work is nowhere near it today, because programs move.
What changes at the boundary
Once spatial data contributes to a regulatory submission, whether as a biomarker result supporting a development decision, a companion diagnostic claim, or an imaging-derived endpoint, the retention question stops being yours alone to answer. FDA guidance on clinical trial imaging endpoints sets process standards covering imaging acquisition, display, archiving, and interpretation, with the stated aim of ensuring that imaging data are obtained in line with the protocol, that quality is maintained within and across sites, and that a verifiable record of the imaging process exists.18 Note that archiving is named explicitly in that list. Retention becomes part of the process standard rather than a budget decision.
The practical consequences are familiar to anyone who has worked in a validated environment. Provenance has to be demonstrable, which means the chain from acquired image through processing to reported result has to be reconstructable from records rather than from the memory of the analyst who did it. Processing has to be reproducible, which means the pipeline version, parameters, and reference data used to generate a submitted result all have to be captured and retained alongside the result. And deletion stops being a routine storage decision, because you cannot remove something you may be required to produce.
The signal to watch for. The boundary is usually crossed quietly. A biomarker that started as exploratory becomes a stratification factor. An imaging readout that was descriptive becomes supportive of an endpoint. Nobody sends a note saying the data governance model has changed. The practical control is to ask, at every study close, whether the result is intended to support or may reasonably be expected to support a submission. If the answer is yes or maybe, the study’s data leaves the research retention policy and enters the regulated one.
Even outside the regulated setting, obligations exist
Research data is not obligation-free simply because it is not regulated. Publicly funded work carries data management and sharing expectations, with data supporting a publication expected to be shared by the time of publication and other data by the end of the project, using established repositories where they exist.19 Journals and consortia add their own requirements. For spatial data specifically, community guidance recommends sharing microscopy images, segmentation masks with documentation of what was segmented, and count matrices deposited in appropriate repositories.6
These are worth checking early rather than at submission, because repository capacity limits are real. Free-tier deposits at general-purpose repositories are typically capped well below what a spatial dataset weighs, and larger allocations may require an institutional arrangement.6 Discovering that at manuscript submission is an avoidable problem.
Conclusion
The pattern we see repeatedly is that spatial biology gets adopted as a scientific decision and shows up as an infrastructure problem eighteen months later, at which point the choices that determined the size of the problem have already been made. The tissue areas, the plex levels, the binning defaults, and above all the unspoken decision to keep everything were all set by people who were making good scientific judgments and had no reason to think they were also setting a storage commitment. There is no villain in this story. There is just a conversation that did not happen at the right time.
The work that pays back is unglamorous and mostly not technical. Measure what your own runs actually produce rather than relying on ranges. Write down what gets kept and who decides. Match storage tiers to how the data is actually opened rather than how old it is. Be honest that the analysis queue is limited by the number of people who can make methodological judgments, and resist the temptation to solve a staffing constraint by buying hardware. Ask platform vendors the specific portability questions before purchase rather than after. And know where the regulatory boundary sits so that the day a study crosses it, someone notices.
Sakara Digital works with pharma and biotech organizations building the data infrastructure behind emerging research technologies, including spatial and multi-omics programs where the scientific ambition has outrun the storage and analysis plan. If you are standing up a spatial capability, or already have one and want an independent view of what it will require over the next three years, we are happy to have that conversation.
For Further Reading
For Further Reading
- Multi-Omics Data Integration: Technology Architecture for Precision Medicine and Biomarker Discovery
- FAIR Data Principles for Life Sciences: Building Findable, Accessible, Interoperable, and Reusable Data Assets
- AI-Ready Data Infrastructure: What Pharma Needs to Build Now
- Data Lakehouse Architecture for Pharma: Unifying Clinical, Manufacturing, and Commercial Data
- Cloud Data Residency for Global Pharma: EU, US, and APAC Requirements
References & Sources
- 10x Genomics. “Archiving Xenium Data.” Xenium Onboard Analysis documentation. https://www.10xgenomics.com/support/software/xenium-onboard-analysis/latest/analysis/xoa-output-archive-data
- 10x Genomics. “Xenium Output Files at a Glance.” Xenium Onboard Analysis documentation. https://www.10xgenomics.com/support/software/xenium-onboard-analysis/latest/analysis/xoa-output-at-a-glance
- 10x Genomics. “Understanding Space Ranger Outputs.” Space Ranger documentation. https://www.10xgenomics.com/support/software/space-ranger/latest/analysis/outputs/output-overview
- 10x Genomics. “Nuclei Segmentation and Custom Binning of Visium HD Gene Expression Data.” Analysis Guides. https://www.10xgenomics.com/analysis-guides/segmentation-visium-hd
- Lab Manager. “How Big Is Spatial Biology Data? Planning Storage and Compute.” https://www.labmanager.com/how-big-is-spatial-biology-data-planning-storage-and-compute-35748
- Jackson H. W. and Pachter L. “A standard for sharing spatial transcriptomics data.” Cell Genomics, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC10435375/
- Moore J. et al. “OME-NGFF: a next-generation file format for expanding bioimaging data-access strategies.” Nature Methods, 2021. https://www.nature.com/articles/s41592-021-01326-w
- Moore J. et al. “OME-Zarr: a cloud-optimized bioimaging file format with international community support.” Histochemistry and Cell Biology, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC9980008/
- Open Microscopy Environment. “OME-Zarr specification, version 0.5.” https://ngff.openmicroscopy.org/0.5/index.html
- Marconato L. et al. “SpatialData: an open and universal data framework for spatial omics.” Nature Methods, 2025. https://www.nature.com/articles/s41592-024-02212-x
- Muhlich J. L. et al. “Stitching and registering highly multiplexed whole-slide images of tissues and tumors using ASHLAR.” Bioinformatics, 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC9525007/
- “Confronting the challenge of cell segmentation in spatial transcriptomics.” Nature Methods, 2025. https://www.nature.com/articles/s41592-025-02717-z
- “Cell simulation as cell segmentation.” Nature Methods, 2025. https://www.nature.com/articles/s41592-025-02697-0
- “Systematic comparison of sequencing-based spatial transcriptomic methods.” Nature Methods, 2024. https://www.nature.com/articles/s41592-024-02325-3
- “Systematic benchmarking of computational methods to identify spatially variable genes.” Genome Biology, 2025. https://link.springer.com/article/10.1186/s13059-025-03731-2
- Grases D., Porta-Pardo E. et al. “A practical guide to spatial transcriptomics: lessons from over 1000 samples.” Trends in Biotechnology, 2025. https://pubmed.ncbi.nlm.nih.gov/40975650/
- Turner J. A. et al. “HuBMAP Data Portal: A Resource for Multi-Modal Spatial and Single-Cell Data of Healthy Human Tissues.” arXiv, 2025. https://arxiv.org/abs/2511.05708
- U.S. Food and Drug Administration. “Clinical Trial Imaging Endpoint Process Standards: Guidance for Industry.” April 2018. https://www.fda.gov/files/drugs/published/Clinical-Trial-Imaging-Endpoint-Process-Standards-Guidance-for-Industry.pdf
- National Institutes of Health. “Writing a Data Management and Sharing Plan.” Grants and Funding policy guidance. https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/dms/writing-dms-plan
- Amazon Web Services. “Secure archive storage: Amazon S3 Glacier storage classes.” https://aws.amazon.com/s3/storage-classes/glacier/
- Amazon Web Services. “Amazon S3 pricing.” https://aws.amazon.com/s3/pricing/
- BioSpace. “Bioinformatics Roles in Increasing Demand, Critical to Industry, Personalized Medicine.” https://www.biospace.com/job-trends/bioinformatics-roles-in-increasing-demand-critical-to-industry-personalized-medicine
- Meyer-Bender M. et al. “Spatialproteomics: an interoperable toolbox for analyzing highly multiplexed fluorescence image data.” Nature Methods, 2026. https://www.nature.com/articles/s41592-026-03155-1
- Bioconductor. “Importing data.” Orchestrating Spatial Transcriptomics Analysis with Bioconductor. https://bioconductor.org/books/release/OSTA/pages/bkg-importing-data.html
- Cell Systems. “What is the main bottleneck in deriving biological understanding from spatial transcriptomic profiling?” 2025. https://www.cell.com/cell-systems/fulltext/S2405-4712(25)00033-X








Your perspective matters—join the conversation.