Why Spatial Data Breaks the Sequencing Playbook

Most life sciences organizations built their research data infrastructure around bulk sequencing, then extended it for single-cell work. That architecture has a recognizable shape. Instruments write FASTQ files to a landing area. A pipeline converts them into count matrices. The matrices are small enough to hand to a biologist on a laptop. The FASTQ files go to cheap archival storage because they are the only thing that cannot be regenerated, and everything downstream can be rebuilt from them by rerunning the pipeline. It is a clean model and it worked for over a decade.

Spatial assays break that model in three specific ways.

The raw data is images, and images behave differently

In sequencing, the archival object is a text file that compresses well and never needs to be looked at by a human. In spatial work, the archival object is a multi-channel, multi-cycle microscopy image of a tissue section, and it is looked at constantly. Pathologists want to see it. Analysts want to overlay segmentation results on it. Reviewers want to check whether a region call is defensible. That means the archive is not write-once, read-never. It is write-once, read-occasionally-but-unpredictably, and that changes which storage tier is appropriate.

The scale is also different in kind. A single benchmark specimen used in the ASHLAR image stitching work was a colon section of roughly 24 mm by 14 mm, imaged as 609 tiles per cycle across nine cycles of tissue-based cyclic immunofluorescence to build a 28-plex image. Each cycle produced roughly 7 GB, and the complete dataset came to 61 GB from a single tissue section.11 That is one section, one specimen, one moderate plex level. A study with sixty sections at that scale is close to 4 TB of image data before anyone has analyzed anything.

The processed output is not small

The second break is that the processed output stops being laptop-sized. Sequencing-based spatial platforms have pushed toward finer capture resolution, and the feature count rises quickly. Visium HD produces feature-barcode matrices at 2 µm native resolution as well as binned versions at 8 µm and 16 µm, and the smallest bins can reach roughly 10.5 million per dataset.3 10x Genomics own guidance on custom binning works through a rasterization example where representing the full data as a dense array of 3,250 by 3,250 by 19,059 values, one byte each, would require 187 GB.4 Nobody stores it that way in practice, which is precisely the point: sparse-aware tooling is not optional at this resolution, and the tools your analysts already know may not be sparse-aware.

Practical guidance on planning for spatial storage makes the same point from the other direction. A single high-definition section can hold over 516,000 spatial features across more than 1,600 genes, roughly 838 million values, which comes to something like 3.35 to 6.7 GB per section if held densely in memory. Aggregating from 2 µm to 16 µm bins reduces the feature count by roughly 69 times, which is why binning choices have such a large effect on what hardware an analyst needs.5

The pipeline is not settled

The third break is the one that matters most and gets discussed least. In sequencing, the value of keeping only raw FASTQ files rests on a stable, well-benchmarked pipeline. You can throw away derived outputs because you are confident you can rebuild them the same way, or better, later. Spatial analysis does not have that stability yet. Segmentation methods are actively changing, benchmarking frameworks are still being proposed, and the choice of method materially changes what cells you think you saw.12 That means “we can always reprocess” is a promise your infrastructure has to be able to keep, because you will probably need to.

61 GB Total image data from a single 24 mm by 14 mm colon section imaged at 28-plex across nine cycles11
7 to 60 GB Archival output directory size for a single Xenium slide, depending on tissue area and panel1
10.5 million Approximate number of 2 µm bins in a single Visium HD capture area3

What Spatial Data Actually Weighs

Vague statements about spatial data being “huge” are useless for planning. The useful version is a set of ranges tied to what the experiment actually did. Vendors publish enough to build that, and the published figures are more helpful than the industry conversation suggests.

Published output sizes

10x Genomics publishes example output directory sizes for Xenium runs, broken out by tissue area and assay type. For RNA-only runs, a core needle biopsy of roughly 0.01 cm² produces about 0.3 GB, a full coronal mouse brain section of roughly 1 cm² produces 29 to 31 GB, and a maximum-area section of 2.35 cm² produces 67 to 72 GB. Adding 12 to 27 protein markers to the same tissue areas raises those to 0.5 to 0.6 GB, 47 to 64 GB, and 110 to 149 GB. Public Xenium datasets in the wild span roughly 3.5 GB to 144 GB.1 At the run level, a single slide with the full imageable area selected lands between 7 and 60 GB, and a four-slide run between roughly 28 and 240 GB.5

Independent guidance on the field as a whole puts contemporary spatial transcriptomics datasets at several hundred gigabytes to a few terabytes each, depending on how much imaging is retained alongside the expression data.6 Multimodal work sits at the upper end, because a single section profiled by more than one method carries the image burden of each.

What you changedEffect on stored volumeWho usually decides
Tissue area imaged (0.01 cm² to 2.35 cm²)Roughly 200-fold range on Xenium RNA-only outputs, from 0.3 GB to 72 GB1The scientist selecting the imageable region at run setup
Plex level and added modalitiesAdding 12 to 27 protein markers roughly doubles output size at the same tissue area1Assay design, usually settled months earlier
Capture or bin resolutionMoving from 16 µm to 2 µm bins raises feature count by roughly 69 times5Pipeline defaults, frequently never revisited
Number of sections per studyLinear, and the one number leadership can actually forecastStudy design and program roadmap
Whether raw images are retainedDetermines whether the archive is tens of gigabytes or tens of terabytes per sample1Nobody, in most organizations

The number that surprises people

The figures above describe the analysis outputs, not the instrument’s internal sensor data. That distinction matters enormously and it is the source of most of the confusion about spatial storage. 10x Genomics states plainly that Xenium internal sensor data runs to roughly tens of terabytes per sample, cannot be reanalyzed after onboard processing has run, and is therefore not practically useful to store.1

Be precise about which “raw” you mean. When someone says a spatial run produces terabytes, ask which files. Instrument sensor data is measured in tens of terabytes per sample and is explicitly not worth keeping, because it cannot be reprocessed. The archival image and transcript outputs, which are the ones you genuinely have a choice about, are measured in tens of gigabytes per slide. Confusing the two produces storage budgets that are wrong by three orders of magnitude in either direction.

Building your own estimate

A sizing model that works looks like this. Start with sections per year multiplied by gigabytes per slide to get an annual archive. Multiply by retention years for a cumulative archive. Then multiply by the number of copies you keep, a multiplier for derived files, and a headroom factor to get provisioned capacity.5 Applied to realistic program sizes, a small program running roughly 40 sections a year accumulates something like 1.4 to 12 TB over five years, which doubles with a second copy. A medium program at 160 sections a year lands at 5.6 to 48 TB, and a large program at 320 sections a year at 11 to 96 TB.5

Those are not frightening numbers on their own. What makes storage budgets go wrong is that storage is billed by terabyte-month, not by terabyte. A program adding 9.6 TB a year accrues roughly 1,730 terabyte-months over five years, which is a very different figure from the 48 TB that appears on the capacity slide.5 If your finance model multiplies final capacity by a per-terabyte rate, it will understate the five-year figure substantially.

The Retention Decision Nobody Makes Deliberately

Here is the sharp question that a spatial program has to answer, usually within its first year: once processed outputs exist, do you keep the raw images?

In most organizations this question is never asked. It is answered by default, and the default is “keep everything,” because nobody wants to be the person who deleted the images from the study that later needed reprocessing. That default is understandable and it is also the most expensive possible answer. It commits the organization to an archive that grows forever, on storage that has to stay reasonably accessible because images get looked at, with no criteria for ever removing anything.

Why the answer is not obvious

The case for keeping raw images is real. Segmentation is the step most likely to be revisited, and the field is actively improving it. Recent work has shown that inaccurate segmentation misattributes large numbers of transcripts to neighboring cells, and newer probabilistic approaches produce better boundaries and better recovery of hard-to-segment immune cell populations.13 Nature Methods has described the field as needing an open benchmarking framework built on large-scale annotated data, reproducible infrastructure, and transcript-informed metrics rather than morphology alone.12 If your two-year-old study used the segmentation method that was standard at the time and the field has since moved on, reprocessing is the only way to bring that study up to current standards. Without the images, you cannot.

The case against keeping everything is equally real. Image archives are the largest component of spatial storage by a wide margin, they grow monotonically, and a large fraction of them will never be reopened. Vendors themselves draw a line here: 10x recommends archiving decoded transcripts, provided in Zarr and Parquet, together with high-resolution morphology images in OME-TIFF, on the grounds that all other outputs are derived from those and can be rebuilt.1 That is a retention policy stated as a vendor recommendation, and it is more specific than what most organizations have written down themselves.

A framework for deciding

The decision is not binary across the whole program. It should be made per study, at study close, against four criteria.

CRITERION 1

Reprocessing likelihood

Is the analysis method for this data type still moving? Segmentation-dependent imaging assays are far more likely to be reprocessed than sequencing-based assays with settled preprocessing. High likelihood argues for keeping images.

CRITERION 2

Sample recoverability

Could the tissue be sectioned and rerun? Cell line and animal model work often can be. Rare patient material, a closed clinical study, or a discontinued program cannot. Irreplaceable samples argue strongly for keeping images.

CRITERION 3

Decision weight

Did this data support a go or no-go decision, a publication, a patent filing, or a regulatory interaction? If a result may need to be defended or reconstructed, the underlying image is part of the record and retention is not really optional.

CRITERION 4

Reuse potential

Is this tissue type, indication, or panel one your organization will keep working on? Reference-quality images from a well-characterized cohort earn their keep. One-off exploratory sections in an abandoned area usually do not.

Two of four pointing toward retention is a reasonable threshold for keeping images on deep archive. Three or four means keeping them somewhere warmer, because you will probably reopen them. Zero or one means keeping only the processed outputs and the metadata needed to explain what was done, and writing down that you made that choice on purpose.

The point is not the specific threshold. It is that the decision gets made at a defined moment, by a named person, using criteria the organization agreed on, and recorded. A spatial program with a written retention policy and an imperfect threshold is in a far better position than one with no policy and a full disk.

What makes reprocessing possible

Keeping images is necessary but not sufficient. Reprocessing requires the image, the metadata that describes how it was acquired, and enough record of the original processing to know what you are comparing against. Vendors have made this easier: Xenium Ranger allows segmentation to be rerun or replaced with a user’s own segmentation results using the archived raw outputs.1 Community guidance on sharing spatial data goes further, recommending that segmentation masks be shared alongside clear documentation of which images were segmented and mappings that link segmented objects back to the other analytical data.6

If you keep the images but lose the acquisition metadata, the panel definition, or the record of which pipeline version produced the outputs you published, you have kept the expensive part and thrown away the part that made it useful. The metadata is small. Treat it as the highest-value object in the archive.

The Analysis Bottleneck Is People, Not Processors

Every conversation about spatial infrastructure eventually turns into a conversation about compute. It is the wrong conversation, and buying compute will not fix the problem it is meant to fix.

The constraint is skilled analyst time

When Cell Systems asked working researchers what the main bottleneck is in deriving biological understanding from spatial transcriptomic profiling, the answer was not hardware. It was the need for deep biological expertise paired with genuine computational skill, applied from experimental design through to interpretation, and the observation that the bottleneck has moved from generating data to analyzing it.25 That is a labor constraint, not an infrastructure constraint.

The labor market makes it worse. Bioinformatics and data science roles are among the hardest categories to fill across pharma and biotech, driven by demand from large-scale genomic datasets, real-world evidence platforms, and computational drug design, all competing for people who can work fluently in both biology and code.22 Spatial work raises the bar further, because it requires image processing expertise on top of standard computational biology. The person who can run a single-cell pipeline competently is not automatically the person who can judge whether a segmentation result is trustworthy.

A useful test for whether your constraint is compute or people. Ask how many completed runs are sitting unanalyzed, and how long the oldest one has been waiting. Then ask what the average utilization of your analysis cluster is over the same period. If the queue is long and the cluster is idle, more compute will change nothing. That pattern is the normal one in spatial programs.

There is no settled pipeline to hand someone

The second half of the bottleneck is that spatial analysis has no universally accepted pipeline in the way that single-cell RNA sequencing broadly does. Each platform brings its own protocols and its own analysis path, which makes it hard to obtain uniformly preprocessed data across a study that used more than one technology. Benchmarking work is underway across several fronts: eleven sequencing-based spatial methods have been systematically compared using shared reference tissues,14 and methods for identifying spatially variable genes have been benchmarked systematically,15 but the results establish that method choice changes the answer rather than establishing which method to use in every case.

This has a direct infrastructure implication that is easy to miss. In a field with a settled pipeline, you can standardize, automate, and hire people to operate a known process. In a field without one, every study involves methodological judgment, which means senior analyst time, which means you cannot scale by adding junior staff or more nodes. Recognizing this changes the investment case. The right spend is often a smaller amount on compute and a larger amount on retaining two or three people who can actually make those judgments.

What compute you do need

None of this means compute is irrelevant. It means the compute requirement is more specific and usually more modest than expected. The dominant requirement is memory rather than processor cores, because the limiting operations involve holding large feature matrices and image arrays in memory. Guidance on planning for spatial workloads recommends provisioning generously for memory rather than oversizing core counts, and notes that a well-specified server with generous memory is often sufficient, with shared high-performance computing needed mainly for segmentation work or large cohorts.5

Frameworks have also adapted. SpatialData provides lazy representation of larger-than-memory data, so that analysts can work with datasets that exceed the memory of the machine in front of them without a cluster.10 That capability shifts a real portion of the compute requirement from hardware to software, and it is free.

What to do instead of buying a cluster. Provision one or two high-memory analysis machines, standardize on tooling that handles sparse and larger-than-memory data, reserve shared high-performance computing for segmentation and cohort-scale reprocessing, and put the difference into analyst headcount or a named external partner. That allocation reflects where the constraint actually sits.

Standards, Formats, and What Portability Actually Requires

Choosing a spatial platform is partly choosing how portable your data will be. That is a strategic decision made by scientists on scientific grounds, which is correct, but the data consequence should be visible to whoever owns the infrastructure.

Where the standards are

The imaging side has a real standard and a credible successor. OME-TIFF is the established format with broad support across bioimaging software.6 OME-NGFF, the next-generation file format published by the Open Microscopy Environment, uses Zarr to store metadata as JSON and binary data as individually addressable chunk files, which makes very large multidimensional images workable over object storage rather than requiring a whole file to be pulled down before anything can be viewed.7 The OME-Zarr specification describes itself as a cloud-optimized complement to OME-TIFF and HDF5, developed with international community support, and is currently at version 0.5.89

The spatial omics side is converging on the same foundation. SpatialData builds directly on OME-NGFF and Zarr to provide a unified multiplatform format with alignment to common coordinate systems, organized around a small set of primitive elements.10 Vendor output has moved in the same direction: Xenium writes morphology images as OME-TIFF and writes transcripts, cell feature matrices, segmentation masks, and analysis results as Zarr and Parquet.2 That is meaningfully better than the situation a few years ago.

Expression data has its own well-established path. Count matrices belong in INSDC-compliant repositories such as GEO, ArrayExpress, EGA, or dbGaP, following the established reporting guidelines for high-throughput sequencing experiments.6 That part of the problem is largely solved.

What open formats do not give you

Here is where organizations get a false sense of security. Writing your images as OME-TIFF and your matrices as Parquet means the bytes are readable by open tools. It does not mean the data is portable in any sense that matters scientifically.

Portability requires four things, and the format only covers the first.

1

Readable bytes in a documented container

OME-TIFF, OME-Zarr, Parquet, and HDF5 all deliver this. Any modern platform will produce at least some of its outputs in one of these. Necessary, and the easiest part.

2

Metadata rich enough to interpret the bytes

Channel definitions, panel composition, pixel size, acquisition settings, tissue annotation, and processing history. Community guidance for spatial data sharing builds on imaging and single-cell metadata standards adapted for spatial context, and this is where most real datasets fall short.6

3

A coordinate system that survives the move

Spatial data is only spatial if the relationship between the image, the segmentation mask, and the expression values is preserved. Frameworks that align modalities to common coordinate systems exist precisely because this alignment is fragile and platform-specific.10

4

An analysis path that does not depend on the vendor’s software

The strongest form of lock-in is not the file format. It is that the only practical route from raw output to interpretable result runs through one vendor’s application. Open toolboxes for multiplexed image analysis and community analysis frameworks reduce this dependence, but only if your analysts actually use them.2324

The practical question to ask a vendor

Rather than asking whether a platform supports open formats, which will always get a yes, ask three concrete questions. Can I export the raw images, the segmentation masks, and the expression matrix in documented open formats with full acquisition metadata attached, without using a licensed application? If I replace my segmentation with my own, can I regenerate every downstream output from the archived files alone? And if I stop paying for the analysis software next year, what specifically can I no longer do with the data I already have?

The answers to those three questions describe your actual lock-in. They are also questions that a scientific evaluation team will not think to ask, which is a good reason for someone from the technology side to be in the platform selection conversation early rather than after the purchase order.

Architecture: Tiers, Cloud, and On-Premises

Once the retention policy exists, the architecture follows from it fairly directly. The mistake is doing it in the other order, buying storage first and letting the policy be whatever the storage happens to permit.

Match tiers to access patterns, not to file age

Most storage tiering policies move data based on how old it is. That is the wrong variable for spatial data, because a two-year-old image from an active indication may be opened weekly while a three-month-old image from a dropped program will never be opened again. Tier by expected access, and set that expectation at study close using the same criteria that drove the retention decision.

TierWhat belongs hereAccess expectationNotes
WorkingActive study images, matrices, and intermediate outputsDaily, by analysts and scientistsFast, expensive, and should be small. Time-box how long a study stays here.
WarmProcessed outputs, segmentation masks, metadata, images from active programsWeekly to monthlyWhere most reopened data lives. Processed outputs are small enough that keeping them warm indefinitely is usually reasonable.
Cold archiveRaw images from closed studies flagged for retentionRarely, but must be recoverableDeep archive tiers price at roughly a dollar per terabyte-month, with standard restore in up to 12 hours and bulk restore in up to 48 hours.20
Public repositoryData supporting publicationsExternalRepository capacity is a real constraint and needs checking before submission, not after.6
DeletedInstrument sensor data, regenerable intermediates, images that failed the retention criteriaNoneRequires a written decision and a record of it. This tier is the one that makes the model work.

Read the fine print on deep archive before committing to it. Deep archive tiers typically carry a minimum storage duration, commonly 180 days, so data written and then deleted early is still billed for the full minimum. They also charge for retrieval by volume. Bulk retrieval is inexpensive per gigabyte but can take up to two days.20 That combination is fine for an archive you expect to open once every few years. It is a poor fit for images a pathologist wants to look at next week.

Where cloud makes sense and where it does not

For image data at this scale, the honest answer is that it depends on one variable more than any other: how often the data crosses the boundary between where it is stored and where it is analyzed.

Cloud is the better answer when analysis happens next to the data. Chunked, cloud-native formats like OME-Zarr exist precisely so that a viewer or an analysis job can read the region it needs from object storage without downloading the whole image.7 If your analysts work in cloud notebooks or a hosted workspace, storing images in object storage and computing beside them avoids moving the large objects at all. Consortium-scale efforts have gone this way for exactly this reason. The HuBMAP Data Portal, which as of June 2026 hosted 9,232 public datasets across 25 data types, 29 organ classes, and 498 donors, pairs its data holdings with collaborative workspaces that provide access to high-performance compute in place.17

Cloud is the worse answer when analysts work locally and pull images down repeatedly. Egress is billed by volume, with the first 100 GB per month free and subsequent tiers starting around nine cents per gigabyte before volume discounts.21 Ten analysts each pulling a few hundred gigabytes a month is a recurring bill that nobody forecast and that grows with adoption. It also gets worse as the program succeeds, which is the wrong shape for a research cost.

On-premises makes sense when you already have capable storage and a network fast enough that scientists do not think about it, when your analysis genuinely happens on local workstations, and when your volume growth is predictable enough to buy ahead. It stops making sense when growth is unpredictable, because the failure mode of on-premises is a hard capacity wall reached on a Friday afternoon.

A hybrid split that works

The pattern that holds up in practice separates the two data classes rather than choosing one location for everything. Keep processed outputs, metadata, segmentation masks, and anything an analyst touches regularly close to where the analysis happens, which for most groups now means cloud object storage adjacent to cloud compute. Put raw image archives on the cheapest durable tier available, wherever that is, and accept a slow restore, because by definition you have already decided this is data you rarely open. Keep one copy of anything genuinely irreplaceable in a second location under different administrative control.

The reason this works is that it puts the egress-sensitive data where it does not need to move and the volume-sensitive data where volume is cheap. It also makes the storage bill legible, because the large line item and the frequently accessed line item are no longer the same line item.

Planning Capacity When the Assay Roadmap Is Uncertain

The most common objection to any of this is that nobody knows what assays the research organization will run next year. That is true and it is not a reason to skip planning. It is a reason to plan in a form that tolerates being wrong.

Plan in bands, not point estimates

Ask the research leaders for three numbers rather than one: how many sections they expect to run next year if the current programs continue as planned, if one new program starts, and if a major initiative adopts spatial across the board. The published sizing ranges convert those directly into storage bands. A program at 40 sections a year, 160 sections a year, and 320 sections a year produces five-year cumulative archives in the ranges of 1.4 to 12 TB, 5.6 to 48 TB, and 11 to 96 TB respectively, before copies.5

Presenting a band rather than a number changes the conversation with finance in a useful way. It makes the width of the band the subject, and the width is driven by decisions the organization can actually control: tissue area, plex, resolution, and retention. That is a far more productive discussion than defending a single forecast that everyone knows is a guess.

Five practical steps

1

Measure what you already have

Before forecasting, get the actual gigabytes per section from your own last twenty runs, split into images and processed outputs. Vendor ranges are a starting point. Your own numbers reflect your tissue types, your panels, and your defaults, and they are usually different.

2

Write the retention policy before you buy capacity

Decide what gets kept, for how long, and who signs off. Without this, every capacity forecast is unbounded, because “keep everything forever” has no upper limit and therefore no meaningful budget.

3

Model in terabyte-months, not terabytes

Accumulating storage bills for every month it exists. A program adding 9.6 TB a year accrues roughly 1,730 terabyte-months across five years rather than the 48 TB that appears on a capacity chart.5 Use the accumulating figure in the budget.

4

Buy in increments that match your uncertainty

If the roadmap is unclear beyond twelve months, do not commit to three years of on-premises capacity. Elastic capacity is worth paying a premium for exactly when the forecast is weak, and worth abandoning once the volume becomes predictable.

5

Review quarterly against actuals

Compare forecast sections to actual sections and forecast gigabytes to actual gigabytes. Two quarters of data will tell you more about your growth rate than any vendor estimate, and will catch a resolution or plex change that quietly doubled your per-section volume.

Watch for the changes that do not look like infrastructure changes

The changes that break a capacity plan rarely arrive labeled as infrastructure decisions. A group switches to a higher-plex panel. A protocol moves to finer binning because a reviewer asked for more resolution. A study adds a second modality on the same sections. Sequencing depth guidance shifts, as it has for formalin-fixed paraffin-embedded Visium work where 100,000 to 120,000 reads per spot is now often the recommendation.16 Each of these is a scientific decision made for good scientific reasons, and each of them changes the storage curve.

The fix is not to give infrastructure a veto over assay design, which would be both unwelcome and wrong. It is to establish a lightweight notification: when a group changes plex, resolution, modality count, or tissue area as a standing default, someone tells the person who owns the storage forecast. That is a five-minute conversation that prevents a surprise a year later.

When the Data Supports a Regulatory Submission

Everything above describes research infrastructure, where you get to make your own retention decisions on scientific and financial grounds. There is a boundary, and crossing it changes the rules. It is worth knowing where the boundary sits even if your spatial work is nowhere near it today, because programs move.

What changes at the boundary

Once spatial data contributes to a regulatory submission, whether as a biomarker result supporting a development decision, a companion diagnostic claim, or an imaging-derived endpoint, the retention question stops being yours alone to answer. FDA guidance on clinical trial imaging endpoints sets process standards covering imaging acquisition, display, archiving, and interpretation, with the stated aim of ensuring that imaging data are obtained in line with the protocol, that quality is maintained within and across sites, and that a verifiable record of the imaging process exists.18 Note that archiving is named explicitly in that list. Retention becomes part of the process standard rather than a budget decision.

The practical consequences are familiar to anyone who has worked in a validated environment. Provenance has to be demonstrable, which means the chain from acquired image through processing to reported result has to be reconstructable from records rather than from the memory of the analyst who did it. Processing has to be reproducible, which means the pipeline version, parameters, and reference data used to generate a submitted result all have to be captured and retained alongside the result. And deletion stops being a routine storage decision, because you cannot remove something you may be required to produce.

The signal to watch for. The boundary is usually crossed quietly. A biomarker that started as exploratory becomes a stratification factor. An imaging readout that was descriptive becomes supportive of an endpoint. Nobody sends a note saying the data governance model has changed. The practical control is to ask, at every study close, whether the result is intended to support or may reasonably be expected to support a submission. If the answer is yes or maybe, the study’s data leaves the research retention policy and enters the regulated one.

Even outside the regulated setting, obligations exist

Research data is not obligation-free simply because it is not regulated. Publicly funded work carries data management and sharing expectations, with data supporting a publication expected to be shared by the time of publication and other data by the end of the project, using established repositories where they exist.19 Journals and consortia add their own requirements. For spatial data specifically, community guidance recommends sharing microscopy images, segmentation masks with documentation of what was segmented, and count matrices deposited in appropriate repositories.6

These are worth checking early rather than at submission, because repository capacity limits are real. Free-tier deposits at general-purpose repositories are typically capped well below what a spatial dataset weighs, and larger allocations may require an institutional arrangement.6 Discovering that at manuscript submission is an avoidable problem.

Conclusion

The pattern we see repeatedly is that spatial biology gets adopted as a scientific decision and shows up as an infrastructure problem eighteen months later, at which point the choices that determined the size of the problem have already been made. The tissue areas, the plex levels, the binning defaults, and above all the unspoken decision to keep everything were all set by people who were making good scientific judgments and had no reason to think they were also setting a storage commitment. There is no villain in this story. There is just a conversation that did not happen at the right time.

The work that pays back is unglamorous and mostly not technical. Measure what your own runs actually produce rather than relying on ranges. Write down what gets kept and who decides. Match storage tiers to how the data is actually opened rather than how old it is. Be honest that the analysis queue is limited by the number of people who can make methodological judgments, and resist the temptation to solve a staffing constraint by buying hardware. Ask platform vendors the specific portability questions before purchase rather than after. And know where the regulatory boundary sits so that the day a study crosses it, someone notices.

Sakara Digital works with pharma and biotech organizations building the data infrastructure behind emerging research technologies, including spatial and multi-omics programs where the scientific ambition has outrun the storage and analysis plan. If you are standing up a spatial capability, or already have one and want an independent view of what it will require over the next three years, we are happy to have that conversation.

For Further Reading