In This Article
- Executive Summary
- The Gap an In-Silico Prediction Falls Into
- Why the GAIP Work Lands in Nonclinical First
- Four Questions Before a Regulator Can Rely on a Prediction
- The Precedents Already Exist: ICH M7, E14/S7B, and Now M15
- Endpoint by Endpoint: Where Prediction Is Good and Where It Is Not
- Why the Applicability Domain Is the Hardest Part
- Data Provenance, Not Just Data Quality
- What to Document Now, Before GAIP Is Final
- Conclusion
- For Further Reading
- References & Sources
Executive Summary
Good Laboratory Practice was written for physical studies. Read 21 CFR Part 58 and you find test articles, test systems, specimens, dosing, and archives of tissue slides. A computational prediction of hERG liability or hepatic clearance has none of those things. It is not a GLP study. It is also not nothing, because sponsors are already using those predictions to kill compounds, pick a lead, set a starting dose range, and decide which animal study to run. That is the gap.
Black Mesa Technology received an ARPA-H award, worth up to $2 million and announced in January 2026, to build a Good AI Practice framework under the DeepMesa project. It sits inside CATALYST, the ARPA-H program whose stated ambition is a future where first-in-human approval can rest on in-silico safety data. Most commentary has treated GAIP as a general AI governance story. It is not. It is a nonclinical data integrity story, and the nonclinical space is where computational prediction is most heavily used and least explicitly governed.
This article takes GAIP out of the abstract and applies it to in-silico ADME-Tox. It walks the regulatory gap in Part 58, sets out the four things a regulator has to be satisfied about before relying on a prediction, works through the endpoints where prediction is genuinely good and where it is not, and gives sponsors a documentation practice they can start now that will map onto whatever GAIP finalizes.
The Gap an In-Silico Prediction Falls Into
Start with the regulation itself. Under 21 CFR 58.3, a nonclinical laboratory study means an in-vivo or in-vitro experiment in which test articles are studied prospectively in test systems under laboratory conditions to determine their safety.5 Every operative noun in that sentence assumes physical matter. A test article is a substance. A test system is an animal, plant, microorganism, or subparts thereof. A specimen is material derived from a test system for examination or analysis.
Now try to apply the rest of Part 58 to a machine learning model that predicts whether a molecule blocks the hERG potassium channel. There is a study director requirement, and you can appoint one. There is a quality assurance unit requirement, and you can staff one. Then it gets difficult. Part 58 requires characterization of the test article, including strength, purity, and stability. It requires records of the test system, including housing and feed. It requires retention of specimens in archives. It requires equipment to be adequately inspected, cleaned, and maintained. None of that translates cleanly.
The honest answer is that the regulation was written in 1978 in response to laboratory fraud involving physical studies, and it has never been rewritten for computation. FDA proposed a substantial revision in August 2016 that would have moved Part 58 to a full quality systems structure, referred to in the proposal as a GLP Quality System, and would have widened the scope of what falls under the rule.6 The comment period, originally set to close in November 2016, was extended to 21 January 2017. A decade later there is still no final rule. Part 58 as it stands today is the 1978 structure with minor amendments.
The practical position a sponsor is in. An in-silico ADME-Tox result is not a GLP study, so it cannot be presented as one. It is also not informal, because it changed a decision that affected which compound went into an animal, and eventually into a person. There is no regulation that tells you what evidence to keep about it. There is also no regulation that says you may keep nothing.
Why “just do not call it GLP” is not a durable answer
Many organizations resolve the tension by declaring computational predictions to be non-GLP, research-grade, or for internal decision support only. That works right up to the point where the prediction becomes part of the story you tell a regulator. It becomes part of that story more often than people expect. A sponsor that used a computational tox screen to deprioritize an analog series will be asked, at some point, why that series was dropped. A sponsor that used a physiologically based model to set a first-in-human starting dose has put a computational result directly on the path to a person receiving a drug.
The pressure is increasing rather than easing. FDA’s April 2025 roadmap for reducing animal testing in preclinical safety studies names computational modeling alongside organ-on-chip and advanced in-vitro assays as the methods intended to take over work animals do today, starting with monoclonal antibodies.10 A 2026 review in Clinical Pharmacology in Drug Development works through what the roadmap will and will not achieve, and returns repeatedly to the same practical question: the alternatives need to be qualified for their intended use before they can carry regulatory weight.22 Qualification requires evidence. Evidence requires records. Records require rules about what to record.
Why the GAIP Work Lands in Nonclinical First
Black Mesa Technology, based in Bedford, Massachusetts, holds an ARPA-H award listed as DeepMesa: Good AI Practice (GAIP) for assured AI-driven ADME-Tox. The award page records funding of up to $2 million, a start date of September 30, 2025, a principal investigator of Charles Fracchia, and management by the Health Science Futures office.1 The company announced the award publicly on January 20, 2026, describing a framework and associated technical methods intended to help organizations use AI in a way that upholds data integrity under 21 CFR, modeled on existing Good Practice standards such as Good Laboratory Practice and Good Manufacturing Practice.4
The detail that matters most is where the award sits. CATALYST stands for Computational ADME-Tox and Physiology Analysis for Safer Therapeutics. ARPA-H describes its ambition plainly: a future in which approval to begin first-in-human clinical trials can be based on in-silico safety data, developed in collaboration with regulators.2 The program funds three technical areas covering data discovery and deep learning methods for drug safety models, living systems tools for model development, and in-silico models of human physiology.3
So the rest of CATALYST is building the predictive models. DeepMesa is building the evidence discipline that would let anyone rely on them. Those are not the same problem, and the second one is the harder one. A model that predicts hepatotoxicity at 85% accuracy is a scientific achievement. A model whose prediction for a specific compound on a specific date can be regenerated, explained, bounded, and defended three years later in front of a reviewer is a regulatory artifact. Very little of the field currently produces the second thing.
What is actually known, and what is not, as of August 2026
Sakara Digital has written about the GAIP work before, and it is worth being precise about the state of play rather than repeating the announcement. What is confirmed is the award, the ceiling, the start date, the program placement, and the stated intent to publish a framework with associated technical methods. What is not confirmed, from ARPA-H or from Black Mesa’s own communications, is a published framework document, a public comment process, or a date. Neither the ARPA-H award listing nor the company’s news page carries a GAIP deliverable beyond the January 2026 announcement.
That matters for planning. A sponsor waiting for a finished GAIP document before changing anything will be waiting through at least one more development cycle, and quite possibly through a first-in-human decision made on partly computational evidence. The practical move is to build the record now using vocabulary that already exists in finalized regulatory text, which is exactly what the last section of this article sets out.
A note on scope. GAIP is not a nonclinical-only framework. It is written to cover AI use across drug discovery and development. But the funding sits inside an ADME-Tox program, the technical partners are building ADME-Tox models, and the regulatory gap is widest in nonclinical. If GAIP proves itself anywhere first, it will be here.
Four Questions Before a Regulator Can Rely on a Prediction
Strip away the framework language and a reviewer looking at a computational result is asking four things. They are not new questions. They are the same questions a reviewer asks about any study, translated into a setting where there is no bench, no animal, and no specimen jar.
Is this compound one the model can speak to?
The applicability domain. A model trained on kinase inhibitors has an opinion about a macrocyclic peptide, and that opinion is worthless. The reviewer needs to know whether the query compound sits inside the chemical space the model was built from, by what method that was determined, and what the threshold was.
Where did the training data come from?
Provenance and quality. Which assay produced each label, under which protocol, in which laboratory, in which units, and how were conflicting values resolved. A model is a compressed restatement of its training set. If the training set cannot be described, the model cannot be defended.
Can this exact prediction be reproduced?
Reproducibility of a specific result, not of the method in general. Model version, weights, descriptor calculation code, software environment, random seeds, and input structure representation all have to be recoverable, or the number in the report is an assertion rather than a result.
Was the model qualified for this specific purpose?
Documented fitness for the stated use. A model qualified to rank compounds within a series is not thereby qualified to support a safety margin. The claim has to be written down before the result is generated, and the evidence has to match the claim.
Those four questions are why “the model is 90% accurate” is not an answer. Accuracy on what set, measured how, for compounds resembling which chemistry, and used for what decision. A reviewer who accepts a bare accuracy figure has accepted a claim with no boundary on it, and reviewers do not do that.
The asymmetry that makes this uncomfortable
There is a genuine asymmetry between a physical study and a computational one, and it runs in an unexpected direction. A physical study is hard to repeat and easy to document. You cannot rerun a 90-day rat study, but you can describe exactly what was done, and the raw data sits in an archive. A computational study is easy to repeat in principle and hard to document in practice. You could rerun the prediction in seconds, but only if you still have the model version, the code, the environment, and the input in the exact form it was submitted, and most organizations do not keep all four.
Six months after a model is retrained on new data, the original prediction may be genuinely unrecoverable. Nobody committed fraud. Nobody was careless in the ordinary sense. The record simply was not designed to survive. That is the failure mode GAIP exists to close, and it is the failure mode most exposed in nonclinical work, because nonclinical models are retrained constantly as new assay data arrives.
The Precedents Already Exist: ICH M7, E14/S7B, and Now M15
It is easy to describe in-silico regulatory acceptance as unprecedented. It is not. There are three places where regulators have already worked out how to rely on a computational result, and each one teaches something specific.
ICH M7: the first time a prediction replaced a test
ICH M7 governs the assessment and control of DNA-reactive mutagenic impurities. In the absence of experimental data, it accepts an in-silico assessment of bacterial mutagenicity using two complementary (Q)SAR methodologies, one expert rule-based and one statistical.11 If neither flags a structural alert, the impurity is treated as being of no mutagenic concern and no further testing is recommended. That is a genuine substitution of computation for a laboratory test, in a binding regulatory guideline.
The design of that acceptance is instructive. M7 does not accept a single model. It requires two, built on different principles, so that a systematic weakness in one is less likely to be shared by the other. It also expects expert review of the predictions where the situation is ambiguous, and a published analysis of that practice found that expert review to be the step that resolves the cases where the two methods disagree or where an alert is present but mitigated by structural context.12 Redundancy plus documented human judgment, not a single model score.
ICH E14 and S7B: qualification tied to a context of use
The E14/S7B question-and-answer document addresses nonclinical and clinical evaluation of QT interval prolongation and proarrhythmic potential, and it explicitly contemplates in-silico modeling as part of that assessment.15 The language it uses is the language that matters: an appropriately qualified proarrhythmia risk prediction model may be used according to its context of use. Not qualified in general. Qualified for a use, and used only for that use.
The Comprehensive in-vitro Proarrhythmia Assay initiative produced the worked example. Its mechanistic in-silico model takes multi-ion-channel pharmacology measured in vitro and simulates a human ventricular cardiomyocyte to classify torsade risk, and it was assessed against a defined training and validation compound set with a pre-specified metric before anyone proposed using it.16 That is what qualification looks like in practice: a stated purpose, a defined compound set, a metric agreed in advance, and a boundary on the claim.
ICH M15: a finalized framework, and what it does not cover
ICH M15, General Principles for Model-Informed Drug Development, was adopted at Step 4 on January 29, 2026, after Step 2 consultation in November 2024.7 It is the most useful document available to a sponsor thinking about computational evidence, and it is finalized rather than draft, which is rare in this area.
M15 defines a six-element assessment structure that a sponsor is expected to fill in and share with regulators: question of interest, context of use, model influence, consequence of wrong decision, model risk, and model impact. Model risk is derived by combining model influence with consequence of wrong decision, and it is what determines how much model evaluation is required. The guideline is direct about this: when model outcomes are the sole source supporting a decision, model influence should be considered high.
| M15 element | What it means for an in-silico ADME-Tox result |
|---|---|
| Question of interest | The decision the prediction is meant to inform, stated explicitly. “Does this compound carry sufficient hERG liability to require a dedicated in-vitro assay before candidate selection?” is a question of interest. “Predict hERG” is not. |
| Context of use | The role and scope of the model, plus a description of the data it was built from and any other evidence contributing to the answer. This is where the applicability domain claim belongs. |
| Model influence | How much weight the prediction carries relative to everything else. A prediction used to triage 4,000 virtual compounds carries low influence. A prediction used in place of an assay carries high influence. |
| Consequence of wrong decision | Severity and likelihood of harm if the prediction is wrong. A false negative on proarrhythmia liability that carries through to first-in-human is a different consequence from a false positive that drops a backup series. |
| Model risk | Influence combined with consequence. Determines the depth of evaluation required. High risk means external validation with independent data may be essential rather than encouraged. |
| Model impact | How far the proposed approach departs from existing regulatory standards, or from expectations where no standard exists. For in-silico ADME-Tox, where no standard exists, this rating is almost always at least medium. |
M15 also sets out model evaluation in three parts: verification that the code and equations are correct, validation comparing the model against data and prior knowledge, and applicability assessment covering whether the data and model are adequate for the specific intended use. It names overfitting explicitly as a method-specific issue to consider for artificial intelligence and machine learning models. It expects a Model Analysis Plan written before the analysis and a Model Analysis Report documenting what was actually done, with any departures justified. It expects supporting files, including the data used and the relevant coding scripts, to be submitted or available for regulatory review.
Where M15 stops. M15 governs modeling evidence that a sponsor submits to a regulator to answer a stated question. It is a submission framework. It says almost nothing about the day-to-day discipline required of a model that sits inside a discovery pipeline generating thousands of predictions a week, most of which never reach a submission and some of which quietly shape what does. That operating layer is the gap GAIP is aimed at, and it is the layer where nonclinical computational work actually lives.
The complementary regulatory documents fill in the AI-specific expectations. FDA’s January 2025 draft guidance on the use of artificial intelligence to support regulatory decision-making sets out a risk-based credibility assessment tied to context of use, and is explicit that it does not address AI used in drug discovery or for operational efficiency that does not affect the reliability of nonclinical or clinical study results.8 That carve-out is precisely the boundary a sponsor has to reason about, because a discovery-stage tox model that changes which compound enters a nonclinical study is arguably on both sides of it. EMA’s September 2024 reflection paper takes a similar risk-based position and singles out nonclinical uses that inform the design of first-in-human studies as warranting development and testing proportionate to that role.9
Endpoint by Endpoint: Where Prediction Is Good and Where It Is Not
Governance conversations tend to treat in-silico ADME-Tox as one thing. It is not. Predictive performance varies enormously by endpoint, and any framework that applies the same evidence expectation to a permeability prediction and a hepatotoxicity prediction will be wrong in both directions. Here is an honest picture, drawn from published evaluations.
| Endpoint | Reported performance | Honest reading |
|---|---|---|
| hERG liability | Structure-based classifiers reached a maximum AUC of 0.86 on 8,337 curated ChEMBL compounds; ligand-based machine learning and graph neural network approaches report AUC values in the high 0.80s to low 0.90s [17][23] | Genuinely useful. Large public dataset, well-defined binary endpoint, decades of medicinal chemistry knowledge behind it. Still bounded: the benchmark restricted itself to molecular weights of 200 to 600 daltons, and performance varied with the choice of protein conformation and docking software. |
| CYP inhibition | Strong on held-out internal test sets, weak on genuinely external ones. A 2025 study evaluating against known approved-drug inhibitors reported recall of 0.27 for CYP2B6 and 0.60 for CYP2C8 [18] | The clearest illustration of the applicability domain problem in the literature. The authors attributed the CYP2B6 result to structural difference between training and external compounds, with mean Tanimoto similarity of 0.39. The model was not broken. It was being asked about chemistry it had never seen. |
| Hepatotoxicity (DILI) | Random forest models averaged 0.631 accuracy in cross-validation and a multilayer perceptron reached a Matthews correlation coefficient of 0.245; on an external set of candidates that had failed clinically for hepatotoxicity, both correctly flagged 90.9% [19] | Weak discrimination overall, decent sensitivity on known-bad compounds. Useful as one input to a weight-of-evidence view. Not usable on its own to clear a compound, and a negative prediction carries very little information. |
| Permeability (Caco-2) | Consensus regression models built on regional and global random forests produced RMSE of 0.43 to 0.51 log units across validation sets [21] | Reasonable for ranking and triage. Half a log unit of error is tolerable when you are sorting a library and intolerable when you are supporting a specific exposure claim. Fit depends entirely on what the number is being used for. |
| Human clearance | Random forest models trained on 1,340 compounds with human intravenous data gave a geometric mean fold error of 3.3 on a quasi-prospective test set of 343 compounds from which structurally similar training compounds had been removed; renally cleared compounds reached 2.3 [20] | The most sobering number in this table, and the most honest study design. Threefold error on clearance is not a basis for a dose decision. The same authors note that performance improves substantially when near neighbors are present in training, which is the whole point. |
Read the right-hand column together and a pattern emerges that should shape every governance decision in this area. Predictive performance is highest exactly where you need it least and lowest exactly where you need it most. A well-populated endpoint on familiar chemistry predicts well, and you probably already had an assay for it. A novel scaffold on a complex endpoint predicts poorly, and that is the compound you were hoping the model would tell you about.
Watch the evaluation design, not the headline number. The clearance study reported threefold error because the authors deliberately stripped structurally similar compounds out of the training set before testing. Most published performance figures do not do that. A random train-test split on a dataset full of analog series measures how well the model interpolates within series it already knows, which is not the question anyone actually has.
Why the Applicability Domain Is the Hardest Part
Of the four questions in the earlier framework, the applicability domain is the one that most often has no answer at all in current practice, and it is the one where the regulatory expectation is clearest.
The OECD principles for the validation of (Q)SAR models, set out in Guidance Document 69, state that a model intended for regulatory use should be associated with five things, while noting the principles are not themselves criteria for regulatory acceptance: a defined endpoint, an unambiguous algorithm, a defined domain of applicability, appropriate measures of goodness-of-fit, robustness and predictivity, and where possible a mechanistic interpretation. The measures principle covers goodness-of-fit, robustness and predictivity together, and dropping robustness is a common misquotation. Appropriate measures of goodness of fit and predictivity, and a mechanistic interpretation where possible.13 The third principle is the one that separates a research model from a regulatory one. It is not enough to report how well a model performs. You have to say for which chemicals that performance claim holds.
There is no single correct method, which is why you must state yours
Published comparisons of applicability domain methods find several families in use, based on descriptor range, geometric distance, leverage, probability density, or similarity to nearest neighbors in the training set, and they do not agree with each other on which compounds fall inside.14 A compound can be inside the domain by a range-based method and outside it by a distance-based one. That is not a reason to skip the question. It is a reason the method and the threshold must be recorded alongside the prediction, because otherwise “in domain” is an unfalsifiable claim.
Conformal prediction has become a practical option here, and the clearance study cited above used it to attach confidence intervals to individual predictions and to assess model applicability.20 The attraction is that it produces a per-prediction statement rather than a global one, which is what a reviewer actually wants. It does not remove the need to state the method. It makes the statement more useful.
The governance rule this leads to
A prediction generated for a compound outside the model’s stated applicability domain is not a weak result to be treated with caution. It is not a result at all, and it should not be recorded as one. Systems that emit a number for every input, with no domain flag, quietly convert “we do not know” into “the value is 4.2” and then that number travels into a slide, a decision, and eventually a memory of a decision.
The fix is unglamorous and entirely achievable: every prediction record carries a domain flag, the method used to set it, and the threshold. Out-of-domain predictions are stored and reported as out of domain, not suppressed and not silently passed through.
Applicability drift over the life of a program
There is a second-order problem that nonclinical teams meet constantly. A model qualified at lead identification, when the series looked one way, is still in use at candidate selection after eighteen months of medicinal chemistry has pushed the series somewhere else. Nobody re-checked the domain. The compounds moved and the model did not.
ICH M15 has language for this. It distinguishes appropriateness of the proposed approach, which is about whether the strategy suits the question, from applicability assessment, which is about whether the data and model are adequate for each intended use.7 Each intended use. Not the program. A practical control is to re-run the domain assessment at each stage gate and record the result, which takes minutes and produces an evidence trail that would otherwise not exist.
Data Provenance, Not Just Data Quality
Data quality gets most of the attention in AI readiness work. Provenance is the harder and more consequential problem for nonclinical models, because a model is a compressed restatement of its training data and a reviewer who cannot see the training data is being asked to accept the compression on trust.
What “the model was trained on public data” actually means
Public bioactivity databases are indispensable and they are heterogeneous by construction. Values for the same compound and the same nominal endpoint are aggregated across laboratories, protocols, incubation times, cell lines, and reporting conventions accumulated over decades. Reviews of in-silico ADME modeling report that models built on curated public data can reach predictive performance comparable to models built on in-house or commercial datasets, and caution that uncurated duplicate records inflate apparent predictability.23 There is a real trade-off between quantity and consistency, and where a team lands on it is a scientific decision that has to be written down.
The provenance record that makes a nonclinical model defensible is not complicated. It is a description, per training set, of where the labels came from, what assay generated them, what filters and deduplication rules were applied, how conflicting measurements were resolved, what was excluded and why, and a fixed identifier for the resulting dataset so that the exact set can be retrieved later. Sponsors already do the equivalent for clinical datasets. Very few do it for the training data behind a tox model.
ALCOA+ applied to a training set
The data integrity principles that govern GxP records translate onto a training set more naturally than people expect, and the translation is a useful exercise for a quality team meeting computational work for the first time.
Who produced each label
Every training record traces to a source assay and a source publication or internal study, not to an aggregate table with no upstream reference.
The measured value survives
The raw measured value and its units are retained alongside any transformed or binarized version used for training. A threshold applied to turn a continuous value into a class label is a processing decision, and it is recorded as one.
When the set was frozen
The training set carries a date and a fixed identifier. A model trained on “ChEMBL” is not reproducible. A model trained on a named, hashed extract taken on a stated date is.
Conflicts resolved by a stated rule
Where sources disagree, the rule that picked a value is written down and applied consistently, rather than resolved case by case by whoever built the set.
The retraining trap. Nonclinical models get retrained whenever a meaningful batch of new assay data arrives, which in an active program is frequently. If retraining overwrites the deployed model without preserving the previous version, every prediction made before the retrain becomes irreproducible. The decision those predictions supported is still in the program. The evidence for it is gone. This is the single most common and most avoidable failure in computational nonclinical work, and it is a versioning problem rather than a science problem.
What to Document Now, Before GAIP Is Final
Everything above points at the same practical conclusion. A sponsor using computational tox and ADME today does not need to wait for GAIP to be published to start building the record GAIP will ask for. The elements are already specified in finalized text: ICH M15 supplies the vocabulary, OECD Guidance Document 69 supplies the model-quality principles, and ALCOA+ supplies the data expectations. What is missing is not the content. It is the habit of writing it down for computational work the way it already gets written down for physical work.
Here is a sequence that takes a mid-size nonclinical group a few weeks rather than a few quarters.
Inventory the predictions that changed a decision
Not every model. Every decision. Walk back through the last two years of program decisions and list the ones where a computational ADME or tox result contributed: a series dropped, a lead picked, an assay skipped, a dose range narrowed. For most groups this list is shorter than expected and more consequential than expected. It is also the list a reviewer would eventually construct, so it is better to have it first.
Write a one-page context-of-use statement per model
Use the M15 elements directly: question of interest, context of use, model influence, consequence of wrong decision, model risk, model impact, each with a short justification. One page. If a model is used for three different questions, it gets three statements. The discipline of writing “this model is not qualified to support a safety margin” is worth more than the page itself.
Freeze and register model versions
Every deployed model gets an entry recording the version identifier, the training set identifier and hash, the descriptor calculation code version, the software environment, the training date, and the person accountable for it. Superseded versions are retained, not overwritten. This is the control that makes historical predictions recoverable, and without it nothing else in this list works.
Attach a domain flag to every prediction
Record the applicability domain method, the threshold, and whether the specific query compound fell inside or outside. Out-of-domain results are stored and reported as out of domain. Where the tooling supports it, add a per-prediction uncertainty estimate through conformal prediction or an equivalent method, because a reviewer asks about the individual compound rather than the average.
Record the human decision alongside the prediction
Who reviewed the result, what other evidence was on the table, what was decided, and why. ICH M7 works because it pairs two independent computational methods with documented expert review. The same pattern applies here. A prediction with no recorded human judgment attached is an output. A prediction with one is a decision record.
Run a re-prediction test twice a year
Pick five predictions from six or twelve months ago and try to regenerate them exactly from the archive. This takes an afternoon and it is the only honest test of whether the record works. Teams that run it the first time usually fail it, and the failure is specific and fixable: a missing environment specification, an unversioned descriptor library, a training set nobody snapshotted.
None of this is wasted if GAIP lands somewhere different. Model versioning, training set provenance, applicability domain flags, context-of-use statements, and decision records are the common content of every framework in this space: M15, the FDA credibility approach, the EMA reflection paper, and the OECD validation principles. A sponsor that builds these records is not betting on a specific framework. It is building the evidence any of them would ask for, and it is doing so while the programs that need it are still running.
The organizational question underneath the technical one
One question decides how hard this is: who owns computational nonclinical models. In many organizations they sit with a computational chemistry or data science group that reports into research, operates outside the quality system by design, and is measured on scientific throughput. That placement is defensible for models used purely to triage virtual libraries. It becomes untenable the moment a model output contributes to a decision that reaches a regulator.
The answer is not to pull computational chemistry into the quality system wholesale, which would slow the science down for no safety benefit. It is to define a boundary. Below the boundary, models are research tools with light documentation. Above it, where a model’s output informs a decision with regulatory weight, the model carries a context-of-use statement, a version record, a domain policy, and a decision trail. Drawing that boundary explicitly is a governance decision that takes a couple of meetings and prevents a category of problem that is very expensive to fix later.
Conclusion
The ARPA-H CATALYST program is built on a proposition that would have seemed implausible ten years ago: that first-in-human approval could one day rest on in-silico safety data. Whether that arrives on the timeline ARPA-H hopes for is genuinely uncertain, and the endpoint performance data in this article suggests the science has some distance to cover, particularly on clearance and on anything involving novel chemistry. But the direction is not uncertain, and neither is the FDA position on reducing animal testing that runs alongside it. Computational nonclinical evidence is going to carry more weight, not less.
The work Black Mesa is doing under DeepMesa matters because it addresses the part of that transition nobody else is funded to solve. Building a better hepatotoxicity model is a scientific problem with a scientific community working on it. Building a discipline under which a specific prediction, made on a specific compound, on a specific date, can be regenerated, bounded, and defended is an evidence problem, and evidence problems are solved by writing rules down and following them. GLP did exactly that for physical studies in 1978. The nonclinical space needs the same thing for computation, and Part 58 in its current form is not going to provide it.
The practical point for a sponsor is that waiting is the wrong move. GAIP is in development and has no published framework document. ICH M15 is final today and gives you the vocabulary. OECD Guidance Document 69 is decades old and gives you the model quality principles. ALCOA+ gives you the data expectations. A nonclinical group that starts recording context of use, model versions, training set provenance, applicability domain flags, and human decisions is building the record that any of these frameworks will ask for, and is building it while the programs that need it are still open rather than reconstructing it under audit pressure two years later.
Sakara Digital works with pharma and biotech organizations bringing computational methods into regulated nonclinical work and figuring out what evidence has to exist around them. If you are using in-silico ADME-Tox to inform real decisions and want an independent view on what to document and where the boundary should sit, we are happy to have that conversation.
For Further Reading
For Further Reading
- Black Mesa GAIP Progress Report: What the Latest Draft Reveals About Industry Direction
- The AI Model Risk Assessment for Pharma: A Structured Checklist
- Building an AI Model Registry: What to Track and Why
- Anchored Knowledge: How to Design AI Systems That Know When They Don’t Know
- Data Lineage in Regulated Industries: From Source to Submission
- Building AI as Scientific Infrastructure: Platform Strategies for Drug Discovery
References & Sources
- ARPA-H. “DeepMesa: Good AI Practice (GAIP) for assured AI-driven ADME-Tox.” Award listing, Black Mesa Technology, Inc., award start September 30, 2025. https://arpa-h.gov/explore-funding/awards/3596
- ARPA-H. “CATALYST: Computational ADME-Tox and Physiology Analysis for Safer Therapeutics.” Program page. https://arpa-h.gov/explore-funding/programs/catalyst
- ARPA-H. “CATALYST program to fast-track safer medicines from lab to patients.” News release. https://arpa-h.gov/news-and-events/catalyst-program-fast-track-safer-medicines-lab-patients
- Black Mesa. “Black Mesa awarded ARPA-H funding to develop a ‘Good AI Practice’ (GAIP) framework.” Press release, January 20, 2026. https://www.einpresswire.com/article/874731311/black-mesa-awarded-arpa-h-funding-to-develop-a-good-ai-practice-gaip-framework
- Electronic Code of Federal Regulations. “21 CFR 58.3: Definitions.” Good Laboratory Practice for Nonclinical Laboratory Studies. https://www.ecfr.gov/current/title-21/section-58.3
- U.S. Food and Drug Administration. “Good Laboratory Practice for Nonclinical Laboratory Studies.” Proposed rule, Federal Register, August 24, 2016. https://www.federalregister.gov/documents/2016/08/24/2016-19875/good-laboratory-practice-for-nonclinical-laboratory-studies
- International Council for Harmonisation. “General Principles for Model-Informed Drug Development, M15.” Final version adopted January 29, 2026. https://database.ich.org/sites/default/files/ICH_M15_Step4_Final_Guideline_2026_0129.pdf
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products.” Draft guidance, January 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological
- European Medicines Agency. “Reflection paper on the use of Artificial Intelligence (AI) in the medicinal product lifecycle.” EMA/CHMP/CVMP/83833/2023, September 2024. https://www.ema.europa.eu/en/documents/scientific-guideline/reflection-paper-use-artificial-intelligence-ai-medicinal-product-lifecycle_en.pdf
- U.S. Food and Drug Administration. “Roadmap to Reducing Animal Testing in Preclinical Safety Studies.” April 2025. https://www.fda.gov/files/newsroom/published/roadmap_to_reducing_animal_testing_in_preclinical_safety_studies.pdf
- International Council for Harmonisation. “Assessment and Control of DNA Reactive (Mutagenic) Impurities in Pharmaceuticals to Limit Potential Carcinogenic Risk, M7(R1).” https://database.ich.org/sites/default/files/M7_R1_Guideline.pdf
- “The importance of expert review to clarify ambiguous situations for (Q)SAR predictions under ICH M7.” Regulatory Toxicology and Pharmacology, 2019. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7510098/
- OECD. “Guidance Document on the Validation of (Quantitative) Structure-Activity Relationship [(Q)SAR] Models.” OECD Series on Testing and Assessment No. 69. https://www.oecd.org/content/dam/oecd/en/publications/reports/2014/09/guidance-document-on-the-validation-of-quantitative-structure-activity-relationship-q-sar-models_g1ghcc68/9789264085442-en.pdf
- “Comparison of Different Approaches to Define the Applicability Domain of QSAR Models.” Molecules, 2012. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6268288/
- U.S. Food and Drug Administration. “E14 and S7B Clinical and Nonclinical Evaluation of QT/QTc Interval Prolongation and Proarrhythmic Potential: Questions and Answers.” https://www.fda.gov/media/161198/download
- Li, Z., et al. “Assessment of an In Silico Mechanistic Model for Proarrhythmia Risk Prediction Under the CiPA Initiative.” Clinical Pharmacology & Therapeutics, 2019. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6492074/
- Creanza, T.M., et al. “Structure-Based Prediction of hERG-Related Cardiotoxicity: A Benchmark Study.” Journal of Chemical Information and Modeling, 2021. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9282647/
- Permadi, E.E., Watanabe, R., and Mizuguchi, K. “Improving the accuracy of prediction models for small datasets of Cytochrome P450 inhibition with deep learning.” Journal of Cheminformatics, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12044814/
- “Machine Learning to Predict Drug-Induced Liver Injury and Its Validation on Failed Drug Candidates in Development.” Toxics, 2024. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11207878/
- Lombardo, F., Bentzien, J., Berellini, G., and Muegge, I. “Prediction of Human Clearance Using In Silico Models with Reduced Bias.” Molecular Pharmaceutics, 2024. https://pubmed.ncbi.nlm.nih.gov/38285644/
- “Reliable Prediction of Caco-2 Permeability by Supervised Recursive Machine Learning Approaches.” Pharmaceutics, 2022. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9610902/
- Fossler, M.J., et al. “The FDA Roadmap to Reducing Animal Testing in Preclinical Safety Studies: Where Will It Lead Us?” Clinical Pharmacology in Drug Development, 2026. https://pubmed.ncbi.nlm.nih.gov/41766281/
- “The Trends and Future Prospective of In Silico Models from the Viewpoint of ADME Evaluation in Drug Discovery.” Pharmaceutics, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC10675155/








Your perspective matters—join the conversation.