A question posed during a debate at the MPS World Summit last week really stuck with me:
“Would you be the first person to take a medicine that had never been tested in an animal?”
Really, this is asking something much more blunt: “Would you trust your life on NAMs?”.
I realised, my answer was “Yes, but.” I could trust NAMs, but only if I had confidence in the approach and the individual systems.
However, “trust NAMs” is a generic statement without much practical meaning. Trust is not a property of the technology. It is a property of the package: how it was designed, how each component was built, and whether the combination is proportionate to the actual risk being taken.
Breaking that down reveals three distinct components of trust that any credible NAMs-based preclinical strategy must address.
1: Trusting the Individual NAM
Before a NAM contributes anything meaningful to a regulatory package, it must be trustworthy on its own terms. This means several things working together.
The Context of Use must be precisely defined. A NAM is not generically useful. It answers a specific biological question, in a specific context, with specific limitations. A liver-on-chip system qualified to detect hepatocellular injury caused by direct cytotoxicity is not the same tool as one qualified to detect cholestatic injury or immune-mediated DILI. A COU written broadly enough to appear to cover all of them is not a strength. It is a red flag. Regulators evaluate fit-for-purpose, not ambition. The FDA’s March 2026 draft guidance, General Considerations for the Use of New Approach Methodologies in Drug Development, is explicit on this point: NAM validation is defined relative to a specific COU, not in the abstract.
Benchmarking data must tell a credible story. How does this NAM perform against a reference dataset? What is the false negative rate for known toxicants? What are the limits of detection? Has performance been demonstrated across drug classes relevant to the proposed COU? These are not academic questions. They are the evidentiary foundation on which a regulator will decide whether the data can inform a decision. Standardisation and inter-laboratory comparisons are still underdeveloped across much of the field. They sit at the heart of this problem.
Quality systems must underpin every data point. Reproducibility should not be assumed from a well-designed assay. To the contrary, regardless of the assay, reproducibility should be assumed absent until demonstrated otherwise. And reproducibility can only be maintained through documented processes, controlled materials, and a quality management system capable of catching drift before it becomes a pattern. A NAM without a quality backbone is an academic prototype, not a product. Variability between operators, sites, or batches does not disappear because the biology is human-relevant. In fact, superior human-relevance may artificially mask variability that undermines the strength of the experiment. Quality can never be ignored.
2: Calibrating Trust to Risk
Before assembling those individual NAMs into a package, developers need to know what standard the package will be held to. That standard is not fixed. It scales with the risk of the clinical program.
For example:
Lower-risk programs carry a proportionally lower evidentiary burden. A biosimilar for a clinically approved monoclonal antibody with an established safety record and a well-characterised target is the clearest example. The clinical comparator provides a safety anchor. In this context, a NAMs package with good COU definitions, reasonable quality controls, and validation against the relevant endpoints may well be sufficient. The bar is meaningful but not extreme.
Moderate-risk programs represent the sweet spot where NAMs can make the most immediate progress. A novel monoclonal antibody against a new target with poor cross-reactivity in standard preclinical species is a good example. Animal models may actively mislead here. A human-relevant NAM package may provide genuinely superior predictivity. But the evidentiary standards must match the increased reliance. Validation breadth, cross-platform benchmarking, and explicit documentation of assumptions are all required. The number and types of NAMs used may increase proportional to that risk.
Higher-risk programs place the highest demands on every layer of the framework. First-in-class small molecules with novel mechanisms, significant metabolic complexity, or unknown immune interactions sit at this end of the spectrum. Quality systems might need to approach or meet GLP standards. Data integrity must be demonstrable under something equivalent to Annex 11 or 21 CFR Part 11 requirements. The breadth of preclinical questions being answered must be fully mapped and comprehensively addressed. The known limitations of the package including systemic responses, complex metabolite toxicity, and immune interactions must be explicitly characterised, not minimised.
This risk stratification is not a new concept. Regulators apply it implicitly every day. What is missing is its explicit application to the design and evaluation of NAMs-based preclinical packages. Importantly, this risk evaluation exists on a spectrum. It is the responsibility of the sponsor and the reviewer to ensure this risk has been adequately appreciated and accounted for in the preclinical strategy.
3: Trusting the Package
With individual NAMs validated and a risk tier established, the final question is whether the package as a whole covers the ground it needs to. Even a collection of rigorously validated NAMs can fail at this level. A panel of NAMs isn’t successful just because it’s diverse. It is successful because it is comprehensive. No single NAM covers the full toxicological landscape of a new drug candidate. The gaps matter as much as what is covered.
Assembling a credible package requires mapping the relevant risk profile of the candidate specifically, not generically. What organ systems are most likely to be affected? What mechanisms of toxicity are plausible given the target, the drug class, and the route of administration? What duration of exposure is relevant?
Each of those questions defines a COU requirement. The package earns trust when the union of the individual COUs comprehensively addresses the risk profile, the gaps are explicitly acknowledged, and the design logic is documented and defensible. This is the foundation of a weight of evidence approach: no single NAM needs to be definitive, but the collective package must be. The FDA’s draft NAMs guidance frames this explicitly, noting that NAMs data should support regulatory decisions through a structured, question-driven weight of evidence assessment rather than as stand-alone substitutes for individual animal studies.
This is not how NAMs packages are always built today. More often, developers include the NAMs they have: the ones they know, the ones that are validated, the ones that are convenient. Gaps are described as future work. That is acceptable for exploratory science. It is not acceptable as the foundation for a first-in-human decision.
What This Means for NAM Developers
The question of whether a NAMs-based preclinical package can be trusted is answerable. But the answer is not “yes, because NAMs are human-relevant.” It is “yes, because this specific package was designed to address this specific risk profile, each component meets these specific quality and validation standards, and the combination was assembled according to a documented logic that a regulator can evaluate and challenge.”
That is a harder case to make than “we used human biology.” It is also the only case that should ever get a new drug into a first-in-human trial.
The field has made remarkable progress on the science. The next frontier is building the frameworks for quality, package design, and risk-stratified standards that turn that science into something a patient, a clinician, and a regulator can actually rely on.
What would it take for you to be the first volunteer for an animal-free medicine?
At InnovApproach Consulting, we help NAM developers and sponsors build preclinical strategies that are scientifically rigorous, regulatory-ready, and designed around the specific risk profile of the program. Get in touch
to discuss your qualification or package design challenges.

No responses yet