Weight-of-Evidence has become a key phrase in drug development. It shows up in nearly every New Approach Methodologies (NAMs) guidance regulators have published over the last two years. Although it has been used by regulators for decades, its impact has changed. Weight-of-Evidence is quickly becoming the cornerstone of evaluating modern preclinical packages.
A term this important ought to be well understood. Unfortunately, it isn’t.
Too many NAMs developers and drug sponsors misinterpret it.
Some groups treat it like a single data point. Feed in your preclinical data, average it out, and read off a score. If it’s safe enough, you can start clinical trials. This reasoning buries organ-specific risks. It doesn’t matter if 99% of tissues are safe. If it is toxic to a key organ like brain or liver, reviewers won’t approve it.
Others treat it like a volume test. Run enough studies, tick enough boxes, and the reviewer is satisfied. This mistakes quantity for reasoning. You can run twenty assays, but if they lack biological relevance, you may as well have done none.
Weight-of-Evidence is neither of these things. It is a structured way to answer a comprehensive set of scientific questions through multiple independent experiments, weighted by regulatory precedent and by the risks specific to your drug. Evaluated together, these experiments provide confidence that the candidate drug is reasonably safe to test in humans. The weighting is key. Not every experiment counts equally, and not every question carries the same risk.
This is where it becomes critical for NAMs.
A NAMs-based preclinical package does not mean replacing one animal study with one NAM experiment. Instead, you assemble several evidence streams, each built to speak to a specific safety question, and you let them work together to tell a comprehensive story.
The Context of Use (COU) defines those questions. A COU sets out exactly what a method is qualified to answer and nothing beyond it. Get the COU right and every method has a job. Get it wrong and you have data with no home and a reviewer with no reason to trust it.
Some evidence streams may address the same question. This is the basis of NAMs complementing traditional animal experiments rather than replacing them. But what happens if these produce contradictory data?
What happens when one test shows a safety risk and another clears it?
This does not mean the package is poorly constructed. Real evidence contradicts itself, and a package that never shows any tension probably isn’t looking hard enough. The mistake is to average the conflict away. You don’t average. You compare.
Validation data and biological relevance decide which stream carries more weight for that question. A well-qualified, mechanistically relevant human model beats a poorly characterised one, even when the poorly characterised one gives the answer you were hoping for. Sometimes the honest conclusion is that neither stream settles it and you need another experiment.
None of this happens by accident. Weight-of-Evidence is not something you bolt on at the end to justify a decision already made. It is something you design from the start.
Whether you are building a single NAM or a full preclinical package, planning your Weight-of-Evidence is no longer optional.

No responses yet