Asking how accurate animal testing is will produce a strong response from many in pharmaceutical research. The most informed will cite statistics. The problem is, depending on who you ask, you can get wildly different numbers leading to starkly different conclusions. Some say animal studies are extremely accurate. Others say a clean animal result tells you almost nothing. Both sides cite peer-reviewed evidence. Even worse, both sides can be technically correct.
This post focuses on accuracy in toxicity studies. It looks at the papers each side points to, why their numbers refuse to reconcile, and a blind spot that further muddies the whole debate. The conclusion is not that one side wins. It is that the question is incredibly challenging to answer conclusively. It is inaccurate to conclude animal testing is extremely predictive in most cases and incorrect to say animal testing provides no value whatsoever. But perhaps most interestingly, the rise of New Approach Methodologies (NAMs) could make the debate irrelevant.
Two viewpoints. One shared dataset problem
Here, we’ll summarise and critique a few of the articles that are most cited by advocates and opponents of animal testing.
The defence of animal testing rests on a series of concordance studies. For example, Olson, et al. surveyed 150 compounds for an industry workshop in 2000 and found that for around 70% of them, at least one animal species had flagged a toxicity that later appeared in humans. Concordance was above 80% for haematological, gastrointestinal and cardiovascular effects, and below 35% for cutaneous and hypersensitivity reactions, with non-rodent species outperforming rodents. They also point out that 94% of concordant toxicities were identified in animals in one month or less. Monticello, et al. built the modern version in 2017, the IQ Consortium DruSafe database of 182 molecules, and reported a positive predictive value (PPV) of about 43% and a negative predictive value (NPV) of about 86%. They split this out across rodent, dog, and nonhuman primate in 12 organ types. They also evaluated likelihood ratios but chose to favour predictive value in the discussion.
Clark and Steger-Hartmann scaled the question up in 2018, mining more than 1.6 million adverse events across 3,290 drugs, and concluded that rat and dog identify human adverse events reasonably well for the specific endpoints those species are used to test, with high concordance in some specific observations (e.g., arrhythmia). Clark and Steger-Hartmann also introduced a more nuanced look by evaluating likelihood ratios and considering that in some cases (particularly negative outcomes) “the absence of animal toxicity is of limited predictivity for the lack of adverse events in humans”. In their abstract they state, “Our study confirmed the general predictivity of animal safety observations for humans…” but offer concessions in other sections. These include, “the lack of animal observation does not generally predict safety in human.” and “In general, many translations are confirmed as predictive, such as QT prolongation and other arrhythmias. Other pairs are statistically related but have low predictive value.”
The critique of animal testing rests on a reframing of much the same kind of data. Bailey, et al. took 2,366 drugs across rat, mouse and rabbit and focused on likelihood ratios rather than raw concordance. Their finding was that the absence of toxicity in animals carries almost no evidential weight for the absence of an adverse reaction in humans, and that the presence of toxicity, while sometimes informative, behaves inconsistently across drug classes. Van Norman’s 2019 review argued that animal models are poor predictors of human drug toxicity, pointing to the high failure rate of drugs in clinical trials. The review argued that perfectly predictive models should change the success rate of drugs from around 10% to above 50%, highlighting the gap of the current system.
To put it bluntly, these conclusions don’t line up.
Why the field cannot agree on a number
Let’s start with the most basic reason for the disagreement. The studies do not measure the same thing.
Olson reported concordance, the share of human toxicities with a matching animal finding. Monticello primarily considered predictive values, splitting the question into how often an animal positive predicts a human positive (about 43%) and how often an animal negative predicts a human negative (about 86%). Bailey focused on likelihood ratios (LRs), which ask how much an animal result shifts the probability of a human result once the base rate is removed. Likelihood ratios are often my preferred metric to evaluate the utility of a tool, but different metrics have different advantages and limitations.
Critically, these metrics are not interchangeable, and treating them as if they were is where much of the public argument fails. A single dataset can produce a reassuring 86% negative predictive value and a damning near-one negative likelihood ratio at the same time. Negative predictive value can be inflated by low frequency. If most drugs are safe for most endpoints, you score a high NPV almost regardless of whether the test is informative. The likelihood ratio controls for that inflation and asks whether the result actually moved your estimate. Bailey’s finding is that a negative animal result barely moves the estimate. Predictive value and likelihood ratio are both valid in specific contexts. They are simply different measurements, and quoting one to rebut the other is a category error that happens constantly.
Study design adds another layer. Concordance can improve or decrease depending on whether animal and human effects are compared at equivalent doses. Endpoint coding involves subjective mapping of findings to standard terms. This may be why Monticello, et al. and Clark & Steger-Hartmann offer such opposing conclusions on NPVs. Equally, concordance is high for some organ systems and low for others, so any single cited number is an average of very unequal number or misrepresents the average entirely.
So “how accurate is animal testing” has no single answer, because it is several questions fit into one conclusion. Accurate at what, measured how, on which drugs, at what exposure, for which organ system? Depending on how exactly you frame the question, the answer could be near 0%, near 100%, or anything in between. That gap is the difference between a tool you trust and a tool you deprioritise.
When 96% accurate means no predictive power
The confusion is not confined to the literature. It gets worse when a figure reaches a slide deck or a press release, where a specific statistic becomes a general claim directed to a non-scientific audience.
Take the statistic I have heard multiple times in public statements. “Dogs are up to 96% accurate in predicting safety outcomes in humans”. This sounds like a firm endorsement of dogs as animal models. It is also a near-perfect example of how these statistics get misrepresented (intentionally or otherwise).
That 96% claim comes from Monticello, et al. It is the negative predictive value (NPV) for pulmonary outcomes in dog. NPV is the proportion of the compounds that showed no pulmonary finding in dogs that also showed none in humans. It is not an overall accuracy figure, and on its own it says little about the dog’s ability to discriminate, only how often a clean result in dogs is matched by a clean result in humans. And it doesn’t describe the dog’s likelihood of identifying toxic outcomes. The other numbers for that endpoint change the picture substantially. For the same pulmonary outcomes, the positive predictive value (PPV) (toxic outcomes) was 0%. The positive likelihood ratio (LR+) was 0. The inverse negative likelihood ratio (iLR-) was 1.
Let’s translate that out of statistics.
The PPV of 0 means the dog never correctly predicted a pulmonary toxicity in humans. The inverse negative likelihood ratio of 1 is statistically identical to a coin toss. It means a negative result did not shift the probability of a human outcome at all. For this endpoint the dog could neither flag a real pulmonary toxicity nor rule one out. It was, in any practical sense, uninformative. Personally, I interpret an LR+ of 0 and an LR- of 1 as indicating that a test is functionally useless or that the sample size was too small to draw a conclusion. It just so happens, if you were trying to defend the utility of animal testing, this “96%” figure is one of the weakest statistics you could cite.
But how can a “useless” test produce a value like 96%? The answer is prevalence. Pulmonary toxicity was rare in this dataset. When an outcome is rare, predicting that it will not happen is right most of the time by default, whatever the test does. It would be like selling a device that predicts if a person would get hit by lightning that day. If it produced the result “no” every time, its negative predictive value would be near 100% despite offering no utility at all. The negative predictive value benefits from this skewed prevalence. It is measuring the rarity of the event, not the accuracy of the dog model. This is exactly why likelihood ratios exist. They are prevalence-independent, so they strip out the flattery and expose the absence of signal underneath.
The same illusion turns up in a completely different dataset. In the Clark and Steger-Hartmann big-data analysis, the preferred-term results for dog liver findings show negative predictive values between 0.82 and 0.95 with negative likelihood ratios between 0.86 and 0.99 (their Table 11). Hepatotoxicity, for instance, shows an NPV of 0.89 next to an LR- of 0.98. The high NPV looks reassuring. The near-one likelihood ratio says the negative result carries almost no information. To its credit, Clark and Steger-Hartmann actively discuss the importance of likelihood ratios in drawing conclusions.
This is similar to the metric problem from the previous section, now made worse by misinterpretation by third parties quoting the study. The same data point is a triumph or a catastrophe depending on which statistic you quote. A high negative predictive value is an easy “win”. The likelihood ratios that contradict it can stay unquoted. That is how a statistic demonstrating no predictive power ends up cited as proof of predictive power.
The blind spot: the most dangerous drugs may never reach humans
A defender of animal testing has an answer to this criticism of the 96% figure. It might be the strongest card in their deck. The reason the dog’s positive column looks so thin, they claim, is that the genuinely dangerous compounds never reach the clinic. The animal test caught them, and they were pulled before a single human was dosed. On this interpretation, the sparse positive data is not evidence of failure, but a demonstration of success.
It is a serious argument. But it is an argument that relies on absence of data, so it is also unprovable for virtually any modern drug.
Almost all concordance studies are built on drugs that reached humans, because a human outcome is the only thing an animal finding can be checked against. A compound that produces severe, disqualifying toxicity in animals usually is not given to people. It is stopped during preclinical studies. It has an animal result and no human result, so it cannot enter a concordance table. The compounds where the animal may have shown the strongest signal, and arguably did its most important work, are systematically missing from the data used to judge the animal test.
This has two consequences. First, the animal test gets no credit for real risks it hypothetically identified. If a preclinical signal correctly stopped a dangerous drug before it harmed a trial participant, that success is excluded from the analysis. You cannot score a true positive when you have deliberately prevented the human outcome that would confirm it. Second, the false positives are just as absent. A safe drug wrongly condemned by a misleading animal signal is also stopped, also never tested, also absent from the record. The lost beneficial drugs that critics point to are, by the same logic, unmeasurable. Potentially worse, because the data are often never released, it is never made clear if this risk could be identified by a single animal (instead of 100) or in a rodent (instead of a dog or primate). So these data cannot be leveraged to advance the 3Rs. The absence of data also prevents us from evaluating how frequently this happens. Are there 10 for every one drug that enters clinical trials or perhaps just a fraction of the whole? These are important, unanswered questions.
This is often described more gently as survivor bias, and Clark and Steger-Hartmann acknowledge it in a dataset built on approved drugs. But survivor bias undersells it. It is not only that the survivors look better than the full population. It is that the single most consequential thing an animal test might do, killing a compound outright, produces no measurable outcome at all.
What this means for a sponsor
Undoubtedly, animal testing has some utility. The high concordance for haematological, gastrointestinal and cardiovascular toxicity appears to be real. Unrelated to direct toxicity, animal studies are often key considerations in establishing first-in-human dose. However, animal data can sometimes be dangerously wrong. Cutaneous and immune-mediated toxicities translate poorly. Human-specific targets, including many monoclonal antibodies, are a known weak spot. Species-specific mechanisms mislead in both directions. In most studies that evaluate it, the likelihood ratio for most organs and tissues is well below being considered diagnostically useful.
With a million ways to measure and interpret a limited dataset, any side can claim victory.
The practical reality is that ICH M3(R2) still currently dictates the use of animal data. Regardless of your position on the utility of animal data, animal experiments are likely to continue in some capacity for the foreseeable future. The only realistic way we have to answer how accurate animal data are is to incorporate newer technologies and see if they “move the needle”.
Where NAMs come in
This is where the argument should turn from a debate about statistics into a plan to move forward. If you are a sponsor, the exact accuracy of the animal model shouldn’t be the end of the argument. It is still your responsibility to build the best possible preclinical package. That decision does not require the accuracy debate to be settled first.
Whatever the true baseline of animal testing is, it has room to improve. We already know animal models regularly translate poorly. NAMs represent dozens of diverse technologies. But integrating them into your preclinical strategy early and strategically can provide insights into the safety of a candidate drug that may reinforce the animal data or provide evidence of limitations.
NAMs can even reach into the structural blind spot this post described. The reason we cannot audit the drugs we stopped is that confirming their human risk would mean exposing humans. A human-relevant in vitro model offers a partial way around that. It lets you interrogate whether an animal-lethal compound would plausibly have harmed a human, without running the trial that ethics correctly forbids. It cannot fully recover the information that was lost, but it can begin to.
So, the scientifically honest position is not to just throw out a number. It is to acknowledge the utility and limitations of animal testing, then work towards whichever permutation of refinement, reduction, and replacement improves human health outcomes. Animal testing is neither the gold standard its defenders imply nor (putting aside ethical considerations) the complete failure its critics describe. It is a partial tool with measurable strengths and known (and unknown) gaps, assessed by often conflicting metrics. But it is, most certainly, insufficient to probe human safety alone. The productive response is not to keep relitigating the percentage. It is to add methods that target the gaps we can already name and build complex, combinatorial packages that are more human-relevant than animal data or any single new method alone.
Positioning NAMs as complements may improve human outcomes, but it does little to advance replacement of animals and we still leave reviewers sifting through potentially contradictory animal and NAMs data. Perhaps the more consequential approach to advance replacement is to target where animals are strongest: haematological, gastrointestinal, QT and arrhythmia signals, and other endpoints that carry the high concordance and positive likelihood ratios, along with non-toxicology applications like dose range-finding studies. A NAM does not need to beat a single global accuracy figure (which is functionally arbitrary). It needs to be validated for a defined Context of Use, measured against human outcomes, and exceed animal predictivity within that context. As those qualified contexts accumulate, the case for animals erodes one endpoint at a time. If targeted NAMs can eventually exceed the predictivity for the limited subset of endpoints where animals perform best, the scientific rationale for using those animals falls away. This will result in natural reductions in numbers of animals killed in testing. As data on the success and predictions of NAMs become more comprehensive, only then could a true path to full replacement be evaluated.
References
Bailey J, Thew M, Balls M. An analysis of the use of animal models in predicting human toxicology and drug safety. Altern Lab Anim. 2014;42(3):181–199. doi:10.1177/026119291404200306
Clark M, Steger-Hartmann T. A big data approach to the concordance of the toxicity of pharmaceuticals in animals and humans. Regul Toxicol Pharmacol. 2018;96:94–105. doi:10.1016/j.yrtph.2018.04.018
Monticello TM, Jones TW, Dambach DM, et al. Current nonclinical testing paradigm enables safe entry to first-in-human clinical trials: the IQ consortium nonclinical to clinical translational database. Toxicol Appl Pharmacol. 2017;334:100–109. doi:10.1016/j.taap.2017.09.006
Olson H, Betton G, Robinson D, et al. Concordance of the toxicity of pharmaceuticals in humans and in animals. Regul Toxicol Pharmacol. 2000;32(1):56–67. doi:10.1006/rtph.2000.1399
Van Norman GA. Limitations of animal studies for predicting toxicity in clinical trials: is it time to rethink our current approach? JACC Basic Transl Sci. 2019;4(7):845–854. doi:10.1016/j.jacbts.2019.10.008
No responses yet