Scientific reviewPosition not yet filled, stated openly.
10 min readLast substantive review
Open access
Executive summary
Animal studies generate hypotheses, not clinical evidence. An overview of 122 systematic reviews found that roughly half of the therapies tested in animals ever reach a human study, and that only about 5 per cent obtain approval. A finding transfers only when mechanism, exposure, outcome measure and population match up, and when the study was randomised, blinded and reported in full.
20primary sources
50 %of them level 1 to 2
4species studied
2007–2024publication years
Key points
Of the therapies tested in animal experiments, an overview of 122 systematic reviews reports that around 50 per cent reach a human study and about 5 per cent reach approval, a median of ten years after the animal experiment.
Part of the gap is biology; a large part is method. In an analysis of 31 systematic reviews, only around 29 per cent of the animal studies reported randomisation and around 35 per cent a blinded assessment.
Voluntary checklists are not enough: in a randomised trial covering 1,689 manuscripts, not a single one met the ARRIVE requirements in full.
Simulations based on 440 preclinical studies suggest that tightly standardised single-laboratory studies often fail to estimate effect sizes accurately; designs with two to four laboratories raised accuracy by up to 42 percentage points.
The horse is model and target species at once, yet the burden of proof stands there too: syntheses of joint and skin treatments in the horse rate most studies as highly prone to bias.
What transferability actually means in research
Transferability is not a feeling. It is a chain of four tests. Internal validity asks whether the measured result can be traced to the treatment at all within the study; randomisation, allocation concealment and blinded assessment are what secure it. External validity asks whether the same result holds outside that one experimental set-up, so at a different age, under different husbandry, in a different sex, in a different season. Construct validity asks whether the model reproduces the disease one actually wants to treat. Predictive validity asks, last of all, whether a positive animal result predicts a positive result in the target species. If any one of the four tests fails, the chain breaks, and nothing in the result itself will show it.
Why is a plausible mechanism not enough?
Because a mechanism establishes only that an effect is possible, not its size and not its direction in a living organism. Between receptor and recovery lie distribution, breakdown, counter-regulation, concurrent disease and time. A shared mechanism is a hypothesis; clinical benefit is a finding. The two must never be fused into a single sentence.
How many animal studies reach human beings
The most robust estimate comes from a 2024 review that pooled 122 systematic reviews covering 54 diseases and 367 therapeutic approaches. Roughly half of the approaches tested in animals reached a human study at all, around 40 per cent a randomised clinical trial and about 5 per cent regulatory approval. The road is long: a median of five years to the first human study, seven to the randomised trial, ten to approval. At the same time, the same analysis found concordance of positive results between animal and human of around 86 per cent.
A broad systematic scoping review of 121 analyses, each comparing at least two species including humans, arrived at a very different picture: the transfer success rates reported there ranged from 0 to 100 per cent. The authors judged the risk of bias in the included work to be high and much of the material to be out of date, and concluded that the success of a transfer can barely be predicted in advance.
A concrete case: when animal and human diverge
A systematic review compared six interventions whose effect in humans had been settled beyond doubt with the corresponding animal data. For the free radical scavenger tirilazad, the animal models reduced infarct volume by around 29 per cent and improved neurological scores by around 48 per cent; in patients with ischaemic stroke, the same substance produced a worse course. Corticosteroids in traumatic brain injury looked favourable in the animal model and showed no benefit in clinical testing. Conversely, thrombolysis and bisphosphonates pointed the same way in both worlds; for antenatal corticosteroids the direction held for respiratory distress in the newborn, while the effect on mortality remained open in the animal model. The direction is therefore often right, but not reliably so.
A success shrinks along the way as well. A systematic review of the oncolytic virus talimogene laherparepvec, approved in 2015, traced five preclinical studies in 150 animals and seven clinical studies in 589 patients. Preclinically, complete regression was reported in up to 100 per cent of injected and 80 per cent of non-injected tumours; the highest complete response rate achieved clinically was 24 per cent, and in one study 0 per cent. Every preclinical study carried a high or unclear risk of bias.
Why effects shrink: bias first
Before reaching for biology, it is worth looking at method. An overview of 31 systematic reviews of animal experiments spanning stroke, spinal cord disease, tumours, multiple sclerosis and emergency medicine found this: only around 29 per cent of the studies reported randomisation, around 15 per cent allocation concealment and around 35 per cent a blinded assessment. In the main analysis, missing randomisation went together with markedly larger effects; in the secondary analyses, randomisation, allocation concealment and blinding lowered effect sizes above all where the outcome was a judgement rather than a measurement.
An analysis of 549 studies in 9,535 animals on acute spinal cord injury puts a precise figure on the price. Across all interventions, an improvement in neurological recovery of around 26 per cent was reported. More than half of the studies, 51.2 per cent, described no blinded assessment, and the effect thus inflated came out 7.2 percentage points higher. The publication bias analysis suggested that between 2 and 41 per cent of the experiments were never published; counting them in lowers the estimated effect by 0.9 to 14.3 percentage points.
Measured consequences of missing methodological control in animal studies, compiled from an overview of 31 systematic reviews and an analysis of 549 spinal cord studies.
Methodological feature
Share of studies reporting it
Observed consequence
Randomised allocation
around 29 per cent
Missing randomisation went together with markedly larger effects
Allocation concealment
around 15 per cent
Where present, smaller effects, above all with subjective outcomes
Blinded assessment
around 35 per cent
Without blinding, the effect in the spinal cord analysis came out 7.2 percentage points higher
Full publication
an estimated 2 to 41 per cent of experiments remain unpublished
Counting them in lowers the effect by 0.9 to 14.3 percentage points
Why effects shrink: then the real species differences
The second block is biology, and it concerns exposure above all. A review of dose extrapolation between species records that simple scaling by body size fails wherever a substance binds strongly to proteins, is extensively metabolised or actively transported, is excreted in the bile, or wherever target structures differ between species in quantity, binding strength and distribution. The same amount per kilogram then does not mean the same concentration at the site of action, and the same concentration does not mean the same response. That is why no dose is ever converted from one species to another on this platform.
The third block is paradoxical: too much order. A simulation study based on 440 preclinical studies of 13 interventions in models of stroke, myocardial infarction and breast cancer compared single-laboratory designs with multi-laboratory designs. Single-laboratory studies mostly failed to estimate the true effect size accurately, and larger animal numbers made the estimate less accurate still, because they measure a narrow special case ever more precisely. Just two to four participating laboratories raised accuracy by up to 42 percentage points, without needing more animals. The bottleneck is not the species; it is the narrowness of the sample.
ARRIVE: why checklists alone are not enough
Since 2010 the ARRIVE guideline has set out the minimum information a publication reporting an animal experiment must contain. The revised version of 2020 arranges 21 items and singles out ten of them as an indispensable core, among them the species and origin of the animals, group size, randomisation, blinding and the full results. Numerous journals and funding bodies endorse it; the authors of the 2020 version note themselves, however, that the hoped-for improvement in reporting quality has failed to materialise since 2010.
Whether that improves the reports has actually been tested in randomised fashion. In a trial covering 1,689 manuscripts submitted to a large journal, half of the author teams were asked to complete an ARRIVE checklist. Not one manuscript in either group met the requirements in full; the only thing that improved was the description of housing conditions, from 52.1 to 74.1 per cent. A comparison of journals with and without official endorsement of the guideline found around 60 per cent of items fully reported in both groups in 2015, and no difference between the groups; it was precisely the bias-relevant information that was poorly reported.
The situation is not hopeless. A cross-sectional analysis of 943 animal experiments taken from the journals that originally published the guideline shows a clear trend across three time windows: the share of poorly reported studies fell from 46.05 to 26.06 and finally to 9.55 per cent. Even so, not a single study was reported in full. Not one of the 943 covered all 38 sub-items, and only in the most recent window did as much as 0.25 per cent reach the grade ‘excellent’, with none at all in the two earlier ones. Reports came out better where the journal made compliance mandatory.
What does that mean for me as a reader?
A completed checklist is not a seal of quality. Check in the full text whether randomisation, blinding and sample size planning really are described. If they are missing, the result is not wrong, but it carries less weight than the abstract suggests. The phrase ‘reported in accordance with ARRIVE’ is no substitute for that check.
The horse: model and target species at once
The horse occupies a rare dual position. For articular cartilage it counts, among the models in common use, as the animal whose cartilage and subchondral bone thickness comes closest to that of the human being; defects can be created and assessed arthroscopically, and loading after treatment can be steered under control. In the tendons the parallel is closer still: the superficial digital flexor tendon of the horse becomes diseased spontaneously and is regarded as the counterpart of human Achilles tendinopathy, tears of the deep digital flexor tendon as the counterpart of rotator cuff tears within the tendon sheath. These diseases arise of their own accord; they are not induced.
Being a good model does not, however, release the horse from the burden of proof. In a randomised, blinded study in 17 horses with an induced lesion of the superficial digital flexor tendon, a locally injected microRNA lowered collagen 3 activity in the tissue after two weeks; the cross-sectional area of the lesion came out smaller and the histological score better than under sham treatment. In a controlled trial in six horses, by contrast, mesenchymal stromal cells produced a temporary increase in inflammation and lesion size between weeks 3 and 6, and after 24 weeks no clear advantage over serum treatment alone could be found.
The same weaknesses apply in equine medicine as in preclinical research. A systematic review with meta-analysis of platelet-rich plasma for joint disease in the horse found 21 publications, of which only five could be pooled; the pooled estimate came out clearly in favour of treatment, yet most studies were rated as highly prone to bias, and the authors call for randomised, blinded comparative studies. A systematic review of the treatment of equine sarcoids included ten studies: the reported regression rates ranged from 28 to 100 per cent per lesion and from 9 to 100 per cent per horse, with a risk of bias running from ‘some concerns’ to ‘critical’. No recommendation for any particular treatment could be derived from it.
Checklist: placing an animal study in ten minutes
The following eight questions can be answered of almost any publication, and they place a result faster than any debate about whether animal experiments make sense.
Which species, which breed, which age, which sex? If the species is already missing from the title, the text is imprecise before it has begun.
Was there randomisation, and was the allocation concealed from the people running the experiment?
Was the outcome assessed blind? And is it measured or judged? Judgements react most strongly to missing blinding.
How many animals, and was the group size justified in advance?
One laboratory or several? Simulations suggest that a single, tightly standardised laboratory often fails to estimate the effect size accurately.
Does the exposure fit the question: dose, duration, route of administration, and was a concentration actually measured?
Did the disease arise spontaneously or was it induced experimentally? Both are legitimate, but they are different statements.
Are there already data in the target species, and what do they show? If so, those weigh more heavily than any result from a model.
Anyone wanting to proceed more systematically will find, in the risk of bias instrument for animal studies available since 2014, ten ordered items adapted from the established tool for clinical trials. It covers selection, performance, detection, attrition and reporting bias, and makes the assessment traceable rather than intuitive.
The sober balance is this: animal research is no proof for the human being, and human research is no proof for the horse. Both produce knowledge that travels in the other direction only as a hypothesis. Keeping that separation clean costs little and gains precision.
Transferability grid: which claim needs which body of data
Original analysis
Constructed by crossing the five common claim types of a health news item with the limits of animal-to-human transfer documented in the syntheses cited here; the column ‘Minimum needed in the target species’ names the smallest body of data below which, under the doctrine of this platform, the claim is not made at all.
Claim type
What animal data can do
What animal data never do
Minimum needed in the target species
Most common fallacy
Mechanism
Make the signalling pathway, cell type and time course visible
Establish the relevance of the pathway to the course of the disease
Evidence that the same structure is present and active in the target species
The mechanism is passed off as an effect
Safety
Deliver coarse warning signals at organ level
Rule out rare and late adverse effects
Systematic recording of adverse events in the target species
‘Unremarkable in animals’ is read as ‘harmless’
Dose
Narrow down a range for the first trial
Set an effective amount for another species
Measured concentration and effect in the target species
Milligrams per kilogram are simply converted across
Clinical benefit
Establish a testable hypothesis
Show the benefit in another species
Randomised, blinded study in the target species
The effect size from the model is quoted as though it held clinically
Prevention at herd level
Isolate individual risk factors
Reproduce the interplay of husbandry, feeding and use
Observational data from the real population
A laboratory finding becomes a husbandry recommendation
Limitations and uncertainty
The figures on transfer scatter enormously: one review arrives at 5 per cent approval alongside 86 per cent concordance, another at a range of 0 to 100 per cent. The two are not measuring the same thing, and neither figure works as a headline.
Meta-research examines publications, not organisms. A missing statement about blinding does not necessarily mean that no blinding took place; part of the measured bias could be a pure reporting problem.
The robust figures on the rate of transfer come from a handful of fields: neurological disease accounts for around a third of the reviews analysed, and in a further analysis 20 of 31 papers concerned stroke alone. No comparable analysis exists for the horse.
The evidence in favour of deliberately heterogeneous study designs rests so far mainly on simulations and on a few experiments in mice, in which reproducibility improved in only about half of the comparisons tested. It has not been tested in large species with small groups.
The syntheses cited draw almost exclusively on English-language literature; work in other languages is systematically under-represented.
No evidence on this page permits statements about individual substances, amounts or treatment decisions in a particular animal or person.
Open questions
Which features of an animal study, specifiable in advance, actually predict later clinical confirmation, and can they be fixed in a binding way before the study begins?
Does a deliberately heterogeneous study design improve reproducibility in large species with small group sizes as well, in the horse for instance?
How high is the rate of transfer within veterinary medicine, that is from the rodent model to the horse? Robust analyses of this are so far lacking.
Can reporting quality be raised effectively by binding editorial processes, now that voluntary checklists have proved ineffective in a randomised trial?
Frequently asked questions
Does that mean animal experiments are useless?
No, that is precisely what the data do not show. The same review that counts only around 5 per cent approvals finds concordance of positive results between animal and human of around 86 per cent, and several treatments taken for granted today, such as thrombolysis in stroke or antenatal corticosteroids, have been confirmed in both worlds. The criticism is not aimed at the animal model as such, but at two things: at studies run without randomisation and blinding, and at the habit of presenting a result from a model as a clinical finding.
Why does one study say 5 per cent and the other 0 to 100 per cent?
Because they count different things. The 5 per cent refers to all the modes of action tested and asks how many of them obtain approval. The range of 0 to 100 per cent comes from a scoping review of 121 analyses, each of which defined very differently what ‘successful transfer’ even means: in one case prediction within a factor of two, in another the share of correctly predicted outcomes, in another the overlap of statistical ranges. Without a common definition, comparing the percentages is pointless.
Can I convert a dose from the horse to the human being?
No, and this site does not do it on principle. Conversion by body weight or body surface area fails as a rule for substances that bind strongly to proteins, are extensively metabolised or actively transported, or are excreted in the bile. On top of that come differences between species in the quantity, binding strength and distribution of the target structures. The same amount per kilogram therefore does not produce the same concentration at the site of action. Questions about use belong in the hands of the attending veterinary surgeon or the attending doctor.
What is gained when a disease arises of its own accord in the horse?
It improves construct validity, that is the question of whether the model reproduces the disease at all. A spontaneously arising tendon disease in the superficial digital flexor tendon of the horse develops over years under real loading, with age, prior strain and concurrent disease, whereas an artificially created defect leaves all of that out. About the transferability of a treatment effect it says nothing, however: with spontaneous disease too, randomised, blinded studies in the target species remain the basis of any statement about efficacy.
How can I tell from the abstract that an animal study is weak?
By three signals. First: the species is not named, or only deep in the text. Second: the words randomised and blinded are missing even though a treatment is being compared. Third: the outcome is a judgement, a rating scale for instance, with no indication of who recorded it. This information is missing more often than one would expect: in an analysis of 943 animal experiments not a single study covered all the items, and even in the most recent window only 0.25 per cent reached the grade ‘excellent’. Missing information does not refute a result, but it lowers its weight.
Does all this hold the other way round as well, from the human being to the horse?
Yes, in both directions and under the same rules. A treatment effect established in humans is, in the horse, a hypothesis to begin with, because digestion, metabolism, loading profile, lifespan and husbandry all differ. The experience of human medicine in designing clinical trials, by contrast, does transfer perfectly well, and that is exactly what specialists from both fields recommend for tendon research. The method, then, transfers more reliably than the result.
Sources
Ineichen BV, Furrer E, Grüninger SL, Zürrer WE, Macleod MR. Analysis of animal-to-human translation shows that only 5% of animal-tested therapeutic interventions obtain regulatory approval for human applications. PLoS Biology, 2024 (Meta-analysis | Multiple species)DOI 10.1371/journal.pbio.3002667 Umbrella review of 122 systematic reviews covering 367 interventions: 50 per cent reach a human study, 40 per cent a randomised trial and 5 per cent approval, with 86 per cent concordance of positive results and median times of 5, 7 and 10 years.
Leenaars CHC, Kouwenaar C, Stafleu FR, Bleich A, Ritskes-Hoitinga M, De Vries RBM, Meijboom FLB. Animal to human translation: a systematic scoping review of reported concordance rates. Journal of Translational Medicine, 2019 (Systematic review | Multiple species)DOI 10.1186/s12967-019-1976-2 Scoping review of 121 reviews: the transfer rates reported range from 0 to 100 per cent; because of high risk of bias and outdated material, the authors consider the cumulative evidence insufficient and the success of transfer unpredictable.
Perel P, Roberts I, Sena E, Wheble P, Briscoe C, Sandercock P, Macleod M, Mignini LE, Jayaram P, Khan KS. Comparison of treatment effects between animal experiments and clinical trials: systematic review. BMJ, 2007 (Systematic review | Multiple species)DOI 10.1136/bmj.39048.407928.BE For six interventions with unambiguous clinical data, animal and human agreed in part (thrombolysis, bisphosphonates; antenatal corticosteroids only for respiratory distress, not for mortality) and in part did not: tirilazad reduced infarct size in animals by 29 per cent and improved neurological scores by 48 per cent, yet in humans was associated with a worse course, and corticosteroids in traumatic brain injury worked only in the animal model.
Hirst JA, Howick J, Aronson JK, Roberts N, Perera R, Koshiaris C, Heneghan C. The need for randomization in animal trials: an overview of systematic reviews. PLoS One, 2014 (Systematic review | Multiple species)DOI 10.1371/journal.pone.0098856 Across 31 systematic reviews, 29 per cent of the animal studies reported randomisation, 15 per cent allocation concealment and 35 per cent a blinded assessment; missing randomisation increased effect sizes significantly.
Watzlawick R, Antonic A, Sena ES, Kopp MA, Rind J, Dirnagl U, Macleod M, Howells DW, Schwab JM. Outcome heterogeneity and bias in acute experimental spinal cord injury: A meta-analysis. Neurology, 2019 (Meta-analysis | Multiple species)DOI 10.1212/WNL.0000000000007718 549 studies in 9,535 animals: mean reported improvement 26.3 per cent; 51.2 per cent without blinded assessment, with an effect inflated by 7.2 percentage points; 2 to 41 per cent of the experiments probably unpublished, and including them lowers the effect by 0.9 to 14.3 points.
Hair K, Macleod MR, Sena ES. A randomised controlled trial of an Intervention to Improve Compliance with the ARRIVE guidelines (IICARus). Research Integrity and Peer Review, 2019 (Randomised trial | Species-independent)DOI 10.1186/s41073-019-0069-3 1,689 randomised manuscripts: the editorial request to complete an ARRIVE checklist produced full compliance in neither group; only the information on housing improved, from 52.1 to 74.1 per cent.
Lalu M, Leung GJ, Dong YY, Montroy J, Butler C, Auer RC, Fergusson DA. Mapping the preclinical to clinical evidence and development trajectory of the oncolytic virus talimogene laherparepvec (T-VEC): a systematic review. BMJ Open, 2019 (Systematic review | Multiple species)DOI 10.1136/bmjopen-2019-029475 Five preclinical studies in 150 animals reported complete regression of up to 100 per cent (injected tumours) and 80 per cent (non-injected); clinically the highest complete response rate was 24 per cent, and 0 per cent in one study, with a consistently high risk of bias.
Peng C, Yang L, Labens R, Gao Y, Zhu Y, Li J. A systematic review and meta-analysis of the efficacy of platelet-rich plasma products for treatment of equine joint disease. Equine Veterinary Journal, 2024 (Meta-analysis | Horse)DOI 10.1111/evj.14042 21 publications included, 5 of them amenable to meta-analysis: the pooled effect favoured platelet-rich plasma, yet most studies carried a high risk of bias; the authors call for randomised, blinded comparative studies and a uniform classification system.
Offer KS, Dixon CE, Sutton DGM. Treatment of equine sarcoids: A systematic review. Equine Veterinary Journal, 2023 (Systematic review | Horse)DOI 10.1111/evj.13935 Ten studies: reported regression rates of 28 to 100 per cent per lesion and 9 to 100 per cent per horse, risk of bias running from ‘some concerns’ to ‘critical’; the evidence is not sufficient to recommend one treatment over another.
Watts AE, Millar NL, Platt J, Kitson SM, Akbar M, Rech R, Griffin J, Pool R, Hughes T, McInnes IB, Gilchrist DS. MicroRNA29a Treatment Improves Early Tendon Injury. Molecular Therapy, 2017 (Randomised trial | Horse)DOI 10.1016/j.ymthe.2017.07.015 Randomised, blinded study in 17 horses with an induced lesion of the superficial digital flexor tendon: intralesional microRNA treatment lowered collagen 3 transcripts, reduced the cross-sectional area of the lesion and improved the histological score compared with placebo.
Percie du Sert N, Hurst V, Ahluwalia A, Alam S, Avey MT, Baker M, Browne WJ, Clark A, Cuthill IC, Dirnagl U, Emerson M, Garner P, Holgate ST, Howells DW, Karp NA, Lazic SE, Lidster K, MacCallum CJ, Macleod M, Pearl EJ, Petersen OH, Rawle F, Reynolds P, Rooney K, Sena ES, Silberberg SD, Steckler T, Würbel H. The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research. PLoS Biology, 2020 (Professional guideline | Species-independent)DOI 10.1371/journal.pbio.3000410 Revised reporting guideline with 21 items, split into an indispensable core of ten items and a supplementary set; the authors note that the hoped-for improvement in reporting quality has failed to materialise since 2010.
Lin Y, Yang F, Shang B, Speich JE, Wan YY, Hashida H, Braun T, Sadoughi A, Puehler T, Lue TF, Zhang K. Reporting quality of animal research in journals that published the ARRIVE 1.0 or ARRIVE 2.0 guidelines: a cross-sectional analysis of 943 studies. Cardiovascular Diagnosis and Therapy, 2024 (Cross-sectional study | Species-independent)DOI 10.21037/cdt-24-413 Across three time windows the share of poorly reported studies fell from 46.05 to 26.06 and then 9.55 per cent, yet none of the 943 studies reported all 38 sub-items, and the share of ‘excellent’ quality stood at 0, 0 and finally 0.25 per cent; mandatory requirements in the instructions to authors went together with better reporting quality.
Leung V, Rousseau-Blass F, Beauchamp G, Pang DSJ. ARRIVE has not ARRIVEd: Support for the ARRIVE (Animal Research: Reporting of in vivo Experiments) guidelines does not improve the reporting quality of papers in animal welfare, analgesia or anesthesia. PLoS One, 2018 (Cohort study | Species-independent)DOI 10.1371/journal.pone.0197882 236 publications from endorsing and non-endorsing journals: in 2015 both groups stood at around 60 per cent of items fully reported, with no difference between the groups; not a single paper reported every item, and bias-relevant information was missing particularly often.
Voelkl B, Vogt L, Sena ES, Würbel H. Reproducibility of preclinical animal research improves with heterogeneity of study samples. PLoS Biology, 2018 (Other | Multiple species)DOI 10.1371/journal.pbio.2003693 Simulations based on 440 preclinical studies of 13 interventions: single-laboratory designs mostly failed to estimate effect sizes accurately, and larger samples made this worse still; designs with two to four laboratories raised accuracy by up to 42 percentage points.
Hooijmans CR, Rovers MM, de Vries RBM, Leenaars M, Ritskes-Hoitinga M, Langendam MW. SYRCLE's risk of bias tool for animal studies. BMC Medical Research Methodology, 2014 (Professional guideline | Species-independent)DOI 10.1186/1471-2288-14-43 Instrument derived from the Cochrane tool, with ten items covering selection, performance, detection, attrition and reporting bias, adapted to the particularities of animal experiments.
Sharma V, McNeill JH. To scale or not to scale: the principles of dose extrapolation. British Journal of Pharmacology, 2009 (Other | Multiple species)DOI 10.1111/j.1476-5381.2009.00267.x Review: scaling by body size fails for strongly protein-bound, extensively metabolised, actively transported or biliary excreted substances, and where species differ in the quantity, binding strength and distribution of the target structures.
McIlwraith CW, Fortier LA, Frisbie DD, Nixon AJ. Equine Models of Articular Cartilage Repair. Cartilage, 2011 (Other | Horse)DOI 10.1177/1947603511406531 Review: the horse comes closer to human cartilage and subchondral bone thickness than the usual models; defects can be created and followed up arthroscopically, and loading after treatment can be controlled.
Ahrberg AB, Horstmeier C, Berner D, Brehm W, Gittel C, Hillmann A, Josten C, Rossi G, Schubert S, Winter K, Burk J. Effects of mesenchymal stromal cells versus serum on tendon healing in a controlled experimental trial in an equine model. BMC Musculoskeletal Disorders, 2018 (Controlled trial | Horse)DOI 10.1186/s12891-018-2163-y Controlled trial in six horses: mesenchymal stromal cells produced temporarily stronger inflammation and larger lesions between weeks 3 and 6; after 24 weeks there was no clear advantage over serum treatment alone.
Smith RKW, McIlwraith CW. "One Health" in tendinopathy research: Current concepts. Journal of Orthopaedic Research, 2021 (Other | Multiple species)DOI 10.1002/jor.25035 Workshop review: the horse’s superficial digital flexor tendon is regarded as the naturally diseasing counterpart of human Achilles tendinopathy, and tears of the deep digital flexor tendon as the counterpart of rotator cuff tears within the tendon sheath; the authors recommend transferring the study methodology of human medicine to equine studies.
von Kortzfleisch VT, Karp NA, Palme R, Kaiser S, Sachser N, Richter SH. Improving reproducibility in animal research by splitting the study population into several 'mini-experiments'. Scientific Reports, 2020 (Controlled trial | Rodent)DOI 10.1038/s41598-020-73503-4 Experimental test in mouse strains: a design deliberately split across several time points improved reproducibility and the rate of correct findings compared with the conventionally standardised design, though only in around half of the strain comparisons examined.
ForschungPferd (2026). Animal studies and human studies: what really transfers. ForschungPferd, English. https://forschungpferd.ch/en/evidence/animal-and-human-studies/