The commercial pitch is compelling. Hamstring injuries cost clubs enormous amounts in wages, medical costs, lost performance and depressed transfer values. If a model could flag high-risk players before the injury, the return would be obvious.

Many systems claim to do this. Very few have held up under independent evaluation, and the reasons say something interesting about the limits of prediction in this domain.

The base rate problem

Start with the arithmetic that undermines most claims.

Suppose a squad of thirty players and twenty significant soft-tissue injuries in a season. Across, say, ten thousand player-sessions in that season, twenty are followed by an injury. That's a rate of 0.2%.

Now imagine a model that's 95% accurate. Sounds excellent. Apply it to ten thousand sessions and it correctly identifies 19 of the 20 injuries — and it also produces roughly 500 false positives.

So of the players flagged as high risk, about 3.6% actually get injured. Which means a coach following the model would rest twenty-six players unnecessarily for every real injury prevented.

This isn't a flaw in any particular model. It's a mathematical consequence of predicting rare events, and it applies regardless of how good the underlying science is.

The validation problem

Many published injury prediction results are developed and evaluated on the same dataset, or on datasets from the same club and season. Performance in those conditions is not evidence that the model works.

With enough candidate variables and a small number of injury events, you will find combinations that separate the injured from the uninjured in your data. Those combinations frequently reflect the idiosyncrasies of that squad, that season, that measurement setup — not a generalisable relationship.

When these models are tested on a different club or a different season, performance typically collapses towards chance. This has been demonstrated repeatedly and it's the single most common failure mode.

The fix is external validation — develop on one dataset, test on a genuinely independent one — and it's done far less often than it should be, partly because independent datasets are hard to obtain and partly because the results are discouraging.

The causality problem

Even a model that predicts well doesn't tell you what to do.

Suppose high sprint volume predicts hamstring injuries. Should you reduce sprint volume? Not necessarily — players with higher chronic sprint exposure may be better conditioned and more robust. The correlation could run either direction and the observational data can't distinguish them.

Worse, the data is contaminated by intervention. Staff already act on their judgement. A player who looked fatigued was rested, so he didn't get injured, so the data records "high fatigue, no injury." The model learns that fatigue is safe, because the dangerous cases were prevented.

This is a general problem with learning from operational data in any domain where humans are already intervening, and it's severe in this one.

The measurement problem

The inputs are also weaker than they look.

GPS-derived metrics, as discussed at length elsewhere, are provider-dependent and definitionally unstable. Wellness questionnaires are self-reported and subject to obvious distortions — players who want to play report feeling fine. Sleep data from wearables has significant measurement error against polysomnography.

And many relevant factors aren't captured at all. Previous injury history is recorded inconsistently. Psychological stress, life circumstances, nutritional status, genetic predisposition — mostly absent.

A model built on noisy inputs, missing the largest causes, predicting a rare event, validated on itself. It's not surprising that the results don't hold.

What does work

Some things have reasonable evidence and they're mostly not predictive models.

Eccentric strengthening. Specific exercises targeting the hamstrings have reasonably strong evidence for reducing injury rates. This is an intervention that works regardless of who you apply it to, which sidesteps the prediction problem entirely.

Gradual load progression. Avoiding sudden large increases in training volume. Simple, well-supported, and requires no model.

Adequate recovery between matches. Fixture congestion is associated with elevated injury rates in multiple studies. This is a scheduling problem more than a monitoring one.

Screening after previous injury. Previous injury is the strongest single predictor of future injury, and it's known without any model. Managing return-to-play properly matters more than predicting the next one.

A better framing

The productive shift I've seen in practitioner thinking is away from prediction and towards monitoring for change.

Rather than asking "will this player get injured," ask "has something about this player changed." A drop in a repeated measure, a deviation from an individual's own established pattern, a subjective report that doesn't match the objective data.

These don't predict anything. They prompt a human to look, and a physiotherapist who knows the player can then assess whether the change means something.

That's a much more modest use of the data and it's the one that seems to hold up. It's also much harder to sell, since the product is "a slightly better-informed conversation" rather than "a risk score." Which is probably why the field keeps producing the other thing.