A new football metric is usually defended by showing that it correlates with winning. That test passes far too easily to distinguish a good measure from a poor one.

Winning teams do almost everything more

Successful sides tend to have more possession, take more shots, complete more passes and win more duels. Any statistic tracking those activities will correlate with results.

The correlation therefore says that the metric is football-shaped rather than that it captures something useful.

Passing the test is close to automatic, which is why nearly every proposed metric passes it.

The causal direction is usually backwards

Teams that go ahead change behaviour, sitting deeper and conceding possession. A statistic measured after that point reflects the scoreline rather than causing it.

Late-match data is particularly contaminated, since the state of the game dominates what both sides are trying to do.

Correlations calculated across full matches mix the effect of performance on results with the effect of results on performance.

Incremental value is the real question

The useful test is whether a metric adds anything once existing measures are accounted for. A statistic that correlates with winning only because it correlates with shot volume contributes nothing new.

Evaluating a metric alongside established ones rather than in isolation answers that question directly.

Many metrics that look impressive alone add almost nothing under this test, which is why it is applied less often than it should be.

Prediction is a stronger standard than description

Explaining past results is easy with enough measures, since any sufficiently flexible combination can be fitted to what already happened.

Predicting future results from past data is much harder and is where weak metrics fail. Holding back a period of data for testing is the standard defence.

Metrics developed and evaluated on the same seasons should be treated as untested regardless of how well they fit.

Stability determines whether a metric is usable

A measure that varies wildly between halves of a season cannot support a decision about a player, whatever its correlation with results.

Splitting the data and checking whether a player's value in one half predicts the other is a simple test that eliminates a large share of proposed measures.

The metrics that have lasted in football all clear that bar, and the ones that appeared briefly and vanished mostly did not.