Football prediction models are commonly evaluated by counting correct results. That test is intuitive and it selects for the wrong properties.

Accuracy ignores how uncertain a match is

Counting correct calls treats a foregone conclusion and a genuinely balanced fixture identically. Both contribute one unit to the score, and the difficulty of the prediction appears nowhere in the measure.

A model that predicts the stronger side in every match will score respectably by this measure while carrying no information beyond the table.

The measure therefore cannot distinguish a useful model from a restatement of what everyone already assumed.

Draws break the accuracy measure entirely

A draw is rarely the single most likely outcome even in matches between evenly matched sides. A model maximising correct calls will therefore almost never predict one, and the three-way structure of football results makes this unavoidable rather than a quirk of any particular method.

Since draws occur regularly in football, an accuracy-maximising model systematically ignores a substantial part of what actually happens.

Any evaluation that pushes models away from predicting a common outcome is measuring the wrong thing.

Proper scoring rules reward honest probabilities

Scoring rules that assess the full probability distribution penalise a model for being confidently wrong and reward it for being appropriately uncertain. The penalty grows steeply with confidence, which is what stops a model from bluffing.

Under those rules, stating a middling probability for a coin-flip fixture is the correct behaviour rather than a failure to commit.

Models compared this way sort very differently from models compared on correct calls, and the ordering is the more informative one.

Calibration is the practical test

A calibrated model is one where matches assigned a given probability resolve at roughly that rate over a long run. Grouping past predictions by their stated probability and comparing against observed frequency is a straightforward diagnostic.

Calibration failures are usually asymmetric, with models overconfident about strong favourites and underconfident about the rest.

Checking it requires many matches, which is why calibration is assessed across seasons rather than within one.

The audience prefers the wrong measure

A single predicted scoreline is easier to publish and to argue about than a distribution over outcomes. Public model comparison consequently defaults to counting hits, because a probability distribution does not fit a headline and cannot be settled by one weekend.

That incentive pushes published models toward confident statements they cannot support, because hedged output reads as evasion.

The models used inside clubs and by trading operations are judged differently, which is most of why they behave so differently from the ones in public.