Football datasets contain the players who made it far enough to be recorded. That filtering shapes every conclusion drawn from them, usually without acknowledgement.
The data only covers those who were chosen
A player appears in match data because a coach selected him, which means selection has already occurred before any measurement is taken. The dataset begins after the most consequential filter has been applied.
Players who were equally capable but not selected leave no record, so the dataset cannot say anything about them.
Every analysis of professional football is therefore conducted on a group that passed through several filters first.
Academy studies are the clearest case
Research on youth development typically follows players who reached a certain level, because those are the ones whose records exist. Records for players released at fourteen are rarely kept in any usable form.
Traits common among that group are then presented as predictors of success, when they may simply be traits that selectors favoured.
The relative age effect is the best-known example, where birth month predicts selection far more than it predicts ability. Older children within an age group are physically ahead, and that advantage is mistaken for talent early enough to become self-fulfilling.
Transfer analysis inherits the same problem
Studies of what makes a signing succeed use players who were actually signed, and those players were chosen because someone expected them to succeed. The rejected alternatives generate no outcome data at all.
The characteristics identified therefore describe recruiter preferences at least as much as they describe drivers of performance.
A model trained on that record learns to imitate past decision-making rather than to improve on it.
Attrition compounds over a career
Players who struggle physically leave the sample through injury, and those who struggle competitively leave through release or demotion. Both exits are gradual, so the filtering happens continuously rather than at a visible moment.
Long-run career data consequently describes an increasingly select group, which makes durability and performance look more connected than they are.
Ageing curves built this way understate decline, because the players who declined most sharply have already dropped out of the dataset.
Partial remedies exist and are rarely used
Following complete intake cohorts regardless of outcome removes much of the bias but requires long and expensive data collection. The players who leave the game are precisely the ones hardest to keep track of.
Statistical corrections for selection exist, though they need assumptions about the selection process that are hard to justify in football.
The most honest practical step is stating clearly which population a finding applies to, which is cheap and still uncommon.