Ask which leagues have good data coverage and you'll get a familiar list. The big five European leagues, obviously. The major second tiers. MLS, Brazil, Argentina, Portugal, Netherlands, Belgium, a handful of others. Perhaps fifty competitions with detailed event data and a smaller number with tracking.
There are thousands of professional and semi-professional leagues in the world. The gap between what's covered and what exists is enormous, and the shape of that gap has consequences.
What determines coverage
Data collection is a commercial activity. Providers cover leagues where somebody will pay for the data, which correlates with broadcast value, betting volume, and the presence of clubs with analytics budgets.
That produces a coverage map that looks a lot like a map of football money. Western Europe, dense. North America, good. South America's major leagues, reasonable. Large parts of Africa and Asia, thin to nonexistent. Lower divisions almost everywhere, absent.
There's also a practical constraint. Event data collection requires reliable video, and reliable video requires broadcast infrastructure. Leagues without consistent televised coverage can't be tracked at all by conventional methods, regardless of demand.
Who this makes invisible
The obvious answer is players in uncovered leagues, but the specific pattern is worth stating.
Players in African domestic leagues, which have produced an enormous number of elite footballers and have almost no systematic data. A player at a club in Ghana or Ivory Coast is assessed essentially the way players were assessed in 1985 — somebody watches him, forms an opinion, and tells somebody else.
Players in most Asian leagues below the top division of the largest countries. Players in Eastern European lower tiers. Players in Central American football. And late developers everywhere, because they're most likely to be playing at levels nobody covers when they're at the age where a club might sign them cheaply.
The pattern is that data coverage tracks existing wealth, and therefore analytics reinforces existing recruitment patterns rather than disrupting them.
The consequence for the market
An interesting economic effect. Players in covered leagues are assessed by many clubs using similar tools, so their valuations converge and their prices reflect a market consensus.
Players in uncovered leagues are assessed by whoever happens to have a scout there, which means fewer bidders, more variance in valuation, and a genuine possibility of finding somebody excellent cheaply.
So the data blind spot is simultaneously a disadvantage for the players in it — fewer opportunities to be noticed — and an opportunity for clubs willing to invest in traditional scouting where the data isn't.
Several clubs have built recruitment models around exactly this, maintaining human scouting networks in regions everyone else has withdrawn from. It's a deliberately contrarian strategy and it seems to work.
What's changing
Broadcast-derived computer vision tracking is the main development, and in principle it removes the installation barrier — you need footage, not infrastructure.
In practice the improvement is partial. The algorithms need reasonably stable camera work, visible pitch markings, and adequate resolution. Footage from a single fixed camera in poor light on a worn pitch is exactly the hardest case, and it's the typical case in the leagues that need it most.
There's been progress on more robust methods and on cheap dedicated camera systems — a couple of fixed cameras and an edge processing unit, at a cost that a club in a developing league might plausibly afford. Some federations have deployed these.
Whether it scales depends on somebody having a commercial reason to pay for it, and the reason has to come from outside those leagues, because the clubs in them mostly can't fund it.
The quality question
One complication worth being honest about. Even with data, comparing across leagues is genuinely difficult.
A player producing excellent numbers in a weak league is producing them against weak opposition. Adjusting for league strength requires estimating relative quality, which is done through the small number of matches between leagues and through player transfers, and both sources are thin and biased.
League strength coefficients exist and they're better than nothing. They're also uncertain enough that a player's projected performance after a large step up has a very wide confidence interval, which is a polite way of saying that these transfers remain a gamble whatever the data says.
Why it's worth caring about
Beyond the recruitment economics, there's a fairness dimension that I don't think gets discussed enough.
A talented seventeen-year-old in a league with data coverage is discoverable by any club in the world with a subscription. An equally talented seventeen-year-old two thousand kilometres away is discoverable only if somebody physically goes to look.
That's a substantial difference in opportunity, driven entirely by where somebody was born, and it's getting larger as scouting shifts towards data. Whatever else football analytics has done, it hasn't flattened that particular hierarchy. If anything it's steepened it.