Buy football data and you are almost certainly buying it secondhand. Only a handful of companies actually watch the matches and write down what happens. Everyone else resells that work, one or more steps down a chain, and a little of the truth leaks out at every step. Knowing where your numbers sit on that chain is the difference between a signing you can defend and one you cannot.
So here is the chain, layer by layer, with no ranking attached. What matters is not which logo you recognise. It is how far that logo sits from the pitch.
Primary collectors
At the top are the companies that pay people to code matches in real time, at the ground or off the broadcast feed, and hold the exclusive league rights that let them do it. Sportradar, Stats Perform (whose event brand is Opta), Genius Sports and Hudl are the names that matter for non-betting data. Everyone below them is, in one form or another, a customer.
This is also where consolidation bites. When two or three houses own the rights to a competition, the price of original data only moves one way, and the clubs furthest from that money feel it first.
Public aggregators
One step down sit the apps you already know. Sofascore, Flashscore and 365Scores license from the collectors and pour the data into consumer products, often with community editors patching the gaps. Useful, fast, free. Not yours to resell, though. Most of them forbid it outright in their terms, which matters more than it sounds, because this is the layer a surprising amount of “professional” data is quietly scraped from. When a feed looks too cheap for what it covers, this is usually why.
Affordable REST APIs
Then come the subscription APIs, the ones with a developer endpoint and a price that undercuts the source many times over. They collect nothing. They wrap data that came from higher up the chain, and a licence here can be a real bargain or a real liability, depending entirely on where that data came from. So ask before you sign. What is your source, and do you have an agreement with them? Some of these vendors resell scraped data with no such agreement. That feed is legally shaky, and it can disappear the week after you have built your season around it.
Open and web-first sources
A different corner of the map gives data away on the open web. FBref, backed by StatsBomb, plus Understat and FootyStats, are the names analysts cut their teeth on. Brilliant for study and useless for a product, because the licences almost never allow commercial reuse, and scraping the sites directly tends to breach their terms too. Learn from them. Do not build on them.
Event and on-ball specialists
This is the layer scouting actually runs on: detailed, hand-coded event data. StatsBomb, Wyscout, InStat and Opta define it. The important thing is that they are collectors in their own right, not resellers. They watch the matches and code them, which is exactly why their data runs richer than anything downstream of it.
Two of those names now answer to one owner. Hudl bought Wyscout in 2019 and InStat in 2022, and Opta has long been the event brand of Stats Perform. The field of independents is closing into a few large houses. Good for consistency, and rarely good for price.
Manually coded?
Coded by whom, exactly? By people. These vendors employ hundreds, sometimes thousands, of analysts who watch matches and tag every event by hand. AI is taking more of that work each season, but a great deal of it is still done by eye, because that is still what keeps the numbers honest.
Tracking and physical data
A separate map measures movement rather than events. Some of the collectors appear here again, alongside specialists you meet nowhere else: Metrica Sports, SkillCorner, Catapult, IMPECT, Track160, Eyeball and PlayerData. Some read the game from video, some from GPS vests, a few from both. They answer the questions event data cannot. How far a player ran, how often he sprinted, and where he was standing when the ball was somewhere else.
Reading the map
One rule sits over all of it: quality erodes with distance from the source. Every hop down the chain is a chance for an event to be miscoded, a gap to be patched by guesswork, or a player ID to drift out of alignment. For a scoreline, none of that matters. For the metric a recruitment call rests on, the distance is the risk.
There is a simple tell. If two cheap APIs carry the exact same wrong nationality for the same player, or the same hole for the same fixture, they are drinking from one upstream well. Two vendors, one source. Working out where your data actually comes from is worth more than any feature comparison, and it is the one thing no vendor’s own page will ever tell you.
So the landscape is not just crowded. It is divided. The same match can reach you through half a dozen vendors at half a dozen prices and half a dozen levels of quality, and some of those routes are not merely unreliable but legally grey. Where you get your data decides two things at once: whether you can trust it, and whether you are even allowed to use it.
Keeping all of that straight is a job on its own. Most clubs end up with a primary feed, a cheaper API to fill the holes, a tracking provider, a stack of spreadsheets and their scouts’ own notes, then spend longer reconciling it than scouting with it. That is one of the problems we built gaffer to solve. Pull the scattered sources into one place you can search, keep the trail of where every number came from, and stay on the right side of the licences you hold. Less time wrestling data. More time deciding who to sign.