Dividing football stats by possession made player comparisons worse, not better
A data analyst built a vector space model of 1,419 footballers using 5.3 million match events to compare player roles, but found that raw statistics were heavily skewed by how much possession each player's team held. An initial attempt to correct this by dividing each metric by team possession share actually worsened the data, because the relationship between a metric and possession is affine — it has an intercept — making division mathematically incorrect. The proper fix was to subtract the fitted linear trend and keep only the residual, which reduced the possession-correlation score from 0.695 to 0.450, compared to a chance baseline of 0.000. The analysis also revealed that even ratio-based metrics like pass completion percentage are significantly contaminated by possession, disproving the assumption that ratios are inherently immune. The key lesson drawn is that any normalisation method should be measured against the specific defect it is meant to remove, not a proxy metric.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in