Reference

Expected Goals (xG): What It Measures and What It Doesn't

Expected goals, usually written xG, is the most quoted number in modern football analysis and one of the most misread. This page sets out what the metric does measure and where its limits sit — a method note, not a verdict on any team or match.

What xG measures

One attempt at a time

Every attempt on goal carries a probability of being converted. Expected goals is the practice of attaching that probability to each attempt. A model looks at a shot — where it was taken from, how it was struck, how it came about — and asks how often shots of that kind have been converted in the past. That share is the attempt’s xG: a number between 0 and 1, from nearly always missed to nearly always converted.

What the model looks at

The exact inputs vary between models, but the core set is well understood. Distance from goal does most of the work: a strike from distance and the same strike from close range are two very different attempts. The angle to goal narrows the target as the shooter moves towards the byline. Headers are converted less often than shots struck with the foot. The build-up enters too: a cut-back creates a different chance from a long ball chased down, and a set piece differs from open play. Most models also carry some representation of the defensive pressure the attempt met — how many defenders stood between the ball and the goal, and how close.

From attempts to a match total

Add the xG of every attempt a team took in a match and the result is that team’s expected goals for the match — a summary of the chances the team created, in volume and in quality. That is the convenience and the limitation: a team that took many thin attempts and a team that carved out a few clear chances can arrive at similar totals. The sum describes what was attempted. It is not a re-run of the match, and it does not claim the same chances would arise again.

Why analysts use it

Goals are rare and noisy. A match finishes with a small handful of them, and the gap between a goal and a near miss can be a deflection, a bounce, a fingertip. Reading form from raw goal tallies therefore means reading noise. Chance quality builds a fuller sample: every attempt contributes, not only the ones that went in, so the picture of how a team is playing accumulates faster than the tally does. Over a run of matches, xG settles towards a team’s underlying level sooner than the goal tally it approximates. How long that takes depends on the team, the model and the noise around both; this page quotes no figure for it, because none is settled.

What xG does not measure

The number belongs to a model

xG is not a physical quantity that anyone observes directly. It is the output of a model fitted to historical shots, and models disagree. One provider, trained on its own events data with its own features, will attach one probability to an attempt; another model, built differently, will attach another. The disagreement is usually small on a single shot and can be meaningful over a match. A quoted xG figure is only meaningful alongside the name of the model that produced it. When two sources give different totals for the same match, that is two estimates from two tools, not a contradiction.

A few matches say nothing about finishing

Over a short run, the gap between a team’s goals and its xG is dominated by variance. Attempts are converted in clumps; a finisher can take his chances in three consecutive matches and miss them in the next three with nothing having changed. Attributing a small-sample gap to a striker being clinical, or wasteful, is a claim the metric does not support at that sample size. Over long enough horizons finishing skill does differ between players; the gap over a handful of matches is mostly noise, and noise tells no stories about anyone.

It only sees attempts

The metric counts attempts. Everything that happens before an attempt exists, and everything that prevents one, is invisible to it. A spell of possession that never produced a shot counts for nothing. A defence so well organised that attempts never materialised counts for nothing. A chance passed up in favour of a safer ball counts for nothing. A team can therefore be described flatteringly — its few attempts were of high quality — while being second-best across the rest of the match, or unflatteringly for reasons that have nothing to do with how well it moved the ball. The number sees the ends of moves, not the moves.

The scoreline is flattened out

Attempts do not take place in neutral conditions. A team that is ahead tends to sit deeper and concede territory, accepting attempts against it in exchange for protecting the lead. A team that is behind pushes forward and carves out attempts of declining quality. A sending-off reshapes a match entirely: a team a player short attacks less and defends more, whatever the plan was. xG records the attempts that took place, but not the situation that produced them, and it carries no memory of the conditions two similar totals were played under.

It is not a result

The word “expected” is where the trouble starts. Expected goals does not mean the match should have finished at those numbers, and it is not a finding that one team deserved more. Summing shot probabilities yields an average; around that average sits a wide distribution of outcomes, because each attempt is an event that either happens or does not. A team that out-creates its opponent and loses is not the victim of an injustice the metric has detected. That is an ordinary event. It takes place most weeks, in both directions, and the metric was never a promise otherwise.

Where that leaves it

Used honestly, expected goals is one input among several. It describes a process over a run of matches — how a team creates chances, and how it gives them up — more cleanly than a goal tally does. As a verdict on a single match it is weak, for every reason above. And it is never a statement about what will happen next: the next match is a fresh set of attempts, each carrying its own probability, and the spread of outcomes around the average is wide enough to embarrass any certainty.

LeagueQuant and forecasts

LeagueQuant publishes method notes like this one. It publishes no forecasts yet and holds no accuracy record, because there is nothing yet to record. When live forecasts begin, each one will be published with a hash recorded before kick-off, so its contents cannot be quietly rewritten afterwards, and each will be evaluated automatically once the match has finished, with misses in the record alongside everything else; when it exists, that record will have been built in public, from a standing start.