Eleven goals in five games. Can they keep it up? Updating a team's scoring rate
Early-season goal averages swing wildly. Starting from last season and updating after every match, the Bayesian way, predicts the rest of the season about twice as well after five games. 25 SPFL seasons show how.
Intermediate Part 4 of Bayesian Thinking Through Football
Contents
The football question
Five games into the 2018/19 Scottish Premiership, Hearts had scored 11 goals: 2.2 a game. Were they really a 2.2-goals-a-game side?
The raw average said yes. Last season said no: they'd scored 1.03 a game. For the rest of the season they scored 0.94.
Early in a season every average is built on a handful of matches, and a handful of matches is mostly luck. The Bayesian answer is to start from what you knew before, and let each match move the estimate a little. Bayesian updating did this with a penalty taker, and Bayes' theorem with a half-time score. Here it runs all season.
The concept
Start each team's scoring rate at a prior: its goals per game last season. Treat that prior as if it were worth k matches. Then after every match, blend it with the goals scored so far:
$$\begin{aligned} &\text{estimate} \\ &= \frac{k \times \text{last season} + \text{goals so far}}{k + \text{matches so far}} \end{aligned}$$
In plain football
- Last season is the team's goals per game last season. k is how much we trust it, counted in matches. With k = 20, last season counts as twenty matches' worth of evidence.
- Goals so far and matches so far are this season's evidence.
- Early on, the prior dominates. By the end of the season, this season's matches outnumber it and take over. It's the same weighted average as Bayesian updating, applied to goals.
For Hearts after five games, with k = 20:
$$\frac{20 \times 1.03 + 11}{20 + 5} \approx 1.26$$
In plain football
- 20 × 1.03: last season, counted as twenty matches at 1.03 goals a game.
- + 11 goals in 5 matches: this season's hot start.
- 1.26: a bit higher than last season, but nowhere near 2.2. It was a lot closer to the 0.94 they managed for the rest of the season.
Teams promoted into the Premiership have no top-flight season to use, so they start at 85% of last season's league average: across these 25 seasons, that's about what promoted sides scored.
A football example
Here's Hearts' 2018/19, match by match:
One team could be luck. So test it on every team in every season from 2001/02 to 2025/26. After a few matches, estimate each team's scoring rate, then compare with what it actually scored for the rest of that season:
| After | Raw average is off by | Bayesian (k = 20) is off by |
|---|---|---|
| 3 matches | 0.52 goals a game | 0.22 |
| 5 matches | 0.41 | 0.22 |
| 10 matches | 0.32 | 0.22 |
| 19 matches | 0.30 | 0.26 |
After five matches, the Bayesian estimate is about twice as close. The gap narrows as the season's evidence piles up, but even at the halfway mark, last season still helps.
How much is last season worth?
The prior's weight, k, is a choice. Try several:
| Last season counts as | Off by, after 5 matches |
|---|---|
| 0 matches (the raw average) | 0.41 |
| 5 matches | 0.26 |
| 10 matches | 0.23 |
| 20 matches | 0.22 |
| 40 matches | 0.22 |
Anything from about 10 matches upwards does well; below that, the handful of new matches still shouts too loud. In the Scottish Premiership, last season's scoring rate is worth roughly half a season of new evidence. It's the same pull seen in where priors come from, where early league leaders drift back towards the pack.
Show the mathsWhy this is the Gamma–Poisson update. Optional.
Goals in a match are close to Poisson with rate λ. A Gamma prior for λ,
$$\lambda \sim \text{Gamma}(\alpha, \beta), \quad \text{mean} = \frac{\alpha}{\beta}$$
with \(\alpha = k \times \text{last season's rate}\) and \(\beta = k\), updates after n matches with G goals to
$$\lambda \mid \text{data} \sim \text{Gamma}(\alpha + G,\ \beta + n)$$
whose mean is \((\alpha + G) / (\beta + n)\): exactly the weighted average above. Gamma and Poisson fit together so neatly that the update is just two additions, which is why this pairing is a workhorse of football models.
Why it matters
- Early-season tables and averages mislead. Five matches of goals is mostly noise; a sensible model leans on last season until this one has said enough.
- The update never stops. Each match nudges the estimate. Nothing needs recalculating from scratch, just a running total.
- It predicts better. Not by a fancy model, but by refusing to overreact to a hot or cold start.
- The same idea runs through football models. Team strengths in prediction models are updated this way, match by match, from a prior.
Limitations
- Last season isn't always relevant. A new manager, a big summer of signings or a sale of the top scorer can make last season a poor prior; a club insider would set it differently.
- Goals scored depend on opponents. A side that played Celtic and Rangers early looks worse than it is. Real models adjust for who each team has played.
- k was chosen with hindsight. It was picked on the same 25 seasons it's tested on. The flat results between 10 and 40 suggest the exact choice matters little.
- "Rest of the season" is noisy too. It's the best available yardstick for a team's true rate, not the truth itself.
Try it yourself
After your team's first five league matches, work out two numbers: goals per game so far, and (20 × last season's goals per game + goals so far) ÷ 25. Write both down. At the end of the season, which was closer?
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:
import csv
from collections import defaultdict
from datetime import datetime
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
def goals(s): # each team's goals scored, match by match, in date order
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
out = defaultdict(list)
for r in games:
out[r["HomeTeam"]].append(int(r["FTHG"]))
out[r["AwayTeam"]].append(int(r["FTAG"]))
return out
data = {s: goals(s) for s in names}
def prior(s, team): # last season's goals per game; a promoted side gets 85% of last season's league average
last = data[names[names.index(s) - 1]]
if team in last:
return sum(last[team]) / len(last[team])
return 0.85 * sum(map(sum, last.values())) / sum(map(len, last.values()))
def bayes(p, scored, k): # the Gamma-Poisson update: a prior worth k matches, plus the goals so far
return (k * p + sum(scored)) / (k + len(scored))
# how far each estimate, after m matches, is from the team's scoring rate over the rest of the season
for k in (0, 5, 10, 20, 40):
line = f"prior worth {k:2} matches:"
for m in (3, 5, 10, 19):
off = [abs(bayes(prior(s, t), g[:m], k) - sum(g[m:]) / len(g[m:])) for s in names[1:] for t, g in data[s].items()]
line += f" after {m:2}: off by {sum(off) / len(off):.2f}"
print(line) # k = 0 is the raw average so far
g, p = data["1819"]["Hearts"], prior("1819", "Hearts")
print(f"Hearts 2018/19: prior {p:.2f}; after 5 matches raw {sum(g[:5]) / 5:.2f}, Bayes {bayes(p, g[:5], 20):.2f}; rest of season {sum(g[5:]) / len(g[5:]):.2f}")
print("raw ", [round(sum(g[:m]) / m, 2) for m in range(1, len(g) + 1)])
print("Bayes", [round(bayes(p, g[:m], 20), 2) for m in range(1, len(g) + 1)])
Further reading
- Conjugate prior, Wikipedia. Why some priors, like Gamma for a Poisson rate, update with simple additions; with a table of the common pairs.
- Gamma distribution, Wikipedia. The shape and the mean α/β of the prior used here.
- Shrinkage (statistics), Wikipedia. Pulling noisy estimates towards a sensible centre, and why it improves predictions.