Top at halfway. Who'll win the league? From beliefs to predictions
Take what we believe about every team, play the rest of the season thousands of times, and count who wins. Tested on 25 SPFL title races, the simulation's favourite won 20; the side top at halfway won 17.
Intermediate Part 5 of Bayesian Thinking Through Football
Contents
The football question
Halfway through 2025/26, Hearts were top of the Scottish Premiership: 41 points to Celtic's 38. Who was going to win the league?
The table says Hearts. But a table is a snapshot of what's happened, not a forecast of what will. To forecast, you need two things the table doesn't have: a belief about how good each team really is, and the fixtures still to play.
This part puts the Bayesian series together. Bayes' theorem showed how evidence updates a belief, and updating a team's scoring rate built those beliefs for every team, match by match. Now we turn beliefs into predictions.
The concept
The method is called simulation, or Monte Carlo after the casino:
- Believe: for every team, a Bayesian estimate of how many goals it scores and concedes, starting from last season and updated with every match so far.
- Allow for doubt: we're not certain of those estimates, so each run of the season draws a plausible strength for every team from the range the evidence allows. A team with little evidence gets a wider range.
- Play: for every remaining fixture, draw the goals from a Poisson distribution using the two teams' strengths and home advantage.
- Count: finish the table, note who's champion, and do it all again. Thousands of times.
The share of runs a team wins the league is its chance of the title. Nothing more mysterious than that.
A football example
Stop 2025/26 at halfway, the first 114 of its 228 matches, and play the rest 10,000 times:
Hearts led, but the simulation made Celtic favourites at 83%. Three reasons, all Bayesian:
- The prior. Celtic had been far stronger the season before, champions with 92 points to Hearts' 52 in sixth; half a season of Hearts outscoring expectations doesn't wipe that out, as in updating a team's scoring rate.
- The game in hand. Celtic had played 18 matches to Hearts' 19. The simulation plays every remaining fixture, so it counts games in hand automatically; the table doesn't.
- The doubt. Nobody's strength is known exactly, so some runs make Hearts better than their record suggests and some make Celtic worse. Hearts' 14% is the share of runs where that doubt, and the luck of the second half, went their way.
Celtic won the league.
Tested on 25 seasons
One season could be luck. So stop every season from 2001/02 to 2025/26 at halfway, simulate the rest, and compare with what happened:
- The simulation's favourite won the title in 20 of 25 seasons.
- The side top of the table at halfway won it in 17 of 25.
And when the simulation gave a side a given chance, did they win about that often?
| The simulation said | Teams | Average chance | Won the title |
|---|---|---|---|
| Under 10% | 265 | 0.2% | 0 |
| 10% to 50% | 10 | 30.7% | 5 |
| 50% to 90% | 11 | 70.8% | 6 |
| 90% and over | 14 | 97.4% | 14 |
The very likely champions all won, and the no-hopers never did. In between, the numbers are small: ten or eleven teams per row, so 5 of 10 against an expected 3 is well within luck. But the 50–90% row hints that the simulation was a little too sure of itself in close races. This check, whether 70% means 70%, is calibration, one of the ways of scoring forecasts in is accuracy the right score?
Races worth a closer look
| Season | Leader | Favourite | Won |
|---|---|---|---|
| 02/03* | Rangers | Celtic 56% | Rangers |
| 04/05* | Celtic | Celtic 68% | Rangers |
| 08/09 | Celtic | Celtic 60% | Rangers |
| 10/11 | Celtic | Rangers 56% | Rangers |
| 11/12 | Rangers | Rangers 67% | Celtic |
| 14/15 | Aberdeen | Celtic 92% | Celtic |
| 21/22 | Rangers | Rangers 89% | Celtic |
| 25/26 | Hearts | Celtic 82% | Celtic |
Leader: top at halfway. Favourite: the simulation's. * Won on the last day of the season.
Some of its misses were close races it called close: 56% and 60% favourites lose a lot. 2010/11 shows it at its best: Celtic were eight points clear, but Rangers had four games in hand, and the simulation made them favourites; they won. 2011/12 shows its blind spot: it knows results, not finances. Rangers went into administration mid-season and were docked ten points. And 2021/22 was simply a big miss: an 89% favourite that didn't win.
Why it matters
- A table isn't a forecast. It ignores games in hand, the strength of the fixtures to come and how much of each team's record is luck.
- Honest uncertainty is part of the answer. "Celtic 83%, Hearts 14%" is more useful than "Celtic will win", and it's testable.
- This is how the pros do it. Title odds, relegation odds and every "supercomputer predicts the table" headline come from simulations like this, with fancier strength ratings.
- Beliefs, evidence, predictions. The whole Bayesian series in one method: start from a prior, update on the evidence, and let the uncertainty flow through to the forecast.
Limitations
- It knows only goals. No injuries, transfers, managers or finances; 2011/12 shows what that costs.
- No points deductions. The simulated tables use results only, so deducted points aren't included. That's why this part sticks to the title: some final placings in these years were changed by deductions, such as Hearts' 15 points in 2013/14.
- 2019/20 was cut short by COVID after 179 of its 228 matches. It's included, with its "halfway" being halfway through the matches actually played.
- The split is simplified. After 33 rounds the Scottish Premiership splits into top and bottom six. The simulation plays the fixtures that actually happened; had earlier results gone differently, the split fixtures would have too.
- Close races are hard. The 50–90% row suggests the simulation is a little too confident when two sides are close; a stronger model would help.
Try it yourself
Take your league at halfway. Give each team a rough chance of winning each remaining home match, drawing and losing, then use dice or a spreadsheet's random numbers to play out the rest of the season five times. How often does the current leader finish top?
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. The simulations use a fixed random seed, so you'll get the same numbers; it takes about fifteen seconds. Then:
import csv
import random
from collections import Counter
from datetime import datetime
from math import exp
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
def season(s): # (home, away, home goals, away goals) in date order
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
return [(r["HomeTeam"], r["AwayTeam"], int(r["FTHG"]), int(r["FTAG"])) for r in games]
data = {s: season(s) for s in names}
def totals(games): # goals for, goals against and matches, per team
gf, ga, n = Counter(), Counter(), Counter()
for h, a, x, y in games:
gf[h] += x; ga[h] += y; n[h] += 1
gf[a] += y; ga[a] += x; n[a] += 1
return gf, ga, n
def table(games, teams): # points, then goal difference, then goals scored (no points deductions)
pts, gd, gs = Counter({t: 0 for t in teams}), Counter(), Counter()
for h, a, x, y in games:
pts[h] += 3 * (x > y) + (x == y); pts[a] += 3 * (y > x) + (x == y)
gd[h] += x - y; gd[a] += y - x; gs[h] += x; gs[a] += y
return sorted(teams, key=lambda t: (pts[t], gd[t], gs[t]), reverse=True), pts
def poisson(lam, rng):
limit, k, p = exp(-lam), 0, rng.random()
while p > limit:
k += 1
p *= rng.random()
return k
def title_chances(s, sims, k=20, seed=1):
"""Stop season s at halfway, then play its remaining fixtures sims times."""
rng = random.Random(seed)
games, last = data[s], data[names[names.index(s) - 1]]
teams = sorted({t for g in games for t in g[:2]})
lf, la, ln = totals(last)
mu = sum(x + y for *_, x, y in last) / (2 * len(last)) # goals per team per match last season
home = sum(x for *_, x, _ in last) / len(last) / mu
away = sum(y for *_, y in last) / len(last) / mu
played, rest = games[:len(games) // 2], games[len(games) // 2:]
gf, ga, n = totals(played)
# priors from last season, worth k matches; promoted sides score less and concede more than average
pf = {t: lf[t] / ln[t] if ln[t] else 0.85 * mu for t in teams}
pa = {t: la[t] / ln[t] if ln[t] else 1.15 * mu for t in teams}
wins = Counter()
for _ in range(sims):
# we aren't sure how good each team is, so each run draws a plausible attack and defence from the Gamma posterior
att = {t: rng.gammavariate(k * pf[t] + gf[t], 1 / (k + n[t])) for t in teams}
dfn = {t: rng.gammavariate(k * pa[t] + ga[t], 1 / (k + n[t])) for t in teams}
sim = [(h, a, poisson(att[h] * dfn[a] / mu * home, rng), poisson(att[a] * dfn[h] / mu * away, rng)) for h, a, *_ in rest]
wins[table(played + sim, teams)[0][0]] += 1
now, pts = table(played, teams)
return {t: wins[t] / sims for t in teams}, now, pts, table(games, teams)[0][0]
said, fav_right, leader_right = [], 0, 0
for s in names[1:]:
chance, now, pts, champion = title_chances(s, 2000)
favourite = max(chance, key=chance.get)
fav_right += favourite == champion
leader_right += now[0] == champion
said += [(p, t == champion) for t, p in chance.items()]
print(f"20{s[:2]}/{s[2:]}: top at halfway {now[0]}, model's favourite {favourite} {chance[favourite]:.0%}, champions {champion}")
print(f"title won by the model's favourite {fav_right} times, by the side top at halfway {leader_right} times, of {len(names) - 1}")
for lo, hi in ((0, 0.1), (0.1, 0.5), (0.5, 0.9), (0.9, 1.01)):
band = [(p, w) for p, w in said if lo <= p < hi]
print(f"model said {lo:.0%}-{min(hi, 1):.0%}: {len(band)} teams, average {sum(p for p, _ in band) / len(band):.1%}, won {sum(w for _, w in band)}")
chance, now, pts, champion = title_chances("2526", 10000)
print("2025/26 at halfway:", [(t, pts[t]) for t in now[:4]], {t: f"{chance[t]:.1%}" for t in now[:4]})
Further reading
- Monte Carlo method, Wikipedia. Simulating something many times to estimate a probability, and where the name comes from.
- Posterior predictive distribution, Wikipedia. Predicting new data while allowing for uncertainty in the parameters, as each simulated season does.
- Poisson regression, Wikipedia. The standard way to model goals with attack, defence and home advantage.