Skip to content

Home, draw or away, with probabilities. A first real model, logistic regression

Three numbers known before kick-off, one model trained by gradient descent, and a probability for every result. Tested on five SPFL seasons it beats every model in the series so far and gets close to the bookmakers.

Intermediate Part 8 of Machine Learning Through Football

Contents

The football question

So far this series has predicted results with lookup tables and simple rules. Can a proper model, trained from scratch, turn a few numbers known before kick-off into a probability for home win, draw and away win, and how close can it get to the bookmakers?

This part pulls the whole series together: features and targets to choose the inputs, training and test data to check it honestly, gradient descent to train it, and is accuracy the right score? to judge it.

The concept

The model is logistic regression, one of the most widely used models in data science. For every match it gives each result a score, and turns the three scores into probabilities.

The score for a home win is a starting value plus a weight times each feature:

$$\begin{aligned} \text{score}_{\text{home}} = \;&b_{\text{home}} + w_1 \times \text{form} \\ &+ w_2 \times \text{this season} \\ &+ w_3 \times \text{last season} \end{aligned}$$

In plain football

  • The features are three numbers for each match, each the home side's figure minus the away side's: points per game over the last five matches, over this season so far, and over last season.
  • The weights, w, say how much each feature matters; the starting value, b, is the score before any of them, which is where home advantage lives.
  • The away-win score has its own starting value and weights. The draw is the baseline: its score is always 0.

Then the scores become probabilities, each result's share of the total:

$$P(\text{home}) = \frac{e^{\text{score}_{\text{home}}}}{e^{\text{score}_{\text{home}}} + e^{0} + e^{\text{score}_{\text{away}}}}$$

In plain football

  • e to the power of a score turns any score, positive or negative, into a positive number.
  • Dividing by the total makes the three probabilities add up to 1. This step is called softmax.
  • A higher home-win score means a higher chance of a home win, at the expense of the draw and the away win.

Before training, each feature is put on the same scale, measured in standard deviations as in player similarity, so the weights can be compared with each other.

Training it

Training means finding the weights that give the lowest log loss on the training matches: every Scottish Premiership match from 2001/02 to 2020/21, 3,900 of them, leaving out each season's opening games until both sides have played five. The method is gradient descent, with six dials this time instead of one: work out which way each weight should move to lower the log loss, move them all a step, repeat.

Steps Training log loss
1 0.9981
10 0.9729
50 0.9712
200 0.9712

It's done in about fifty steps: this valley is smooth, and the bottom is easy to find.

What it learned

Start Form This season Last season
Home vs draw +0.52 +0.04 +0.37 +0.29
Away vs draw +0.18 −0.01 −0.21 −0.36
  • The starting values are home advantage. In an even match the home win starts ahead of the away win.
  • This season and last season both matter, with similar weight. A side that's been better this season and last is favoured, as you'd expect.
  • Recent form barely matters. Once the model knows a team's points per game over the season so far and last season, its last five results add almost nothing: weights of +0.04 and −0.01. That's the finding of does form matter more than underlying performance?, learned by a model on its own.

Here's what those weights mean, as the home side gets stronger on this season and last season together:

The fitted model's probabilities. In an even match it says home 43%, draw 26%, away 31%. The draw is most likely when the teams are evenly matched, but even there it never becomes the single most likely result.
Home side is Home win Draw Away win
2 sd weaker 9% 19% 72%
1 sd weaker 22% 25% 53%
Even 43% 26% 31%
1 sd stronger 66% 20% 14%
2 sd stronger 82% 13% 5%

Testing it

The real exam: the five seasons it never saw, 2021/22 to 2025/26, the same 990 matches used to score every model in is accuracy the right score?:

Model Accuracy Log loss Brier
Knows nothing (a third each) 47.2% 1.099 0.667
Base rates 47.2% 1.056 0.638
Form + table position (lookup) 53.5% 0.997 0.594
Logistic regression 54.8% 0.950 0.562
Bookmaker 56.2% 0.932 0.550

From "knows nothing" to the bookmaker, log loss falls by 0.167. The lookup table got about 60% of the way; logistic regression gets about 90% of the way, with three numbers per match. The bookmakers, with team news, injuries and a whole market behind their prices, are still ahead, but not by much.

It did better on the test seasons (0.950) than on training (0.971). That's not a warning sign: the most likely reason is that recent seasons, with two clubs so far ahead, have simply been easier to predict. The gap that would worry us, overfitting, runs the other way.

And, as is accuracy the right score? predicted, it never picks a draw: 644 home wins and 346 away wins, no draws. Its draw probabilities do their work in the log loss, not in the picks.

Try it yourself in the Logistic Regression Match Predictor: set the three gaps and watch the probabilities move.

Why it matters

  • This is what a real model looks like. Features, weights, a way to turn scores into probabilities, and training by gradient descent. Bigger models add more of each, not a different idea.
  • The weights are readable. Logistic regression tells you what it learned: here, that the season's evidence matters and last week's form barely does.
  • Probabilities beat picks. It improves accuracy by 1.3 points over the lookup table, but log loss by 0.047: most of what it adds is better-judged confidence.
  • Simple, well-chosen features go a long way. Three numbers took it most of the way to the bookmakers.

Limitations

  • Three features. No injuries, suspensions, transfers or manager changes, which is exactly where the bookmakers' edge comes from.
  • Straight lines only. Each feature adds to the score in a straight line; real effects can bend. More flexible models can capture that, at more risk of overfitting.
  • No teams by name. The model sees only the gaps between two sides, not who they are.
  • One league. Trained and tested on the Scottish Premiership only.

Try it yourself

For your team's next match, work out the three gaps: points per game over the last five, this season and last season, home side minus away side. Using the table of probabilities above as a rough guide, where does the match sit? Then compare your answer with the bookmakers' odds.

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It runs in a few seconds:

import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log, sqrt

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
FEATURES = ["recent form", "this season so far", "last season"]
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return games

def points_per_game(games):
    pts, n = Counter(), Counter()
    for r in games:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            n[team] += 1
    return {t: pts[t] / n[t] for t in n}

# three features for every match, each the home side's figure minus the away side's, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
    last = points_per_game(season(s_last))
    promoted = 0.85 * sum(last.values()) / len(last)
    history = defaultdict(list)
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(history[h]) >= 5 and len(history[a]) >= 5:
            rows.append((s, [
                (sum(history[h][-5:]) - sum(history[a][-5:])) / 5,                   # points a game, last five
                sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]),  # points a game this season
                last.get(h, promoted) - last.get(a, promoted),                          # points a game last season
            ], r["FTR"]))
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            history[team].append(p)

train = [r for r in rows if r[0] < "2122"]   # 2001/02-2020/21
test = [r for r in rows if r[0] >= "2122"]   # 2021/22-2025/26
k = len(FEATURES)
mean = [sum(x[j] for _, x, _ in train) / len(train) for j in range(k)]
sd = [sqrt(sum((x[j] - mean[j]) ** 2 for _, x, _ in train) / len(train)) for j in range(k)]
scale = lambda x: [(x[j] - mean[j]) / sd[j] for j in range(k)]  # in standard deviations, so the weights compare
train = [(scale(x), y) for _, x, y in train]
test = [(scale(x), y) for _, x, y in test]

def predict(w, x):  # a score for each result, turned into probabilities that add up to 1
    score = {c: w[c][0] + sum(wj * xj for wj, xj in zip(w[c][1:], x)) for c in RESULTS}
    e = {c: exp(v) for c, v in score.items()}
    return {c: e[c] / sum(e.values()) for c in RESULTS}

def log_loss(w, data):
    return sum(-log(predict(w, x)[y]) for x, y in data) / len(data)

# gradient descent on log loss; the draw is the baseline, so its weights stay at zero
w = {c: [0.0] * (k + 1) for c in RESULTS}
for step in range(1, 201):
    grad = {c: [0.0] * (k + 1) for c in "HA"}
    for x, y in train:
        p = predict(w, x)
        for c in "HA":
            miss = p[c] - (c == y)
            grad[c][0] += miss
            for j in range(k):
                grad[c][j + 1] += miss * x[j]
    for c in "HA":
        w[c] = [wj - 1.0 * gj / len(train) for wj, gj in zip(w[c], grad[c])]
    if step in (1, 10, 50, 200):
        print(f"step {step:3}: training log loss {log_loss(w, train):.4f}")

for name, c in (("home win", "H"), ("away win", "A")):
    print(f"{name} vs draw: start {w[c][0]:+.2f}, " + ", ".join(f"{f} {wj:+.2f}" for f, wj in zip(FEATURES, w[c][1:])))
picks, right, brier = Counter(), 0, 0
for x, y in test:
    p = predict(w, x)
    pick = max(RESULTS, key=p.get)
    picks[pick] += 1
    right += pick == y
    brier += sum((p[c] - (c == y)) ** 2 for c in RESULTS)
print(f"test, {len(test)} matches: accuracy {right / len(test):.1%}, log loss {log_loss(w, test):.3f}, Brier {brier / len(test):.3f}, picks {dict(picks)}")
for gap in (-2, -1, 0, 1, 2):  # this season and last season both `gap` standard deviations in the home side's favour
    p = predict(w, [0, gap, gap])
    print(f"home side {gap:+} sd stronger: home {p['H']:.0%}, draw {p['D']:.0%}, away {p['A']:.0%}")

Further reading

  • Logistic regression, Google for Developers. How the model turns a score into a probability, and why it's trained on log loss.
  • Multinomial logistic regression, Wikipedia. The version for more than two outcomes, like home, draw and away.
  • Softmax function, Wikipedia. The step that turns three scores into three probabilities that add up to 1.

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.