Skip to content

Is accuracy the right score? Evaluating a model

Accuracy counts how many results a model called right, and hides almost everything else. A confusion matrix shows where the mistakes are, probability scores show how confident it was, and five SPFL seasons show why draws break accuracy.

Beginner Part 7 of Machine Learning Through Football

Contents

The football question

A model calls 53.5% of results right. The bookmakers call 56.2%. Is that all you need to know?

Accuracy, the share of results a model gets right, is the number everyone quotes. It's simple, and it's what training data and test data and the parts after it used to compare models. But it only asks one question: did the model's favourite win? It says nothing about where the model goes wrong, or how confident it was.

The concept

Evaluating a model means choosing how to score it, and there are two big gaps in accuracy:

  • Where are the mistakes? A model that's 50% accurate could be wrong in all sorts of ways. A confusion matrix lays them out: what the model picked against what actually happened.
  • How sure was it? Most models give probabilities, not just a pick: 60% home win, 25% draw, 15% away win. Accuracy throws those away. Probability scores use them.

Both are tried here on the same test as underfitting vs overfitting: Scottish Premiership models trained on 2000/01 to 2020/21 and tested on the 990 matches of 2021/22 to 2025/26. In those matches there were 467 home wins, 227 draws and 296 away wins.

Where the mistakes are

Here's the bookmaker's favourite (the shortest Bet365 price) against what actually happened:

Right answers on the diagonal, in green: 385 home wins and 171 away wins called correctly. Everything off it is a miss. The draw row is empty: the bookmaker never once made a draw the favourite.

The diagonal holds the right answers: 385 + 171 = 556 of 990, which is the 56.2% accuracy. But the matrix shows something accuracy hides completely. The draw row is empty. In 990 matches, a draw was never the bookmaker's most likely result, so all 227 draws count as misses.

It isn't just the bookmaker. None of the other models here picks a draw either, including the form and table-position model from underfitting vs overfitting. Draws happen nearly a quarter of the time, yet on accuracy they're almost impossible to call.

The bookmakers aren't wrong about draws

Look at the draw probabilities instead of the picks, grouping matches by how likely the bookmaker made a draw:

Draw chance Matches Said Happened
Under 25% 378 18.2% 15.6%
25–28% 302 26.6% 27.5%
28%+ 310 29.3% 27.4%

When the bookmaker makes a draw more likely, draws do happen more often, and at close to the rate it says. The draw prices are good. They just never go above 33.2%, so a draw is never the single most likely result. Accuracy only sees the pick, so it gives the bookmaker no credit at all for pricing draws well. Probability scores do.

Scoring the probabilities

Two scores judge the probabilities themselves. Both are explained in full, with worked examples, in evaluating prediction models:

$$\text{log loss} = -\ln(p_{\text{actual}})$$

In plain football

  • p actual is the probability the model gave to what really happened.
  • Said 60% and it happened: a small penalty, 0.51. Said 5% and it happened: a big one, 3.0. Said 0% and it happened: an infinite one.
  • Averaged over every match, lower is better. A model that says a third each for home, draw and away every time scores 1.099: that's "knows nothing".

The Brier score adds up the squared gaps between each of the three probabilities and what happened (1 for the result that came in, 0 for the others). Lower is better again; "knows nothing" scores 0.667.

Model Accuracy Log loss Brier
Knows nothing (a third each) 47.2% 1.099 0.667
Base rates (always the league's usual split) 47.2% 1.056 0.638
Form + table position 53.5% 0.997 0.594
Bookmaker 56.2% 0.932 0.550

Two things stand out. The first two rows have identical accuracy, because both always end up picking a home win, but the probability scores separate them: knowing the league's usual split of home wins, draws and away wins is worth something, and accuracy can't see it. And on these four models the order is the same on every score. That isn't always so: among the eleven models in evaluating prediction models, the rankings on accuracy and log loss disagree.

The forecast that said "impossible"

The form + table position model works out its probabilities from how often each result happened in each group of training matches. Done naively, a small group can say 0%: one group held just two training matches, both away wins, so it put the chance of a home win at nothing.

On the test seasons, that naive version said 0% three times about results that then happened. Accuracy didn't notice; a wrong pick is a wrong pick. Log loss did: a single 0% for something that happens makes it infinite, and the model is scored as infinitely bad.

The fix is to start every group with one imaginary home win, draw and away win before adding the real matches, the same "no opinion yet" starting point as Bayesian thinking. No group can then say 0%, and the model's log loss comes out at the 0.997 in the table. Its accuracy doesn't change at all, which is the point: accuracy never saw the problem.

Pick the score for the job

Who's asking What they need Score
A pundit calling results The most likely result Accuracy, and a confusion matrix to see the misses
A bettor Prices that are right, including for draws Log loss or Brier, and a check that 30% means 30%
A club planning a season Expected points from every fixture Probabilities, scored on log loss or Brier

Even after all the care over training and test data and cross-validation, accuracy might not be the best number to judge a model by. It depends what the model is for.

Checking that 30% really means 30% is called calibration; who'll win the league? runs that check on 25 seasons of title chances.

Why it matters

  • Accuracy hides where the mistakes are. A confusion matrix shows them; here it shows a whole result, the draw, that no model ever picks.
  • Accuracy throws away confidence. Two models with the same accuracy can be very different forecasters.
  • Impossible is a strong word. A model that says 0% will eventually be embarrassed; log loss makes sure you find out.
  • The right score depends on the decision. Picking results, pricing them and planning a season each need a different yardstick.

Limitations

  • Probability scores need probabilities. A model that only names a winner, like "the side higher in the table", can only be scored on accuracy.
  • Five seasons is still a sample. Each accuracy here is good to about ±3 points; the log loss and Brier gaps are more stable, but not exact.
  • Bookmaker prices carry a margin. Turning odds into probabilities means scaling them to add up to 1, which removes the margin only approximately.
  • Calibration needs lots of matches. With a few hundred matches per band, "said 27%, happened 27.5%" could easily have been a few points out either way.

Try it yourself

Before the next round of fixtures, write down your probabilities for each match: home win, draw, away win, adding up to 100%. Afterwards, score yourself two ways: how many favourites won, and your average log loss (the calculator's −ln of the probability you gave to what happened). Did you ever make a draw your favourite? Would it have helped?

Reproduce the analysis

The results and odds files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:

import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import log

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}

def season(s):  # one row per match, with features known before kick-off
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    for r in games:
        r["day"] = datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y")
    games.sort(key=lambda r: r["day"])
    pts, res, gd, gf, rows = defaultdict(list), defaultdict(list), Counter(), Counter(), []
    for r in games:
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(pts[h]) >= 5 and len(pts[a]) >= 5:
            gap = (sum(pts[h][-5:]) - sum(pts[a][-5:])) / 5
            r["form"] = 4 - ((gap <= -1) + (gap < -0.4) + (gap <= 0.4) + (gap < 1))  # 0 away far better ... 4 home far better
            table = sorted(pts, key=lambda t: (sum(pts[t]), gd[t], gf[t]), reverse=True)
            r["pos_h"], r["pos_a"] = table.index(h) + 1, table.index(a) + 1
            d = r["pos_a"] - r["pos_h"]  # places the home side is above the away side
            r["pos"] = (d >= -6) + (d >= -2) + (d > 2) + (d > 6)  # 0 away far higher ... 4 home far higher
            r["last2"] = "".join(res[h][-2:])
            r["weekday"], r["month"] = r["day"].strftime("%a"), r["day"].month
            rows.append(r)
        hg, ag = int(r["FTHG"]), int(r["FTAG"])
        gd[h] += hg - ag; gd[a] += ag - hg; gf[h] += hg; gf[a] += ag
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            pts[team].append(p)
            res[team].append({3: "W", 1: "D", 0: "L"}[p])
    return rows

RESULTS = "HDA"

def group_probabilities(train, prior):  # share of each result in each form + table-position group
    seen = defaultdict(Counter)
    for r in train:
        seen[r["form"], r["pos"]][r["FTR"]] += 1
    def predict(r):
        c = seen[r["form"], r["pos"]]
        return {k: (c[k] + prior) / (sum(c.values()) + 3 * prior) if c or prior else 1 / 3 for k in RESULTS}
    return predict

def bookmaker(r):  # Bet365 odds turned into probabilities that add up to 1
    inverse = {k: 1 / float(r["B365" + k]) for k in RESULTS}
    return {k: inverse[k] / sum(inverse.values()) for k in RESULTS}

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
train = [r for s in names[:21] for r in season(s)]  # 2000/01-2020/21
test = [r for s in names[21:] for r in season(s)]  # 2021/22-2025/26
base = Counter(r["FTR"] for r in train)
models = {
    "knows nothing": lambda r: {k: 1 / 3 for k in RESULTS},
    "base rates": lambda r: {k: base[k] / len(train) for k in RESULTS},
    "form + table position": group_probabilities(train, prior=1),
    "bookmaker": bookmaker,
}
for name, predict in models.items():
    picks, loss, brier = Counter(), 0, 0
    for r in test:
        p, actual = predict(r), r["FTR"]
        picks[max(RESULTS, key=p.get), actual] += 1
        loss -= log(p[actual])
        brier += sum((p[k] - (k == actual)) ** 2 for k in RESULTS)
    right = sum(picks[k, k] for k in RESULTS)
    print(f"{name:22} accuracy {right / len(test):.1%}  log loss {loss / len(test):.3f}  Brier {brier / len(test):.3f}")
    for k in RESULTS:  # the confusion matrix: rows are picks, columns what happened (H, D, A)
        print(f"   picked {k}:", [picks[k, a] for a in RESULTS])

raw = group_probabilities(train, prior=0)
print("forecasts of 0% that happened:", sum(raw(r)[r["FTR"]] == 0 for r in test))

band = lambda d: "under 25%" if d < 0.25 else "25-28%" if d < 0.28 else "28% and over"
said, drew, n = Counter(), Counter(), Counter()
for r in test:
    d = bookmaker(r)["D"]
    said[band(d)] += d; drew[band(d)] += r["FTR"] == "D"; n[band(d)] += 1
for b in ("under 25%", "25-28%", "28% and over"):
    print(f"bookmaker draw {b:13} {n[b]} matches: said {said[b] / n[b]:.1%}, happened {drew[b] / n[b]:.1%}")

Further reading

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.