Goal or miss? The Bernoulli distribution
A penalty has two outcomes and nothing in between. That simple idea, the Bernoulli trial, is the building block of expected goals and of most football statistics.
Beginner Part 1
Statistics, maths and machine learning, each one introduced through a football question.
A penalty has two outcomes and nothing in between. That simple idea, the Bernoulli trial, is the building block of expected goals and of most football statistics.
Beginner Part 1
One penalty is a Bernoulli trial. Five penalties, each with the same chance, is the Binomial distribution, and it shows why missing one in five is normal, not a slump.
Beginner Part 2
A team averages two goals a game. How likely is it to score exactly three? The Poisson distribution turns an average into a probability for every score, and it sits underneath most football prediction models.
Beginner Part 3
Poisson counts goals. The Exponential distribution times the wait between them, and shows why "we're due a goal" isn't how probability works.
Beginner Part 4
A striker scores with one shot in five. How many shots until his first goal? The Geometric distribution answers it, and shows why a three-match drought is often just bad luck.
Beginner Part 5
The Geometric distribution waits for the first goal. The Negative Binomial waits for the third, or the fifth. It shows how long a hat-trick really takes, and it has a second job modelling goals that vary more than Poisson allows.
Intermediate Part 6
A match has three possible results, not two. The Multinomial distribution handles any number of outcomes, and shows why a team's "expected" record over ten games almost never happens exactly.
Beginner Part 7
A striker scores 8 of his 10 penalties. Calling him an 80% taker is overconfident. The Beta distribution describes how sure we can really be about a probability, and how that changes as the evidence grows.
Intermediate Part 8
The Exponential distribution waits for one goal. The Gamma waits for several. It shows why a team that averages two goals a game gets its third before full time only about one match in three.
Intermediate Part 9
Ranking players from fastest to slowest tells you the order, not how unusual anyone is. The Normal distribution, and its standard deviation, measures how far a player stands out from the rest.
Beginner Part 10
The Uniform distribution says every outcome is equally likely. It's the fairest-sounding model in statistics, and testing it against 28,016 goals shows football isn't that fair.
Beginner Part 11
Transfer fees, wages and market values can't go below zero, cluster low and have a few enormous outliers. The Log-Normal distribution describes them, and shows why the average fee is a poor guide to a typical one.
Intermediate Part 12
A player's performance can be written as a list of numbers in a fixed order. That list is a vector, and it's the first step to comparing players, finding replacements and feeding football into machine learning.
Beginner Part 1
One player's performance is a vector. Stack several players together and you have a matrix, the grid that almost all of data science starts from. Multiply it by a set of weights and every player gets a rating.
Beginner Part 2
Two players' stats are two vectors, and the distance between them measures how alike they are. It's how recruitment teams shortlist replacements, and it only works once every stat is put on the same scale.
Beginner Part 3
A new signing scores his first four penalties. Bayesian thinking combines what you believed before with what you've just seen, and shows why a perfect start should nudge your opinion, not replace it.
Beginner Part 1
Early-season league tables mislead. Over 550 Scottish team-seasons, only a third of a team's early form carried on. Bayesian thinking predicted exactly how much, using a prior measured from the data.
Intermediate Part 2