Same style, different volume? The dot product and cosine similarity
Distance says how far apart two players' numbers are. Cosine similarity asks whether they point the same way, which compares playing style rather than volume. It only works once the stats are standardised.
Intermediate Part 4 of Linear Algebra Through Football
Contents
The football question
A squad player does everything your first-choice midfielder does, in the same proportions, just less of it. Do they play the same way?
Distance says no: his numbers are all smaller, so he sits a long way off. But style isn't volume. What we want to know is whether the two profiles have the same shape. That's what cosine similarity measures.
The concept
Think of each player's vector as an arrow. Distance asks how far apart the arrow tips are. Cosine similarity asks whether the arrows point in the same direction:
- 1: exactly the same direction, the same balance of stats.
- 0: at right angles, no relationship.
- −1: exactly opposite.
It's built from the dot product.
A football example
The same two midfielders as last time, with passes, tackles, shots and chances created:
- Player A: (68, 7, 3, 5)
- Player B: (64, 6, 2, 6)
The dot product
$$\begin{aligned} A \cdot B &= (68 \times 64) + (7 \times 6) \\ &\quad + (3 \times 2) + (5 \times 6) \\ &= 4{,}430 \end{aligned}$$
In plain football
- Multiply each stat by the matching stat of the other player, then add them up.
- It's big when both players are high on the same things, and small when one is high where the other is low.
- On its own it mixes up style and volume: busier players give bigger dot products. So it needs scaling.
Cosine similarity
Divide the dot product by the length of each arrow, which takes the volume out:
$$\cos\theta = \frac{A \cdot B}{\lVert A \rVert \, \lVert B \rVert}$$
$$= \frac{4{,}430}{68.61 \times 64.59} \approx 0.9997$$
In plain football
- The length of a player's arrow is Pythagoras again: \(\sqrt{68^2 + 7^2 + 3^2 + 5^2} \approx 68.61\) for Player A.
- Dividing by both lengths means only the direction counts. Double every one of a player's stats and his cosine similarity with anyone doesn't change.
- 0.9997 is almost exactly 1: by this measure, A and B point the same way.
That's what fixes the squad-player problem. A player with half of Player A's numbers in every column is 34.3 away by distance, but has a cosine similarity of exactly 1: the same style, played at a lower volume.
The catch: everyone looks the same
Now compare Player A with the rest of the midfield from part 2:
| Compared with Player A | Raw cosine similarity |
|---|---|
| Player B | 0.9997 |
| Player 2 | 0.9999 |
| Player 3 (attacking) | 0.9952 |
| Player 4 (attacking) | 0.9952 |
Every one of them is about 1, including the two attacking midfielders who play nothing like him. The reason is the same scale trap as last time. Passes are so much bigger than the other stats that every midfielder's arrow points almost straight along "passes", and all the arrows look parallel.
Standardise first
Put each stat on the same scale first. For each one, subtract the midfield's average and divide by how much it usually varies: a z-score, as in the Normal distribution. Now each number says "above or below this midfield's average, by how much". Then compare the directions again:
| Compared with Player A | Standardised cosine similarity |
|---|---|
| Player 2 | 0.71 |
| Player B | 0.71 |
| Player 4 (attacking) | −0.72 |
| Player 3 (attacking) | −1.00 |
In plain football
- Player A passes and tackles more than this midfield's average, and shoots and creates less: a deep-lying midfielder.
- Player 2 and Player B lean the same way, so they score about 0.7: a similar style.
- Player 3 is his mirror image: fewer passes and tackles, more shots and chances. At −1, he's the opposite type of midfielder.
Standardised, cosine similarity finally says something useful: who plays the same role, whatever their volume.
Show the mathsThe general formulas, and where the angle comes from. Optional.
For two vectors of length n, the dot product and length are
$$\mathbf{a} \cdot \mathbf{b} = \sum_{j=1}^{n} a_j b_j$$
$$\lVert \mathbf{a} \rVert = \sqrt{\mathbf{a} \cdot \mathbf{a}}$$
Geometrically, \(\mathbf{a} \cdot \mathbf{b} = \lVert \mathbf{a} \rVert \lVert \mathbf{b} \rVert \cos\theta\), where \(\theta\) is the angle between the arrows. Rearranging gives cosine similarity:
$$\cos\theta = \frac{\mathbf{a} \cdot \mathbf{b}}{\lVert \mathbf{a} \rVert \, \lVert \mathbf{b} \rVert}$$
When every entry is positive, as raw counts always are, the angle is at most 90°, so the cosine can't go below 0. Standardising puts entries either side of zero, which is what lets similarity run all the way down to −1.
Distance or direction?
| Use | When you're asking |
|---|---|
| Distance | Who produces the same numbers: similar output as well as similar style? |
| Cosine similarity | Who plays the same way, whatever their minutes or volume? |
Recruitment often wants both: a player with the same style (high cosine similarity) who also produces enough (close in distance).
Why it matters
Cosine similarity is one of the most used measures in data science. Scouting tools use it to find players with similar styles. Recommendation systems use it to find songs or films like the ones you already enjoy. AI search engines use it to match a question with the documents most likely to answer it: every piece of text becomes a long vector, and the closest in direction wins.
Limitations
- Standardise first, or the biggest stat decides everything and every player looks alike.
- The comparison group matters. "Above average" here means above this four-man midfield. Against a whole league, the numbers would change.
- Style isn't quality. A player who does very little, but in the same proportions as a star, scores a perfect 1. Pair cosine similarity with a measure of volume before signing anyone.
- The features still decide the answer, as with distance.
Try it yourself
Take the three midfielders from last time's exercise. Standardise each stat across the three, then work out the cosine similarity between each pair. Does it agree with the distances, or does it find a style match that distance missed?
Further reading
- Dot products and duality, 3Blue1Brown. Why the dot product measures how much two arrows point the same way.
- Vector dot product and vector length, Khan Academy. The calculation, step by step.
- Cosine similarity, Wikipedia. The definition and its uses, from text search to recommendations.