At a glance
What went into it
For every race, the model ranks every entrant and predicts the full finishing order. It's validated race by race and scored with ranking metrics, then compared across model types to see which ordering holds up best.
The journey
From a question to a ranking model
-
The question
“Who wins?” is the wrong question
Race prediction is usually framed as picking the winner and graded on accuracy. But what a race produces is a ranking, and the interesting problem is the whole order: who makes the podium, who scores points, who falls back.
So I framed it as learning to rank. Each race is a group, the model orders the drivers in it, and it's graded on how close that order is to what actually happened.
-
The data
Seventy-one seasons of results
I built the dataset from the public Ergast F1 archive: race results, race details and qualifying, joined by race and driver. That's 24,680 driver results across 1,021 Grands Prix, from 1950 to 2020.
Real race data is messy. Qualifying records only exist from 1994 onward, and 43% of entries didn't finish. Drivers who retired are ranked behind every finisher in their race, so the order stays complete.
-
Features
Describing each driver going into the race
Each entry is described by what's known going into the race: the starting grid position and whether the driver starts from the pit lane, the qualifying result, and recent form, meaning the average finish over each driver's and each team's previous three races.
-
Models
Two ways to produce an order
I compared two approaches. A gradient-boosted model predicts each driver's finishing position, and the order follows from sorting. A learning-to-rank model (XGBoost) is trained to order the drivers within each race directly. A neural-network variant was set up for a second pass.
Every model is validated with five-fold cross-validation grouped by race, so a race is never split between training and testing.
-
Scoring
Scoring it like a ranking
Instead of accuracy, each model is scored with NDCG and MAP, the metrics search engines use for ranked results, plus Spearman correlation between the predicted and actual order. NDCG@10 asks how good the predicted top ten is, weighting the front of the grid most.
-
What I found
The grid is most of the story
The starting grid carries the most weight, followed by driver and team form. Qualifying adds a little on top, since it largely decides the grid already.
That changes what “good” means. The real bar for a race model isn't a random guess, it's the simple rule that everyone finishes where they started. A model is only useful where it beats that rule.
What I foundPick the baseline before you pick the model. For finishing orders, “finish = grid” is the one to beat, and it's harder to beat than it looks.
-
Next
The rebuild
This was one of my earlier projects, and I'm rebuilding it with what I've learned since. The new version pulls richer data from the FastF1 API, including lap times, race-control messages, circuit layouts and championship standings.
It will compute every feature strictly as of race morning, train on past seasons and test on the next one, report the grid baseline next to every result, and use a modern ranking objective.
The live pick at the top of this page is the first piece of it: it runs on this season's data and keeps score against the pole-sitter after every race.