← All projects

04 · Machine learning · 2023

F1 prediction

Ranking a Grand Prix finishing order

A model that predicts the full finishing order of a Formula 1 race and is judged the way a ranking should be, not just on whether it guessed the winner.

Complete · rebuild planned

Year
2023
Data
F1 results, 1950–2020
Models
Gradient boosting, learning to rank
Scored with
NDCG, MAP, Spearman

Live demo

Who wins the next race?

A live pick for the next Grand Prix on the calendar. It updates as the weekend unfolds and keeps score against the pole-sitter all season. Then, purely for fun, bet against it.

LiveLoading the calendar…

Just for fun

Bet against the model

Loading the grid…

Lap pace
1.0×

Why a ranking?

A race result is an ordering of about twenty drivers. “Who wins?” throws most of that away. A model that gets P1 right and scrambles the rest of the grid isn't a good model, and accuracy alone can't tell you that.

So every model here is scored on the whole order, the way search engines score a page of results: getting the top places right counts most, but every position counts.

At a glance

What went into it

71seasons of race results, 1950 to 2020
1,021Grands Prix
24,680individual driver results
0.878NDCG@10 in cross-validation (1.0 is a perfect order)

For every race, the model ranks every entrant and predicts the full finishing order. It's validated race by race and scored with ranking metrics, then compared across model types to see which ordering holds up best.

The journey

From a question to a ranking model

  1. The question

    “Who wins?” is the wrong question

    Race prediction is usually framed as picking the winner and graded on accuracy. But what a race produces is a ranking, and the interesting problem is the whole order: who makes the podium, who scores points, who falls back.

    So I framed it as learning to rank. Each race is a group, the model orders the drivers in it, and it's graded on how close that order is to what actually happened.

  2. The data

    Seventy-one seasons of results

    I built the dataset from the public Ergast F1 archive: race results, race details and qualifying, joined by race and driver. That's 24,680 driver results across 1,021 Grands Prix, from 1950 to 2020.

    Real race data is messy. Qualifying records only exist from 1994 onward, and 43% of entries didn't finish. Drivers who retired are ranked behind every finisher in their race, so the order stays complete.

  3. Features

    Describing each driver going into the race

    Each entry is described by what's known going into the race: the starting grid position and whether the driver starts from the pit lane, the qualifying result, and recent form, meaning the average finish over each driver's and each team's previous three races.

  4. Models

    Two ways to produce an order

    I compared two approaches. A gradient-boosted model predicts each driver's finishing position, and the order follows from sorting. A learning-to-rank model (XGBoost) is trained to order the drivers within each race directly. A neural-network variant was set up for a second pass.

    Every model is validated with five-fold cross-validation grouped by race, so a race is never split between training and testing.

  5. Scoring

    Scoring it like a ranking

    Instead of accuracy, each model is scored with NDCG and MAP, the metrics search engines use for ranked results, plus Spearman correlation between the predicted and actual order. NDCG@10 asks how good the predicted top ten is, weighting the front of the grid most.

  6. What I found

    The grid is most of the story

    The starting grid carries the most weight, followed by driver and team form. Qualifying adds a little on top, since it largely decides the grid already.

    That changes what “good” means. The real bar for a race model isn't a random guess, it's the simple rule that everyone finishes where they started. A model is only useful where it beats that rule.

    What I foundPick the baseline before you pick the model. For finishing orders, “finish = grid” is the one to beat, and it's harder to beat than it looks.

  7. Next

    The rebuild

    This was one of my earlier projects, and I'm rebuilding it with what I've learned since. The new version pulls richer data from the FastF1 API, including lap times, race-control messages, circuit layouts and championship standings.

    It will compute every feature strictly as of race morning, train on past seasons and test on the next one, report the grid baseline next to every result, and use a modern ranking objective.

    The live pick at the top of this page is the first piece of it: it runs on this season's data and keeps score against the pole-sitter after every race.

Contact

Hiring for a software role?

I'm looking for full-time software, backend and systems roles. Email is the fastest way to reach me.

pranch555@gmail.com