The last goal of the tournament
Sunday evening, July 19, MetLife Stadium near New York. The World Cup final is goalless after ninety minutes. Argentina have not managed a single attempt on goal. Messi’s only contribution in the entire match will be one blocked shot in the 115th minute. Deep in stoppage time, Enzo Fernández picks up a second yellow card. His first one, absurdly, was for complaining to the referee. Argentina start extra time a man down.
Thirty-nine seconds into the second period of extra time, Nico Williams heads a hooked cross back into the middle, and Ferran Torres smashes it in from close range. Spain win 1-0 after extra time.
My prediction said Spain 1-0.
In my pool, that is worth ten points: the exact score, delivered through extra time. It settled a five-week title race by twelve points. First place among just over a hundred players. 750 euros on a 13 euro entry.
Here is the strange part. Serious statisticians will tell you that a football match is a terrible instrument for finding out which team is better. One well-known paper estimates that the chance the best team actually won the 2006 World Cup was below one in three.[1] And yet the model we had built gave this exact outcome, Spain 1-0 through extra time, a probability of 17 to 20 percent before kickoff. It was the eighth exact score our system hit in the knockout rounds alone.
Both things are true at once. Football is mostly noise, and we still finished clearly above what a perfectly informed forecaster should expect to score. This post is the full story of how: the setup, the methods, the failures, and at the end an honest calculation of how much was skill and how much was the coin landing my way.
The pool
The contest is a classic score prediction pool, the kind known in the German-speaking world as a Tippspiel. Just over a hundred players, 13 euros in, prizes of 750, 325 and 125 euros for the top three. For every match of the World Cup you submit one exact score before kickoff. Scoring is tiered, and only the highest tier counts:
- 10 points for the exact score
- 7 points for the correct winner with the correct goal difference
- 4 points for the correct winner only
- 0 for the wrong winner
Group games are judged after ninety minutes. Knockout games are judged on the score after extra time, and the penalty shootout is ignored entirely. A tie that is still level after 120 minutes counts as a draw, no matter who goes through. On top of the 104 match bets sit a few bonus bets submitted in advance: group winners, the four semifinalists, the top scorer’s team, and the champion.
About the field: most of the other players genuinely live football. Season tickets, fantasy leagues, strong opinions about back three versus back four. I am not one of them. I enjoy a World Cup like anyone else, but I could not have told you who plays right back for Paraguay, and I had no ambition to learn. What I do have is a programming and data analytics background and, this summer, curiosity about a specific question: how far can process and statistics carry you against a hundred people who actually know the sport?
One idea from the academic literature shaped everything that follows, and it has nothing to do with football. A prediction pool is a relative game against a crowd, not an absolute game against a bookmaker. Research on tournament betting pools shows that the best strategy depends on what everyone else is picking and on how many people you have to beat, and that maximizing your expected points is not the same thing as maximizing your chance of finishing first.[2][3] That distinction ends up driving almost every strategic decision in this story.
What the science says about predicting football
Before claiming any edge, it is worth establishing how little edge is available.
Two statisticians, Skinner and Freeman, treat a football match as an experiment for identifying the better team, and find it badly underpowered.[1] By their analysis, only five of the sixty-four matches at the 2006 World Cup produced scorelines convincing enough to give better than 90 percent confidence in the result. Their half-joking remedy, keep playing extra time until the goal difference reaches a chosen level of confidence, will turn out to be surprisingly relevant later in this post.
The comparative evidence points the same way. Across more than 300,000 games in five major leagues, football showed the highest rate of upsets of any sport studied, and about a quarter of English top-division matches ended in a draw.[4] Close to half of all goals involve at least one element of chance, deflections, rebounds, shots off the woodwork, defensive errors.[5][6] And a study of 1,503 league seasons across several team sports finds luck substantially present even in the most competitive championships, which, as the authors note, partially explains why complex prediction models barely beat simple ones.[7]
If the signal is mostly noise, three responses follow logically, and they became the skeleton of the whole project:
- Optimize the scoring system, not your football opinions. In a tiered exact score contest, the question is never “who wins?” but “which single scoreline maximizes expected points?”, and that is a computation, not a debate.
- Account for luck explicitly, so that you neither “learn” from noise after a bad day nor congratulate yourself after a lucky one.
- Play the crowd, not just the game. Behind when it matters, pick differently from the field. Ahead, mirror it.
The rest of this post is those three ideas in action, with a language model doing most of the heavy lifting.
My co-analyst
The AI in question is Claude, Anthropic’s assistant. I want to be precise about what it did and did not do, because “I won with AI” can mean anything.
The honest beginning is that the quality did not appear by itself. The first sessions produced competent but shallow match previews, and the early chat logs literally contain me pushing back: ‘this is not four games’, ‘it is every group’, ‘take your time’. Depth emerged through repeated nagging. The turning point was writing the nagging down.
We defined a fixed three-stage research routine for every round: evidence gathering with a mandatory case against our own pick, then quantitative modeling, then pool strategy. It went into the project’s permanent instructions together with a checklist and a required sign-off line at the top of every delivery, so that a skipped stage would be visible at a glance. From then on the full pipeline ran by itself, every round, without being asked. Of everything I learned about working with these models, this stuck with me most: the default behavior is negotiable, but only if you write the negotiated behavior down somewhere the model reads it every single session.
The second structural piece was memory. A chat has no durable state, so we built one: a single markdown file, the dossier, holding the bet log, a forensic report on every match, profiles of all teams, the growing rule set, standings and strategy notes. After every round it was updated, checked and re-uploaded as the project’s single source of truth, growing from about 330 lines in mid-June to 823 by the end.
Keeping a living document intact for five weeks needed some genuinely boring engineering. Every change was applied by a small script that refused to run unless it found its target text exactly once in the file, and every update ended with an automatic scan for leftover outdated numbers. Boring, until you consider what a strategy document that silently drifts away from reality would have cost.
The division of labor never changed. Claude did the research, the modeling, the post-match analysis and the file maintenance. I made every final pick, held veto power over every proposed rule, and after each round typed in the official point totals from the pool’s leaderboard. That last, least glamorous job earned its keep in week two: cross-checking my reported totals against the logged bets revealed that one of my bets had been recorded wrongly in the dossier, a Brazil 3-0 written down as 2-0. It had been an exact hit, a six point difference, found by simple addition rather than by anyone’s memory. In a contest that would eventually be decided by twelve points, arithmetic hygiene was not optional.
How a prediction actually got made
Every round of matches ran through the same three stages.
Stage one was evidence. Current odds from several bookmakers, cleaned of the built-in bookmaker margin so they become honest probabilities. Confirmed lineups and injury news. The referee. Venue, roof, altitude, weather. What each team still needed from the match. The reasoning behind expert picks, not just their numbers. And then, always, the case against our own lean, argued properly before the pick was allowed to survive.
Stage two was the model. A Poisson score model in the tradition of Maher[8] and Dixon and Coles[9], fitted for each match to those cleaned odds and to the market’s expected goal total. It produces a probability for every possible scoreline, hence the expected points of any pick under the 10, 7, 4, 0 system, plus sensitivity checks: what if the total is a bit higher, the favourite a bit stronger. For knockout games it all ran through an extra-time adjustment I will come back to, because getting it wrong nearly derailed the campaign.
Stage three was pool strategy. Are we chasing or leading, and what does that mean for this pick? Two habits grew here during the tournament. First, name the single question the pick really turns on, the hinge, and spend the analysis budget there. Second, write down switch conditions in advance, “if this player is missing and the live odds flip, change the pick to that score”, and check them mechanically an hour before kickoff.
Why trust bookmaker odds over our own reads or over fancier models? Because the research says to. In the best-known open football prediction competition, run in 2017 on a database of over 200,000 matches, the winning machine learning entry got about 52 percent of match outcomes right,[10][11] and probabilities backed out of bookmaker odds afterwards beat every submitted model on the competition’s official measure.[12] The 2023 edition asked for exact scores, and there machine learning still trailed rule-based team ratings, with the organizers’ own simple baselines outperforming far more complex entries.[13][14] So we treated the cleaned odds as the best available estimate of who wins and by how much, and spent our effort on what odds do not price: which exact score to pick, the mechanics of the scoring rules, and the behavior of the crowd.
The core idea of exact score picking is simple enough to state in two sentences. Under tiered scoring, expected value concentrates on the single most probable scoreline, because that one pick maximizes the 10 point and the 7 point tier at the same time. And the most probable scorelines in football are boring on purpose: across the 964 World Cup matches played before this tournament, 1-0 had happened 182 times, about 19 percent, and 2-1 another 152 times, about 16 percent.[15] A second structural idea follows from the goal difference tier. The scores 1-0, 2-1 and 3-2 all share the same goal difference, so a “favourite by one” pick collects seven points on any one-goal win. We called this the by-one family, and it worked like armor against chaos: three separate times during the tournament, a 1-0 ticket paid seven points on a match that finished 3-2.
One worked example, chosen because it shows the machine at its most transparent. Round of 32, Colombia against Ghana. Stage one: Ghana’s deep block had frustrated England to a 0-0, but breaking down how they had actually scored all tournament showed essentially nothing from open play: a genuinely toothless attack, the one profile where betting on a clean sheet is defensible. Stage two: the odds-fitted model put Colombia’s most likely goal count at exactly one and flagged 1-0 at 19.0 percent, the highest exact score probability of that round. Stage three: we were leading by then, so take the most likely line and resist the temptation of 2-0. Result: Colombia 1-0, goal in the 14th minute, Ghana without a shot on target, a hot keeper and a disallowed Colombian goal keeping it at one. Ten points, exactly as priced. Most picks were messier. The point of the routine was to make as many as possible look like that before kickoff.
Falling to 21st
The shape of the campaign is best shown as a table:
| Date | Standing | What happened |
|---|---|---|
| Jun 13 | 1st | three exact scores in the first four games |
| Jun 17 | 21st, 20 behind | the all-draw day and its aftermath |
| Jun 19 | 8th, 15 behind | grinding |
| Jun 23 | 4th, 14 behind | grinding |
| Jun 26 | 3rd, 5 behind | the draw plays paying off |
| Jun 27 | 2nd, 3 behind | group stage over |
| Jun 29 | 1st, 3 ahead | two knockout exacts flip the race |
| Jul 11 | 1st, 19 ahead | the biggest lead |
| Jul 15 | 1st, 9 ahead | the semifinal scare |
| Jul 19 | 1st, 12 ahead | final score: 609 points |
The opening days put me briefly on top of the table. Then came June 15: four fixtures, four draws. Uruguay 1-1, Spain 0-0, Iran 2-2, Belgium 1-1, against four picks that all named a winner. Zero points, on a day so draw-heavy that the whole field bled with me. By the time the first round of group games closed, I had fallen to 21st, roughly twenty points back.
It was the most useful failure of the tournament, for two reasons. First, it exposed a bad habit: three of the four losses came from betting on clean sheets against teams with a real scoring threat. The correction became the very first rule in our rulebook. Second, it posed the strategic question that the pool research answers: the more opponents you face, the further the optimal picks deviate from the consensus.[2] Far behind, that logic bites hardest. Maximizing expected points just mirrors the field and freezes your deficit in place; you need outcomes where you score and the crowd does not.
Our tool for that was the draw play: a deliberate draw pick in a genuine coin-flip match that the field would predictably bet a winner in, used at most once or twice per round, and only where the bookmaker’s draw price was fat and neither team badly needed to win. The record while chasing: five attempts, three hits, two misses. Every hit banked seven points while the favourite-backing crowd took zero, and those three hits, 21 points in total, were the measurable difference in the climb from 21st place to 2nd.
The best of them was also the most intense one. Algeria against Austria on the last group matchday. A draw sent both teams through, and the unanimous view before the game, ours included, was a careful, low-scoring standoff. It finished 3-3, with an apparent winner in the 93rd minute cancelled by an equalizer in the 95th, scored by a substitute who had been on the pitch for 61 seconds. The draw ticket collected its seven points anyway. The lesson went into the rulebook in its correct, humbler form: when both teams are happy with a draw, that tells you the likely outcome, and tells you nothing at all about how wild the game will be.
A rulebook written in losses
The most valuable thing the project produced was not any single pick but a numbered rule set, seventeen rules by the end, every one of them tied to a specific, dated failure or confirmation in our own log. Three principles kept it honest.
After every round, each result got a written reconstruction and an explicit split into foresight and luck before anything was allowed to change. A pattern needed at least two independent confirmations before it became a rule, and several rounds formally ended with the conclusion “no new rule”, on purpose. And nothing entered the rule set without my sign-off.
Four rules carried the campaign and show how the method worked.
The mechanism test. Before betting on any clean sheet, break down how the opponent has actually scored: open play versus set pieces, own goals and one-off wonder strikes; chances they create themselves versus chances that depend on service from teammates; records inflated by weak opponents. The founding cases were our own misses. Sweden’s scoring record was pumped up by a single 5-1 against Tunisia, and their strikers depended on supply, so a ball-dominant France starved them and shut them out while we had predicted them to score. Bosnia’s “scored in every game” record turned out to be built on corners, an own goal and a teenager’s screamer, and they produced next to nothing against the USA even with a man advantage.
One night of the round of 32 drew the boundary. Ghana, toothless by mechanism, delivered the bankable clean sheet and our exact 1-0. Cape Verde, carrying genuine self-made threats, scored twice against Argentina and broke a clean sheet we should never have counted on. The question is never “do they score a lot?” but “by what route do they score at all?”
Deference to the sharpest odds. Whenever our hand-built view of a team disagreed clearly with the best-priced markets, the market was right, and we could measure it. We had marked Mexico’s attack down for squad rotation; following the sharp odds instead turned a planned draw pick into Mexico 1-0 and saved four points when Mexico won 3-0. The same day, we held a “Bosnia’s attack is cold” opinion against odds that said otherwise; Bosnia won 3-1, and the market’s implied score would have outscored ours by three. From then on, when odds and opinion collided, the opinion lost.
The independent-threat rule. In a knockout tie with a heavy favourite, if the underdog has any self-made way of scoring, pick the favourite to win with a goal conceded rather than to nil. Whether the threat is big does not matter; that it exists is enough, because one goal is all it takes to kill a clean-sheet pick. This rule earned its number the hard way, with five confirmations, one of them from the wrong side of the line: we picked Argentina 1-0 against Switzerland in defiance of the rule’s letter, after our own analysis had named, in writing, the one Swiss player capable of scoring on his own. He equalized.
The greed tax. For a heavy favourite, take the model’s most likely winning margin and never reach for the flashier blowout. The tournament ran a clean natural experiment on this. The two times we broke the rule, picking 0-3 where the model said 0-2, and 3-0 where it said 2-0, cost exactly three points each, because the favourite conceded and won by the modest margin both times. The two times the discipline was hardest to keep, it paid in full: England 2-0 instead of a tempting 3-0 hit the exact score, and Brazil 2-1 instead of the flashy 3-1 hit the exact score and was worth six points more than the greedy version. There is a nice large-scale echo of this. Janning Vygen, the founder of Kicktipp, Germany’s most popular prediction-pool platform, has said that what they observe is players with a lot of knowledge tipping too riskily, convinced they can convert that knowledge into points, and reaching for an underdog that does not deliver.[16] Our log reproduces his observation in miniature.
The bug in our own rulebook
On June 30, nineteen days and 76 bets into the tournament, we discovered a mistake. Not in a pick. In our own strategy document. It stated that knockout matches were judged after ninety minutes. The pool’s actual rule, which every other player had understood correctly from day one: knockouts are judged on the score after extra time, penalties excluded, so a 90-minute draw that gets decided in extra time turns into a recorded win, and only a tie still level after 120 minutes counts as a draw.
To be clear about what kind of error this was: nobody had been scored a single wrong point, and the people I was playing against knew their own rules perfectly well. The bug lived only inside our model and notes, and it had already quietly shaped the first four knockout picks. Left uncaught, it would have kept steering every knockout decision from a wrong map of the payoffs.
The fix went in the same day, and the model was rebuilt around the real rule: decisive 90-minute scores stand, drawn ones get extra-time goals added before recording. How expensive would the bug have been? In the benchmark calculation later in this post, the best possible pick’s expected value in a typical knockout match is about half a point higher under the correct rule than under the wrong one, mostly because the rule shifts probability from draws to narrow wins and changes which score is worth targeting. Over 32 knockout bets, that is roughly fifteen points of bad decisions avoided, in a race that was decided by twelve.
What the corrected model then did with the rule shaped everything that followed. It declared the draw nearly dead as a knockout pick, since a draw ticket now only pays when a game runs all the way to penalties. And it made narrow favourite wins even more attractive, because the extra-time path funnels into exactly those scorelines. Belgium 3-2 after a 2-2 at ninety, decided by a penalty at 124 minutes and 44 seconds, the latest goal in World Cup history, paid our 2-1 ticket seven points. A quarterfinal night on which both matches stood 1-1 after ninety produced fourteen points, including an England 2-1 exact score created by extra time itself. And the final’s 0-0 at ninety became the recorded 1-0 we had bet.
There is a nice piece of irony in this. Skinner and Freeman’s joke cure for football’s randomness was to keep playing extra time until the goal difference becomes statistically meaningful. A knockout rule that counts the score after extra time is a small dose of exactly that. The academic joke had been sitting in our pool’s rulebook all along, known to everyone. Our only achievement was to stop misreading it in time.
Leading is a different game
The race flipped on the first knockout weekend, and it flipped on two exact scores. Canada 1-0, a 91st minute winner in a match whose whole shape the model had called: toothless opponent, safe clean sheet, modest margin. And Brazil 2-1 against Japan, the hardest exact of the tournament, which needed three separate judgments to be right at the same time: that Japan would score, and through precisely the defensive weakness we had flagged; that the temptation of a bigger Brazil win should be resisted; and that the game would be tight and slow, which it was, to the point that Brazil’s coach was openly saving Neymar for extra time.
Those twenty points carried me from 2nd to 1st. It helped that the weekend’s two big shocks, Germany and the Netherlands both going out on penalties, counted as recorded draws and wiped out practically the entire field along with us.
From that day on, the strategy inverted, exactly as the pool research says it should for a leader. No more contrarian picks. Mirror the crowd’s favourites, so that shared losses cost nothing in relative terms. And compete on the one axis that is pure upside: exact score precision, where hitting the ten is worth three, six or ten points against the field’s neighboring picks, and missing costs nothing you were not already risking. The knockout engine, market-based margins, mechanism-tested clean sheet calls, the by-one family, the greed tax, the extra-time adjustment, produced seven exact scores in the knockout rounds before the final, and the lead grew on almost every round for three weeks.
The stress test came at the semifinals, where both of my picks scored zero. This is where writing things down in advance proved its worth. Both losses had been priced before kickoff as explicit failure probabilities, 50.7 percent for one match and 54.2 percent for the other, about a 28 percent chance that both would land. Both pre-agreed switch conditions were checked when the lineups came out, and correctly stayed silent. So when the cushion collapsed from 19 and 25 points down to 9 and 21, the written record said: this is the leader’s accepted coin-flip cost, not a broken process. Nothing changed. It stung anyway.
For what it is worth, the one round of the entire tournament in which my closest rivals gained ground was this one, and the arithmetic implies that second place hit an underdog exact score in a semifinal. A perfectly played chaser’s move, straight out of the same theory I had used to climb.
The last weekend
The final week turned into an exercise in applied probability more than football.
First, a decisive discovery about the bonus bets. Every player still within range held the identical champion and top scorer combination, Spain and France, except one, whose picks were already dead. That meant every remaining bonus outcome would move the leaderboard in lockstep, with zero effect on the gaps. For completeness: the bonus bets supplied 170 of my final 609 points. All four semifinalists correct, the Spain champion bet and the France top scorer bet both landing, the latter through Mbappé becoming the first player to win a second Golden Boot. Not every advance bet had worked; a Portugal group-winner ticket had died quietly back in June.
With bonuses neutralized, the pool reduced to a duel: me against second place, nine points ahead, two hidden score picks each. We listed every combination of results by which they could still pass me, with probabilities attached, which surfaced a pleasing property: if their Saturday gamble failed while I banked even four points, the gap would exceed the maximum possible single-game swing and the pool would be mathematically over a day early. Overall chance of being caught, by our estimate: somewhere between 12 and 20 percent.
Saturday delivered the priced chaos. The third-place match, which our model gave a 41 percent chance of scoring me nothing, produced France 4-6 England. Ten goals, the most in any World Cup match since 1982 and the most ever in a third-place game. England 4-0 up at halftime with a rotated lineup full of players auditioning for next season, a Saka hat-trick, and Mbappé scoring twice on the way to the all-time World Cup goals record. My France 2-1 scored zero, exactly as budgeted, and the math held: the worst Saturday outcome still left Sunday’s ceiling intact.
Sunday was the pick the whole system had been building toward. Spain 1-0 carried a computed expected value of 3.34 points and an exact-hit probability of 17 to 20 percent, the biggest single-score target since that Colombia match, against an honestly stated 52 percent chance of scoring nothing at all, the largest coin flip we accepted all tournament. Three things we had written down in advance all happened. Spain’s shot volume compressed to exactly one goal, twenty shots against a keeper having the night of his life, with two more goals disallowed. Argentina, whose famous late-game scoring we had argued was a product of chasing games rather than a permanent quality, never got to chase, because Spain never sat back, and finished with almost nothing: no attempt before extra time. And the extra-time pipeline carried a 0-0 into a recorded 1-0. Spain ended the tournament having conceded a single goal in seven matches, the fewest by any world champion ever. Both of my rivals took exactly seven points from the weekend. Final margin: twelve.
How lucky did I get?
Winning proves very little on its own, so afterwards I rebuilt the scoring mathematics from scratch and measured the campaign two ways. First against expectations: Poisson score distributions calibrated to cleaned bookmaker probabilities in four typical match categories, from heavy favourite to coin flip, mixed to match my actual fixtures, with the extra-time adjustment for knockouts. Second against reality: I replayed each simple strategy against the actual recorded results of all 104 matches, favourite defined by the betting line, to see what it would truly have scored on this tournament.
The two columns answer different questions. Expected points are the long-run average a strategy would earn if the tournament were replayed many times. The realized column is what it earned on the one tournament that actually happened, and the gap between the columns is luck.
| Strategy | Expected per game | Actually scored on these 104 games |
|---|---|---|
| Always pick 1-1 | 1.92 | 204 |
| Favourite 2-1, every match | 3.30 | 400 |
| Favourite 1-0, every match | 3.40 | 406 |
| Perfectly calibrated best pick | 3.40 | not replayable |
| This campaign | close to the row above, by design | 439 |
| Perfect foresight | 10.00 | 1040 |
Five findings.
The first one genuinely surprised me: at World Cup-typical goal totals, the boring favourite 1-0 is the mathematically best pick across almost the whole range of favourites. With match totals around 2.6 goals, 1-0 maximizes expected points from roughly even matches all the way up to favourites with about a 75 percent win probability; only for extreme mismatches does 2-0 take over, and only when the expected total climbs toward three does the best pick move to 2-1 or 2-0. The reason is subtle. Even against a 75 percent favourite, 2-0 is the more likely exact score, 15.5 against 14.1 percent, but 1-0 still wins on expected points because winning by exactly one goal remains the most likely margin.
The practical consequence is humbling: the folk wisdom “always tip the favourite 1-0” matches the perfect picker on average. Any real edge must live in knowing when to leave the default, match by match. The record suggests the system did exactly that: eleven of my eighteen exact scores were not 1-0.
Second, the realized column shows something the expected column cannot: this particular tournament broke heavily toward favourites. They won 70 of the 104 matches, 24 ended drawn, and only 10 were genuine upsets. One match in eight finished exactly 1-0 for the favourite. As a result, even the purely mechanical favourite 1-0 strategy would have scored 406 points, 53 above its own long-run expectation, about one and a half standard deviations of pure good fortune, available to anyone consistently backing favourites. Our 439 therefore splits into two parts: a favourite-friendly tailwind of roughly 53 points that we shared with every disciplined favourite backer, and 33 points of our own, earned by deviating from mechanical play on identical fixtures. That 33 is the fairest single number for the campaign’s edge.
Third, the exact score rate ran hot, but stayed inside the plausible range. Eighteen exacts in 104 tries is 17.3 percent against a baseline near 12.5 percent, about a ten percent chance of happening by luck alone. In the knockouts, eight in thirty-two is 25 percent against the probabilities our own model had assigned before each match, which averaged about 13.5 percent; the chance of doing that well by luck is around six percent. Strong, not miraculous.
Fourth, we came close to our own ceiling. Adding up every mistake our written post-match reports classified as avoidable, points that better use of our own rules would have earned, gives a best-possible-play total of roughly 460 to 480 match points. The realized 439 captures 91 to 95 percent of that. The full outcome distribution, for the record: 18 exact scores, 17 correct goal differences, 35 correct winners, 34 zeros. Some points on 67 percent of all bets.
Fifth, and most important for honesty: the winning margin lives deep inside luck’s error bars. With a per-game standard deviation around 3.4, two equally skilled players running this exact strategy over 104 matches would typically differ by about 49 points on pure chance. I won by twelve, over a runner-up on 597 and a third place on 585.
So the sober summary reads: the process ran near its own ceiling, the tournament gave every favourite backer a tailwind, and on top of that our deviations earned 33 points on the real fixtures. Beating well-calibrated odds on average is impossible if the odds are efficient, so those 33 points must themselves be some mixture of genuinely better score selection and further good luck; with only 104 matches the two cannot be cleanly separated, and our own match reports, read honestly, contain plenty of both. It is Skinner and Freeman’s conclusion about a single match, scaled up to a single pool.
What I’m left with
Three things surprised me, and none of them are about football.
The first is where the edge actually came from. Going in, I assumed the value would be in analysis, in some clever read of form or tactics. In hindsight, the biggest single edge was taking the scoring system seriously as a piece of mathematics: computing which score maximizes expected points under the exact rules, what the extra-time recording does to those numbers, and when the boring pick is the right one. The second biggest was refusing to have opinions where a liquid betting market already had better ones. Almost none of it required knowing anything about football that a search engine could not supply in thirty seconds.
The second is what it took to make an AI genuinely useful for something like this. Not clever prompting in the moment, but infrastructure around the model: a permanent instruction set it re-reads every session, a document that serves as its memory, and a human who checks the totals and owns the decisions. Once that existed, the quality was remarkably consistent for five straight weeks. Before it existed, the same model produced pleasant, shallow previews.
The third is that the hardest part was emotional, not technical. Doing nothing after the all-draw day. Doing nothing after scoring zero points across both semifinals. Trusting numbers written down before kickoff over the very loud feeling, afterwards, that something must be wrong. The written record made that possible, because it kept proving that the bad days had been priced in.
Whether any of this survives contact with a harder game is an open question, and I intend to find out. The same setup is currently being pointed at a season-long Bundesliga fantasy league, where the target, weekly player performance, is notoriously noisy: a purpose-built AI agent from the research literature managed to reach only around the top percentile of 2.5 million human fantasy players.[17] If the noise wins there, you will read about that too.
Two footnotes, so nobody takes the wrong message. This was a fixed-stake pool among colleagues and friends, thirteen euros, once. It is not sports betting, and nothing here is betting advice; the studies in the science section above are very clear about what happens, on average, to people who try to out-predict football for money. And I remain what I was in June: someone who still could not name Paraguay’s right back, and no longer needs to.
A note on the numbers
All campaign figures come from a log kept during the tournament itself: every bet recorded before kickoff, every result analyzed in writing within a day, every standing checked against the pool’s official leaderboard. The headline numbers: 609 total points, made up of 439 from the 104 match bets and 170 from bonus bets; first place out of just over a hundred entrants, twelve points ahead of second and twenty-four ahead of third; eighteen exact scores, ten in the group stage and eight in the knockouts, two of those through extra time.
The realized strategy totals were computed by replaying each fixed strategy against the actual recorded results of all 104 matches, with the pre-match favourite defined by the betting line. Nine matches were close enough to coin flips that the favourite call is debatable; flipping any or all of them shifts the mechanical strategies by a few points and changes no conclusion. The expected-points comparison uses simple, independent Poisson score distributions calibrated to cleaned bookmaker win probabilities in four typical match categories, mixed to reflect my actual fixture list, with the extra-time adjustment applied to knockout games. It is deliberately basic. Known simplifications: no correction for the slight correlation between low scores, estimated category weights, and matches treated as independent in the standard deviation calculations. None of these change the ranking of the strategies or the size of the conclusions in any way that matters.
One disclosure, since transparency is the point of this post: Claude drafted the analyses, ran the models and maintained the records throughout the tournament, and helped draft this text. Every pick, every rule, and every claim above was reviewed and decided by me.
Skinner, G. K., & Freeman, G. H. (2009). Soccer matches as experiments: how often does the ‘best’ team win? Journal of Applied Statistics, 36(10), 1087-1095. Preprint: arXiv:0909.4555. ↩︎ ↩︎
Clair, B., & Letscher, D. (2007). Optimal strategies for sports betting pools. Operations Research, 55(6), 1163-1177. ↩︎ ↩︎
Kaplan, E. H., & Garstka, S. J. (2001). March Madness and the office pool. Management Science, 47(3), 369-382. A qualitative precursor: it maximizes expected score and sets opponent behavior aside; the dependence on the crowd’s picks and the pool size is modeled explicitly by Clair and Letscher. ↩︎
Ben-Naim, E., Vazquez, F., & Redner, S. (2006). Parity and predictability of competitions. Journal of Quantitative Analysis in Sports, 2(4). Preprint: arXiv:physics/0608007. ↩︎
Lames, M. (2018). Chance involvement in goal scoring in football, an empirical approach. German Journal of Exercise and Sport Research, 48, 278-286. ↩︎
Brechot, M., & Flepp, R. (2020). Dealing with randomness in match outcomes: How to rethink performance evaluation in European club football using expected goals. Journal of Sports Economics, 21(4), 335-362. Cited for the broader influence of randomness on match outcomes. ↩︎
Aoki, R. Y. S., Assunção, R. M., & Vaz de Melo, P. O. S. (2017). Luck is hard to beat: The difficulty of sports prediction. Proceedings of the 23rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 1367-1376. Preprint: arXiv:1706.02447. ↩︎
Maher, M. J. (1982). Modelling association football scores. Statistica Neerlandica, 36(3), 109-118. ↩︎
Dixon, M. J., & Coles, S. G. (1997). Modelling association football scores and inefficiencies in the football betting market. Journal of the Royal Statistical Society: Series C (Applied Statistics), 46(2), 265-280. ↩︎
Dubitzky, W., Lopes, P., Davis, J., & Berrar, D. (2019). The Open International Soccer Database for machine learning. Machine Learning, 108(1), 9-28, and the accompanying special issue on machine learning for soccer. ↩︎
Hubáček, O., Šourek, G., & Železný, F. (2019). Learning to predict soccer results from relational data with gradient boosted trees. Machine Learning, 108(1), 29-47. ↩︎
Robberechts, P., & Davis, J. (2019). Forecasting the FIFA World Cup: Combining result- and goal-based team ability parameters. In Machine Learning and Data Mining for Sports Analytics (MLSA 2018), Springer, 16-30. ↩︎
Yeung, C., Bunker, R., Umemoto, R., & Fujii, K. (2024). Evaluating soccer match prediction models: a deep learning approach and feature optimization for gradient-boosted trees. Machine Learning, 113. ↩︎
Berrar, D., Lopes, P., & Dubitzky, W. (2024). A data- and knowledge-driven framework for developing machine learning models to predict soccer match outcomes. Machine Learning, 113, 8165-8204. ↩︎
Fjelstul, J. (2022). The Fjelstul World Cup Database (v1.0), CC-BY 4.0. Scoreline counts computed from its match table: all 964 men’s World Cup matches, 1930-2022. ↩︎
Vygen, J. (2024). Quoted in: “Seriös, wenig Risiko und den Außenseiter nicht vergessen — So gewinnen Sie das Tippspiel zur Fußball EM.” Kicktipp GmbH press release, 6 May 2024. Wayback Machine copy, since the original page is offline; the key quote is also reproduced at fussballtippspiel.com. ↩︎
Matthews, T., Ramchurn, S. D., & Chalkiadakis, G. (2012). Competing with humans at fantasy football: Team formation in large partially-observable domains. Proceedings of the 26th AAAI Conference on Artificial Intelligence, 1394-1400. ↩︎