A goldfish living in a Toronto storefront tank has beaten five of the most powerful artificial intelligence models at predicting World Cup 2026 results. Swimbappé, the fish, logged roughly 80% accuracy midway through the group stage by swimming toward one of two flags representing competing nations. By contrast, ChatGPT, Claude, Gemini, Copilot, and Perplexity, all tasked by Dutch research bureau Onder with forecasting every match of the tournament, have fallen short of that mark.
The experiment, tracked on a live ScoreGPT scoreboard, has become one of the more talked-about side stories of this summer's tournament. It also raises an awkward question for the AI industry: if billion-parameter models cannot reliably call a football match, what does that say about their real-world reasoning?
How the Fish Calls the Shots
Swimbappé operates from a tank fitted with two flags at opposite ends. His handlers release him, and whichever flag he swims toward first is registered as his pick. Draws are ruled no-contests because the setup offers no third option. Midway through the group stage, the fish had racked up sixteen correct calls against four misses. That is a better hit rate than any of the five AI tools being graded by Onder's ScoreGPT tracker.
The fish's method is, of course, random. But randomness with a simple binary choice can outperform overfitted models that misread form, injuries, and momentum. Swimbappé does not overthink. The machines, apparently, do.
Where the AI Models Stumbled
Before the June 11 opener, Onder gave each model the same brief: weigh form, rankings, and injuries, then predict all 104 matches. The answers were sealed and are being marked in real time.
The models' best moment came before kickoff. Four of the five correctly named Spain, France, England, and Argentina as the semifinalists. All four teams made it that far. But the same models also predicted Spain would beat France in the final, a scenario that became impossible when the two sides met in the semifinals instead.
Individual match predictions have been weaker. The tools flagged Brazil as the tournament's biggest disappointment and Norway as the dark horse, yet none predicted Norway would reach the round of 16. Germany's elimination by Paraguay was a miss for the entire ScoreGPT panel, which had backed a comfortable German win. When USA Today asked Copilot to call four matches in a single day, all four ended in draws, an outcome the model had not considered for any of them.
A Global Zoo of Forecasters
Animal predictions at major football tournaments are not new. The tradition traces back to Paul the Octopus, who called eight consecutive matches correctly at the 2010 World Cup, final included. That run spawned a global cottage industry of animal oracles.
This year, a hawk named Shawk in Dubai has forecast more than a dozen matches. Thai zoos have deployed multiple animals, with celebrity pygmy hippo Moo Deng making headlines on July 14 by picking France over Spain and England over Argentina in the semifinals. In Mexico City, a duck named Merlin served more as a talisman than a forecaster, waddling through the hosts' opening-day celebrations and meeting officials at the National Palace on June 22.
The animal-versus-algorithm narrative is partly spectacle, partly sincere curiosity about whether data-driven models can outperform instinct, or in this case, chance.
What This Says About AI Hype
The ScoreGPT experiment is a lighthearted stunt, but it lands at a moment when AI companies are pitching large language models as reasoning engines capable of complex judgment. A tournament like the World Cup is noisy, emotional, and unpredictable. That is precisely why it is a useful stress test. If models trained on vast sports datasets cannot outpredict a fish with a three-second memory, it suggests that real-world uncertainty still breaks algorithmic confidence in ways that marketing decks rarely acknowledge.
The gap between macro accuracy (picking semifinalists) and micro accuracy (calling individual matches) also mirrors a broader pattern in AI deployment: models often sound authoritative at a distance and fall apart on the specifics.
What Happens Next
Moo Deng's semifinal verdict has been public since July 14. Swimbappé is scheduled to make his final pick before Sunday's kickoff. Onder will continue updating the ScoreGPT tracker through the closing matches, and the AI industry will be watching, if only to see whether it can close the gap on a goldfish before the trophy is lifted.