ALASTAIR NORMAN
Article

The Wisdom of the Crowd

AI in practiceTechnology adoptionProduct strategy

Thesis

Individually, almost everyone guessing the weight of a cow, the contents of a jar, or the outcome of an election is wrong. Collectively, under the right conditions, the same group converges on something startlingly close to the truth. That is not a curiosity or a party trick: it is a specific statistical mechanism with known conditions for working and known ways of breaking, and it is now being used deliberately, not by accident, in prediction markets and in the way modern AI models are built.

Counting Jelly Beans

I first came across the idea some years back watching The Code, a BBC series in which Professor Marcus du Sautoy explored how mathematics shows up in everyday life. In one segment, he set a jar of jelly beans down in an office and asked everyone there to guess how many were inside. In total, 160 people offered a number.

The guesses ranged wildly, from around 400 to 50,000. Only a handful landed anywhere near the true count of 4,510. But averaged across all 160 people, the collective guess came out at 4,514, just 0.1% off. Nobody in the room was close. The room, taken together, was almost exact.

The Ox at Plymouth

The same effect has a much older, and much better documented, origin story, if you like. In 1906, the statistician Francis Galton attended a livestock fair in Plymouth where visitors could pay a small fee to guess the weight of an ox, with the closest guess winning a prize. Around 787 valid entries were submitted.

Galton, who assumed the average person had no real ability to judge the weight of an ox, collected the tickets afterwards expecting to prove the crowd was foolish. Instead, the median guess came to 1,207 lbs against an actual dressed weight of 1,198 lbs, an error of under 1%. He published the result in Nature under the title "Vox Populi," and it has been the founding case study for the wisdom of crowds ever since.

Why Averaging Beats Guessing

Neither result is magic. It works because individual errors, when they are independent of each other, tend to cancel out rather than compound. If your guess is too high and mine is too low for unrelated reasons, averaging us together gets closer to the truth than either of us alone. Do that across a large enough group and the errors that remain are the ones everyone shares, not the ones any single person made.

That "independent" is doing most of the work. Writer James Surowiecki, who later popularised the term "wisdom of crowds," identified the conditions under which it holds: the group needs a genuine diversity of information and opinion, people need to form their guesses independently rather than off each other, judgement needs to be decentralised so no single source dominates it, and there needs to be a mechanism to aggregate everyone's input into one answer. The jelly bean jar and the Plymouth ox both had all four almost by accident: private, written, simultaneous guesses with no discussion, tallied afterwards by someone with no stake in the outcome.

When the Crowd Gets It Wrong

Take any of those conditions away and the effect degrades quickly. A crowd stops being wise the moment people start guessing off each other instead of independently: that is a bank run, a meme stock rally, or a social media pile-on, not a Plymouth ox. Prediction markets show this in miniature. A 2026 Bloomberg analysis of Polymarket and Kalshi found they generally outperformed opinion polling on high-profile races, where trading volume is high and information is genuinely dispersed, but stumbled on down-ballot contests, where few people trade, liquidity is thin, and a handful of large bets can move the price more than any new information does. Diversity and independence are not guaranteed by putting a number on something; they have to actually be present.

Prediction Markets Put a Price on the Guess

What Polymarket, Kalshi and similar platforms add to Galton's basic method is a financial incentive to be right rather than merely be heard. A study comparing Polymarket to national polling around the 2024 US election found the market repriced faster and more accurately around discrete events, such as the assassination attempt on Donald Trump in Pennsylvania, than survey-based forecasts did. That is the aggregation mechanism doing its job in close to real time: thousands of independent, self-interested guesses, continuously repriced, rather than 787 tickets counted once at the end of a fair.

It is worth being open about what this is not. As the Columbia Journalism Review has argued, a bet is not a poll: traders skew towards people with money and a risk appetite, markets can be moved by a small number of large positions, and a price reflects what people are willing to stake, not necessarily what the population as a whole believes. Prediction markets inherit the wisdom of crowds' mechanism, but also inherit its fragility whenever the crowd placing bets is neither diverse nor independent enough to earn the name.

AI Runs the Same Trick on Itself

The same mathematics now shows up inside AI systems, deliberately engineered rather than discovered at a fair. Ensemble methods train several models independently and combine their outputs, on the same logic that a group of independent guessers beats any one of them. Large language models increasingly use a version of this internally: sample the same question multiple times with some randomness, let each attempt reason independently, and take the answer that comes up most often, a technique generally called self-consistency. It is Galton's median, run in software, on a crowd of one model's own attempts instead of a crowd of fairgoers.

The Method, Not the Magic

None of this depends on any one guess, trader, or model call being clever. It depends on having enough independent, diverse attempts and a sound way of combining them. That is why the same result shows up in an office jelly bean jar in the 2010s, a Plymouth livestock fair in 1906, a prediction market pricing an election in real time, and a language model checking its own working. Once you see it as a method rather than a coincidence, the interesting question stops being "was the crowd right?" and becomes "were the conditions actually in place for it to be?"


Drawing on Francis Galton's original 1907 account of the Plymouth ox-weighing contest, Ken Wallis's revisiting of Galton's forecasting competition, and on Bloomberg's, arXiv's, and the Columbia Journalism Review's reporting on how well modern prediction markets live up to it.