majjha.
← All research

Are prediction markets accurate? What the record actually shows

Ask whether prediction markets are accurate and you will mostly collect adjectives. Uncanny, from the people selling them. Noise, from the people dismissing them. Both answers are cheap, and neither is necessary — the question has been accumulating a published record since 1988.

The case for taking these prices seriously is strong. It is also smaller and more conditional than the enthusiasm suggests, and the honest version includes both halves.

The election record

The Iowa Electronic Markets opened in 1988 — a small academic exchange run by the University of Iowa, where traders buy and sell contracts on election outcomes. Thirty-eight years later it is still running, and it is the field’s reference dataset.

The central study is Berg, Nelson and Rietz (2008). They compared the market’s vote-share prices against 964 national polls across the five U.S. presidential elections from 1988 to 2004. The market was closer to the eventual result 74% of the time. More than 100 days before the vote — the range where a forecast is actually useful — the market beat the polls in every one of the five elections.

That is the headline statistic of the field, and it deserves its reputation.

Beyond elections

Wolfers and Zitzewitz (2004), surveying the field in the Journal of Economic Perspectives — election markets, sports, economic data releases — found market forecasts fairly accurate and better than most moderately sophisticated benchmarks.

In 2008, twenty-two scholars — economists and legal scholars among them, several Nobel laureates — co-signed a short statement in Science citing mounting evidence that prediction-market forecasts carry lower prediction error than conventional forecasting methods, and urging regulators to clear a path for them.

Companies ran the experiment internally. Cowgill and Zitzewitz (2015) studied years of private prediction markets inside Google, Ford, and a large materials company they call Firm X — markets forecasting demand, deadlines and product outcomes. Despite thin participation and weak incentives, the markets improved on the firms’ own expert forecasts by as much as a 25% reduction in mean squared error.

Where the record cuts the other way

The margin is smaller than the reputation. Goel, Reeves, Watts and Pennock (2010) measured the gap across more than 7,000 NFL games, nearly 20,000 baseball games and about 100 film releases. In football — the deepest market of the set — the market was roughly 3% more accurate at predicting final scores than a simple three-parameter statistical model, and about 1% better than a poll of enthusiasts. In baseball, the market and the simple model effectively tied. The markets came out ahead or level — and nearly all of the predictive power was available to anyone with a spreadsheet.

Even the election headline has a serious critique. Erikson and Wlezien (2008) pointed out that a poll is not a forecast: it records preferences on the day it is taken, not a prediction of election day. Correct for that, they argued, and properly projected polls beat the market’s prices. The 74% figure survives as evidence that the raw market beats the raw poll; it does not survive as proof that markets beat everything you could build from polls.

Structured alternatives can beat markets. Atanasov and colleagues (2017) ran a randomized comparison inside a multi-year geopolitical forecasting tournament — more than 2,400 forecasters, 261 questions. Market prices beat the simple average of forecasters’ judgments. But teams of forecasters whose judgments were statistically aggregated — recency-weighted, performance-weighted, recalibrated — beat the market, and their edge was largest early in long-duration questions.

And the biases are real, documented, and live at the extremes. Page and Clemen (2013) tested calibration directly. Close to resolution, prices are reasonably well calibrated. Far from resolution, they are biased in a specific direction: low-probability events priced too high, near-certainties priced too low. In their data, a long-dated 20 was, on average, a real-world 15. Cowgill and Zitzewitz found the corporate analogue — an optimism bias inside Google and Ford — and found that it shrank as traders gained experience and the less skilled stopped trading. Thin, young, long-dated markets carry the least trustworthy prices.

The pattern is not a relic of small academic markets. A 2025 study of more than 300,000 contracts on a large licensed U.S. venue found the same pair of facts: prices sharpen as markets approach close, and low-price contracts win far less often than their prices imply (Bürgi, Deng and Whelan 2025, working paper). A 2026 decomposition of 292 million trades across 327,000 contracts on two of the category’s largest venues found calibration structured rather than uniform — a horizon effect, domain-specific biases, political prices chronically compressed toward 50 (Le 2026, preprint).

The record is still being written — and it includes a fresh public miss. On September 29, 2024, the Iowa markets read an 85.7% chance of a Democratic popular-vote win, and their thinly traded vote-share market showed a 9-point Democratic margin — numbers the venue’s researchers published in a peer-reviewed forecasting collection (Gruca and Rietz 2025). Five weeks later, the Republican side won the popular vote. The same paper reports the venue’s election-eve record across nine presidential cycles, 1988 through 2020: an average absolute error of 1.34 percentage points. Both facts are true at once. Near resolution, on an active book, these prices have been excellent. Far out, on a thin book, they can fail in public.

The widest study of the 2024 cycle covered more than 2,500 markets across four venues in the final five weeks — some $2.4 billion in transactions (Clinton and Huang 2025, working paper). It found the spread the older literature predicts. Accuracy varied widely from venue to venue. Identical contracts traded at different prices, and the gaps peaked in the last two weeks. Scale arrived; the old lessons held.

What “accurate” even means for a probability

A probability cannot be graded one question at a time. A market quoting 70 on an event that fails to happen was not necessarily wrong — three in ten of its 70s are supposed to fail. The grade only exists in aggregate: collect every question the market priced at 70, and check whether about seven in ten resolved yes. That property is called calibration, and it is the honest meaning of accuracy for a forecast stated as a number.

This is what separates a market price from a pundit. Not that the price is always right — the record above says plainly that it is not — but that the price can be checked. A market leaves a dated, public trail of exactly what it believed and when. Punditry does not.

Does the money matter?

One more result, because it surprises almost everyone. Servan-Schreiber, Wolfers, Pennock and Galebach (2004) ran the direct test during the 2003 NFL season: one exchange where traders risked their own dollars against one where they traded purely for play. The two sets of forecasts were equally accurate. Whatever makes these prices informative, it is not primarily the money. It is the mechanism — people stating beliefs as numbers, in public, where every statement gets scored.

The verdict

Read end to end, the record says something more useful than either camp’s adjectives.

The advantage is real and it is modest. Markets have beaten polls, internal experts and simple benchmarks more often than not — and a well-built model or a well-aggregated team lands in the same neighborhood. What the market adds is not magic: it is a standing, public number — dated, gradable, and open to correction by anyone who disagrees. The failure modes are documented — extremes, long horizons, thin books — and 2024 re-taught them in public.

Enthusiasm is not evidence. The record is.

So hold every venue to the standard the researchers set: publish the number, date it, let it be graded. Ask for the record.

Sources

  1. Berg, J., Nelson, F., Rietz, T. (2008). “Prediction market accuracy in the long run.” International Journal of Forecasting 24(2): 285–300. sciencedirect.com
  2. Wolfers, J., Zitzewitz, E. (2004). “Prediction Markets.” Journal of Economic Perspectives 18(2): 107–126. aeaweb.org
  3. Arrow, K. J., et al. (2008). “The Promise of Prediction Markets.” Science 320(5878): 877–878. science.org
  4. Cowgill, B., Zitzewitz, E. (2015). “Corporate Prediction Markets: Evidence from Google, Ford, and Firm X.” The Review of Economic Studies 82(4): 1309–1341. academic.oup.com
  5. Goel, S., Reeves, D. M., Watts, D. J., Pennock, D. M. (2010). “Prediction Without Markets.” Proceedings of the 11th ACM Conference on Electronic Commerce. dl.acm.org
  6. Erikson, R. S., Wlezien, C. (2008). “Are Political Markets Really Superior to Polls as Election Predictors?” Public Opinion Quarterly 72(2): 190–215. academic.oup.com
  7. Atanasov, P., et al. (2017). “Distilling the Wisdom of Crowds: Prediction Markets vs. Prediction Polls.” Management Science 63(3). pubsonline.informs.org
  8. Page, L., Clemen, R. T. (2013). “Do Prediction Markets Produce Well-Calibrated Probability Forecasts?” The Economic Journal 123(568): 491–513. academic.oup.com
  9. Bürgi, C., Deng, W., Whelan, K. (2025). “Makers and Takers: The Economics of the Kalshi Prediction Market.” Working paper, UCD Centre for Economic Research WP25/19. ucd.ie
  10. Le, N. A. (2026). “Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets.” Preprint, arXiv:2602.19520. arxiv.org
  11. Gruca, T. S., Rietz, T. A. (2025). “Iowa Electronic Markets: Forecasting the 2024 US Presidential Election.” PS: Political Science & Politics 58(2). cambridge.org
  12. Clinton, J. D., Huang, T. (2025). “Prediction Markets? The Accuracy and Efficiency of $2.4 Billion in the 2024 Presidential Election.” Working paper, SocArXiv. ideas.repec.org
  13. Servan-Schreiber, E., Wolfers, J., Pennock, D. M., Galebach, B. (2004). “Prediction Markets: Does Money Matter?” Electronic Markets 14(3). tandfonline.com