Where prediction markets fail: thin books, longshot bias, and the edge of the evidence
Every instrument has conditions under which it stops working. For prediction markets those conditions are named, measured, and published — by the same researchers who established that the prices are worth reading at all (Are prediction markets accurate?).
This piece is about the same record’s other column.
The literature that established the advantage has spent just as long mapping where it disappears — and the map is specific. The failures are not scattered embarrassments. They cluster in named, measurable conditions that have replicated across venues, decades and question types. Publishing that map is not a concession to the category’s critics. Knowing where an instrument fails is what it means to know the instrument.
Prediction markets fail in three documented, recurring conditions — and win by less than advertised in a fourth. Prices on low-probability events tend to run too high and near-certainties too low — the longshot bias — and prices far from resolution are less reliable than prices close to it. Thin books, where few traders participate, aggregate little information; and even where markets work well, their edge over a strong statistical model or a professionally aggregated forecaster team often narrows to a few percentage points, or disappears. The published record itself is uneven — concentrated in U.S. elections, thin on other regions and question types, and in places still awaiting peer review — so some regions of that failure map are better evidenced than others.
The bias with a name
The most consistent failure in the record has a direction: across most of the data, the improbable is priced too generously and the near-certain too cheaply.
The field’s own surveyors flagged it early. Reviewing the evidence two decades ago, Wolfers and Zitzewitz named performance on low-probability events the exception to the markets’ generally good record, and showed the overestimation of unlikely outcomes in the academic Iowa exchange’s own data (2006a, working paper). Page and Clemen (2013) later measured the shape directly: near resolution, prices track realized outcomes well; far from resolution, low-probability contracts trade above what their outcomes justify and high-probability contracts below.
The bias did not retire with the early venues. A 2025 study of more than 300,000 contracts on a large licensed U.S. venue found low-price contracts winning far less often than their prices imply (Bürgi, Deng and Whelan 2025, working paper). A 2026 decomposition of 353 million trades on two of the category’s largest venues found the miscalibration structured rather than uniform — organised by horizon, varying by subject, with political prices compressed toward 50 (Le 2026, preprint).
Why it happens is its own literature. The leading explanation is behavioral, not structural: people misperceive small probabilities. The pattern has been documented at scale well outside prediction markets — across 6.4 million U.S. horse race starts, misperception rather than risk preference best explained the same bias (Snowberg and Wolfers 2010).
A structural explanation is available but does not fit. Manski (2006) showed that under heterogeneous beliefs a market price only bounds the average belief rather than equaling it. But the bound is tightest at the extremes. Prices near zero or one are the most informative about what traders on average believe; prices near 0.5 the least. That is the opposite of where the longshot bias lives. Follow-up analysis reached the same conclusion from the other direction: under plausible conditions a market price closely tracks the central belief of the people trading it, and the structural pricing biases Manski’s bound admits are generally small, though the same paper shows risk preferences can produce the longshot shape where they do bite (Wolfers and Zitzewitz 2006b, working paper). That is what anchors Where a market price comes from. The extremes are miscalibrated despite being the range where the price should track belief most closely — which is why the behavioral account carries the weight.
And the record refuses to make the bias a law. A direct test on the Iowa exchange’s real-money markets over financial questions — monthly horizons, repeatedly reinitialized — found no longshot bias at all (Berg and Rietz 2019). The conditions matter, which is the point of this piece.
The practical reading is short. A 4 is not a 4 the way a 50 is a 50. At the extremes, the record says the quoted number is more likely to overstate the improbable outcome than to understate it.
Thin books
A market price is manufactured from trading (Where a market price comes from). It follows, and the evidence confirms, that a book with little trading manufactures little.
The laboratory isolates the mechanism: thinness held fixed, the question’s difficulty as the dial. With the same groups of three traders throughout, the standard continuous market mechanism performed well on a single binary event. It performed worst of the four designs tested once the question became three correlated events across eight securities. Traders concentrated on a few of them, prices in the neglected ones went badly wrong, and a structured iterative poll beat the market outright (Healy, Linardi, Lowery and Ledyard 2010). Three traders were enough while the board was small. Thinness bites through attention, and attention did not stretch. How a young venue answers the thin-book problem deliberately, rather than simply enduring it, is its own subject (Market makers, disclosed).
The field’s clearest public instance came in 2024. The Iowa Electronic Markets’ vote-share market — thinly traded, long-dated — showed a 9-point margin for the side that went on to lose the popular vote, a miss the venue’s own researchers published alongside the market’s excellent election-eve record (Gruca and Rietz 2025). Same venue, same cycle, same mechanism: the difference was depth and distance.
Corporate prediction markets supply the honest counterweight. Cowgill and Zitzewitz (2015) studied thin internal markets at Google, Ford and a third firm, and found them biased in documented ways — yet still accurate enough to improve on the firms’ own expert forecasts. Thinness degrades a market; it does not automatically sink it below the alternatives.
Long horizons
Distance in time behaves like thinness. Page and Clemen’s calibration failures concentrate far from resolution. Le’s 2026 decomposition (preprint) finds the same horizon structure at venue scale. And in the head-to-head tournament evidence, the market’s disadvantage is largest exactly there: Atanasov and colleagues (2017) found statistically aggregated forecaster teams beat market prices most clearly early in long-duration questions.
The pattern reads less like a flaw than a boundary condition. A market price summarizes what trading has processed so far; on a question with months to run, most of the information that will decide it has not arrived, so the price is a real estimate built on less evidence than it will eventually have.
The margin against alternatives
The failure map has a fourth region, and it is the one enthusiasm handles worst: even where markets work, the margin over good alternatives is modest.
Goel, Reeves, Watts and Pennock (2010) measured it across thousands of sports and film markets: a few percent over a simple statistical model in the best case, a tie in others. Erikson and Wlezien (2008) showed that properly projected polls beat raw market prices in the election record. But the correction cuts both ways. Debias the market prices for the longshot effect too and they beat the debiased polls again, most clearly early in the cycle and in uncertain races (Rothschild 2009). Atanasov’s aggregated teams did it with structured elicitation.
None of that unwinds the record in the markets’ favor — the same studies show markets beating naive benchmarks and matching sophisticated ones while remaining always-on, continuously updated and public. It bounds the claim. The published edge is a habit of small wins. It is not a different category of knowledge.
The edge of the evidence
There is one more limit, and it is not a failure condition at all: the places where the record simply runs out.
The accuracy literature is concentrated — by venue, by country, by question type. Elections dominate. U.S. markets dominate. The venue-scale studies putting numbers on the modern category are working papers and preprints, most dated 2025 and 2026, not yet through peer review. A 2026 study of what venues in under-covered regions can even list found the tradable questions skew toward what can be cleanly worded and sourced, not what matters most locally (Adegbenro 2026, preprint) — the record inherits that skew. Even the intuitive fixes carry caveats. One long-circulated working paper found that on short-horizon event contracts, the more liquid securities were not the better-calibrated ones (Tetlock 2008, working paper). Depth is the best-documented defense in the record, and still not a guarantee.
So the honest statement of the evidence has edges in both senses: documented failure conditions inside the map, and unmapped territory beyond it. The accuracy and calibration studies behind this map — sample, setting, result and publication status, kept current — are indexed in the evidence index.
Reading the map
The failure literature is sometimes wielded as a verdict against the category. It reads better as an operating manual.
Trust the number most where the record says it has been reliable: active books, near resolution, mid-range prices, questions with real information flow. Discount it at the edges the same literature drew: longshots, long horizons, thin books, and any comparison where a strong model or a disciplined team is available. Those are not secrets extracted from reluctant venues. They are findings the field published about itself, in journals, for anyone to read.
Carry the map. A number whose failure conditions you know is worth more than a number you trust.
Sources
- Wolfers, J., Zitzewitz, E. (2006a). “Five Open Questions About Prediction Markets.” Working paper, NBER Working Paper 12060. nber.org
- Wolfers, J., Zitzewitz, E. (2006b). “Interpreting Prediction Market Prices as Probabilities.” Working paper, NBER Working Paper 12200. nber.org
- Page, L., Clemen, R. T. (2013). “Do Prediction Markets Produce Well-Calibrated Probability Forecasts?” The Economic Journal 123(568): 491–513. academic.oup.com
- Bürgi, C., Deng, W., Whelan, K. (2025). “Makers and Takers: The Economics of the Kalshi Prediction Market.” Working paper, UCD Centre for Economic Research WP25/19. ucd.ie
- Le, N. A. (2026). “Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets.” Preprint, arXiv:2602.19520. arxiv.org
- Snowberg, E., Wolfers, J. (2010). “Explaining the Favorite–Long Shot Bias: Is it Risk-Love or Misperceptions?” Journal of Political Economy 118(4): 723–746. journals.uchicago.edu
- Manski, C. F. (2006). “Interpreting the Predictions of Prediction Markets.” Economics Letters 91(3): 425–429. Also NBER Working Paper 10359. nber.org
- Berg, J. E., Rietz, T. A. (2019). “Longshots, Overconfidence and Efficiency on the Iowa Electronic Market.” International Journal of Forecasting 35(1): 271–287. ideas.repec.org
- Healy, P. J., Linardi, S., Lowery, J. R., Ledyard, J. O. (2010). “Prediction Markets: Alternative Mechanisms for Complex Environments with Few Traders.” Management Science 56(11): 1977–1996. ideas.repec.org
- Gruca, T. S., Rietz, T. A. (2025). “Iowa Electronic Markets: Forecasting the 2024 US Presidential Election.” PS: Political Science & Politics 58(2). cambridge.org
- Cowgill, B., Zitzewitz, E. (2015). “Corporate Prediction Markets: Evidence from Google, Ford, and Firm X.” The Review of Economic Studies 82(4): 1309–1341. academic.oup.com
- Atanasov, P., et al. (2017). “Distilling the Wisdom of Crowds: Prediction Markets vs. Prediction Polls.” Management Science 63(3). pubsonline.informs.org
- Goel, S., Reeves, D. M., Watts, D. J., Pennock, D. M. (2010). “Prediction Without Markets.” Proceedings of the 11th ACM Conference on Electronic Commerce. dl.acm.org
- Erikson, R. S., Wlezien, C. (2008). “Are Political Markets Really Superior to Polls as Election Predictors?” Public Opinion Quarterly 72(2): 190–215. academic.oup.com
- Rothschild, D. (2009). “Forecasting Elections: Comparing Prediction Markets, Polls, and Their Biases.” Public Opinion Quarterly 73(5): 895–916. academic.oup.com
- Adegbenro, A. (2026). “What Prediction Markets Can See: Market Formation, Settlement Legibility, and the Geography of Tradable Uncertainty in Africa and Latin America.” Preprint, arXiv:2606.17503. arxiv.org
- Tetlock, P. C. (2008). “Liquidity and Prediction Market Efficiency.” Working paper. business.columbia.edu