In the 1980s, psychologist Philip Tetlock began investigating the accuracy of expert predictions on geopolitical events, discovering that even well-regarded specialists often diverged widely in their forecasts. This finding questioned the reliability of expert judgment. Tetlock’s subsequent research culminated in his 2005 book, “Expert Political Judgment: How Good Is It? How Can We Know?,” which revealed that specialists were only marginally more accurate than random chance.
However, Tetlock maintained that more effective forecasting was possible. His later work, including the 2015 publication “Superforecasting,” identified a small group of individuals—dubbed “superforecasters”—who consistently made better predictions by employing particular analytical habits and probabilistic reasoning. These insights have since encouraged a broader practice of publicly sharing testable forecasts.
In recent years, interest has grown in comparing predictions made by human experts to those generated by artificial intelligence systems, especially large language models (LLMs). Projects like ForecastBench, run by the Forecasting Research Institute led by Tetlock, have tracked forecasting accuracy across humans and AI. According to the latest data, specialized AI forecasters score nearly as well as top-rated human superforecasters, with accuracy indices around 68.8 and 68.9 respectively. Regular humans and untuned LLMs trail with scores closer to 62.8 and just over 61.
These scores are measured on a scale where a perfect oracle would achieve 100, while a forecaster assigning neutral 50-50 probabilities to all outcomes would score 50. The results suggest both humans and AI currently outperform random guessing but remain far from infallible. However, direct comparisons between AI and human forecasts are complicated by differences in the questions posed—some events are inherently easier to predict than others—and the methods used to adjust for question difficulty may be statistically imperfect.
Beyond accuracy, the broader implications of relying on AI for forecasting raise concerns. Experts emphasize that forecasts do not merely describe potential futures; they can influence decisions and shape outcomes. The dominance of forecasts produced by AI systems developed and controlled by powerful corporations could have unforeseen effects on governance and public trust.
Moreover, the value of forecasting lies not only in the end prediction but also in the deliberative process it fosters. Engaging in structured forecasting exercises encourages open-mindedness, reduces political polarization, and promotes a deeper understanding of complex scenarios. These human elements contribute to cultivating better judgment and decision-making skills.
Given these factors, some analysts caution against fully delegating forecasting tasks to automated systems. While AI’s growing capabilities are notable, the nuanced judgment, collaboration, and accountability inherent in human forecasting remain vital components for interpreting and responding to an uncertain future.
