344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support132
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
A ForecastBench project finds specialized AI forecasters score about the same as “superforecasters”, but the comparison has important limits.
In short: New benchmark results suggest specialized AI tools can make forecasts about future events about as accurately as top human forecasters, but the results come with major caveats.
Researchers are starting to test whether AI can make reliable predictions about real world events, like elections or ceasefires. The Financial Times points to ForecastBench, a project from the Forecasting Research Institute, which is led by psychologist Philip Tetlock.
ForecastBench gives forecasters a score based on how well their probability estimates match what later happens. Think of it like weather forecasts. Saying there is a 70 percent chance of rain is only “good” if, over time, days with a 70 percent prediction actually rain about 70 percent of the time.
On this scale, ForecastBench reports that the average “superforecaster” score is 68.9. The best specialized AI forecasters are scoring 68.8, which is very close. Ordinary human forecasts average 62.8, and general purpose AI chatbots, called large language models (systems trained on lots of text), score a bit over 61.
The FT notes several reasons to be cautious. The human superforecasters answered one set of questions in 2024, while the AI systems are answering different questions now. Some questions are simply easier than others, and ForecastBench says its adjustment for question difficulty becomes less reliable over time.
There is also a broader concern about what happens if powerful companies control highly accurate forecasting tools. And even if the forecasts are accurate, relying on AI could reduce the human value of the process, like missing the learning that comes from doing the thinking yourself rather than pressing a button.
Source: Financial Times