RSO-F’s first submission to ForecastBench has placed above evaluated configurations of GPT-5.5 from OpenAI, Claude Opus 5 from Anthropic and Grok 4.20 from xAI on the preliminary dataset leaderboard.

The entry is listed as RSO-F-ST. We submitted its forecasts in the August 2026 round. The September snapshot covers 241 resolved dataset questions from that submission’s seven-day horizon.

What the result covers

ForecastBench’s preliminary leaderboard scores resolved dataset questions. It provides an early view of performance before a model qualifies for the main tournament leaderboard, which also includes market questions.

The comparison uses a difficulty-adjusted score. Entries can cover different questions and submission histories, so the named systems have not all been evaluated on the same seven-day set. Our seven-day results are the basis for RSO-F-ST’s current position on that broader preliminary board.

Forecasting further ahead

The thirty-day results will test RSO-F over a longer horizon. We expect performance to decline from this early result. The next evaluation will show how well the model’s forecasts hold as events unfold.

We will publish those findings and continue testing RSO-F in future rounds, using the results to improve the model. Our goal is to help people anticipate real-world events and make better decisions.

Explore RSO-F

From the archive. The date refers to the period described.

ForecastBench preliminary leaderboard

Leaderboard snapshot, September 2026

How ForecastBench works