Skip to main content

Overview

Metaculus uses sophisticated scoring algorithms to evaluate forecast accuracy. Scoring incentivizes truthful probability estimates and rewards forecasters who make bold, accurate predictions.

Score Types

Metaculus provides multiple scoring methods, each measuring different aspects of forecasting skill. From scoring/constants.py:4-10, the available score types are:

Peer Score

Measures how much better you forecast compared to the community

Baseline Score

Measures raw forecasting accuracy against a uniform baseline

Spot Peer

Peer score evaluated at a specific point in time

Spot Baseline

Baseline score evaluated at a specific point in time

Peer Score

Peer scores measure how much better (or worse) your forecasts are compared to the community prediction.

Algorithm

From scoring/score_math.py:152-200, peer scores use logarithmic scoring:

Key Properties

Peer scores are relative: A positive peer score means you beat the community, while a negative score means the community was more accurate.
  • Scale: Typically ranges from -50 to +50 points
  • Zero sum: Peer scores across all forecasters sum to approximately zero
  • Requires multiple forecasters: Need at least 2 forecasters (formula includes num_forecasters / (num_forecasters - 1))

Time-Weighted

Peer scores are weighted by coverage - the fraction of the forecasting period your prediction was active:

Baseline Score

Baseline scores measure raw forecasting accuracy against a naive baseline prediction.

Algorithm

From scoring/score_math.py:62-108, baseline scoring:

Baseline Assumptions

Baseline: Uniform distribution over all available optionsFor a binary question, the baseline is 50/50. For a 4-option multiple choice, each option gets 25%.Maximum score: 100 points (achieved with 100% confidence in the correct outcome)Minimum score: Negative infinity (as probability approaches 0 for the correct outcome)

Key Properties

  • Scale: 0 to 100+ points (higher is better)
  • Absolute: Scores don’t depend on other forecasters
  • Always calculable: Can be computed even with only one forecaster
  • Incentivizes boldness: Confident, accurate predictions earn more points

Spot Scores

Spot scores evaluate forecasts at a single point in time rather than across the entire forecasting period.

Spot Scoring Time

From questions/models.py:428-441, the spot scoring time is determined by:
Priority order:
  1. Explicit spot_scoring_time if set
  2. cp_reveal_time (when community prediction is revealed)
  3. actual_close_time
  4. scheduled_close_time

Spot Baseline Score

From scoring/score_math.py:111-149:
If you don’t have an active forecast at the spot scoring time, you receive 0 points for that question.

Use Cases

Spot scores are ideal for:
  • Questions where the resolution can affect forecasting (CP hidden until reveal)
  • Live forecasting events with synchronized scoring times
  • Questions that resolve quickly after closing
  • Preventing score manipulation by rapid forecast updates

Score Model

Scores are stored per user, per question, per score type (from scoring/models.py:17-57):

Coverage

The coverage field tracks what fraction of the forecasting period you participated in:
  • Coverage = 1.0: You had an active forecast for the entire question period
  • Coverage = 0.5: You forecasted for half the question period
  • Coverage = 0.0: You never made a forecast
Early forecasting is rewarded! Making predictions early and maintaining them gives you higher coverage and more scoring opportunities.

Archived Scores

Historical scores that can’t be recalculated are archived (from scoring/models.py:59-94):
Archived score types (from scoring/constants.py:13-14):

Default Score Type

Each question designates a default score type (from questions/models.py:91-102):

Question Weighting

Questions can have different weights when aggregating scores (from questions/models.py:90):
  • Weight = 1.0: Standard question (default)
  • Weight > 1.0: More important question (higher contribution to leaderboards)
  • Weight = 0.0: Question excluded from scoring entirely
  • Weight < 1.0: Less important question

Unsuccessful Resolutions

Some questions cannot be resolved normally:
  • Ambiguous: Resolution criteria cannot be clearly applied
  • Annulled: Question was flawed or should not have been asked
From questions/constants.py:4-6:
Questions resolved as ambiguous or annulled are typically excluded from scoring and leaderboards.

Scoring Best Practices

Forecast early and update regularly to maximize your coverage. Even a simple initial forecast can boost your score potential.
Over many forecasts, events you predict at 70% should happen about 70% of the time. Track your calibration!
The scoring system rewards confident, accurate predictions. If you have strong evidence, don’t be afraid to make extreme forecasts.
Scores are weighted by time, so updating your forecast when you learn new information helps your score.
To beat the baseline, you need to do better than uniform/naive probabilities. Consider what “uninformed” would predict.

API Reference

Scores API

Explore the full Scores API documentation

Questions

Understand question types and structure

Forecasting

Learn how to make predictions

Leaderboards

See how scores aggregate into rankings

Tournaments

Compete for prizes using these scoring rules