Overview
Metaculus uses sophisticated scoring algorithms to evaluate forecast accuracy. Scoring incentivizes truthful probability estimates and rewards forecasters who make bold, accurate predictions.Score Types
Metaculus provides multiple scoring methods, each measuring different aspects of forecasting skill. Fromscoring/constants.py:4-10, the available score types are:
Peer Score
Measures how much better you forecast compared to the community
Baseline Score
Measures raw forecasting accuracy against a uniform baseline
Spot Peer
Peer score evaluated at a specific point in time
Spot Baseline
Baseline score evaluated at a specific point in time
Peer Score
Peer scores measure how much better (or worse) your forecasts are compared to the community prediction.Algorithm
Fromscoring/score_math.py:152-200, peer scores use logarithmic scoring:
Key Properties
Peer scores are relative: A positive peer score means you beat the community, while a negative score means the community was more accurate.
- Scale: Typically ranges from -50 to +50 points
- Zero sum: Peer scores across all forecasters sum to approximately zero
- Requires multiple forecasters: Need at least 2 forecasters (formula includes
num_forecasters / (num_forecasters - 1))
Time-Weighted
Peer scores are weighted by coverage - the fraction of the forecasting period your prediction was active:Baseline Score
Baseline scores measure raw forecasting accuracy against a naive baseline prediction.Algorithm
Fromscoring/score_math.py:62-108, baseline scoring:
Baseline Assumptions
- Binary/Multiple Choice
- Numeric/Date/Discrete
Baseline: Uniform distribution over all available optionsFor a binary question, the baseline is 50/50. For a 4-option multiple choice, each option gets 25%.Maximum score: 100 points (achieved with 100% confidence in the correct outcome)Minimum score: Negative infinity (as probability approaches 0 for the correct outcome)
Key Properties
- Scale: 0 to 100+ points (higher is better)
- Absolute: Scores don’t depend on other forecasters
- Always calculable: Can be computed even with only one forecaster
- Incentivizes boldness: Confident, accurate predictions earn more points
Spot Scores
Spot scores evaluate forecasts at a single point in time rather than across the entire forecasting period.Spot Scoring Time
Fromquestions/models.py:428-441, the spot scoring time is determined by:
- Explicit
spot_scoring_timeif set cp_reveal_time(when community prediction is revealed)actual_close_timescheduled_close_time
Spot Baseline Score
Fromscoring/score_math.py:111-149:
Use Cases
Spot scores are ideal for:- Questions where the resolution can affect forecasting (CP hidden until reveal)
- Live forecasting events with synchronized scoring times
- Questions that resolve quickly after closing
- Preventing score manipulation by rapid forecast updates
Score Model
Scores are stored per user, per question, per score type (fromscoring/models.py:17-57):
Coverage
Thecoverage field tracks what fraction of the forecasting period you participated in:
- Coverage = 1.0: You had an active forecast for the entire question period
- Coverage = 0.5: You forecasted for half the question period
- Coverage = 0.0: You never made a forecast
Archived Scores
Historical scores that can’t be recalculated are archived (fromscoring/models.py:59-94):
scoring/constants.py:13-14):
Default Score Type
Each question designates a default score type (fromquestions/models.py:91-102):
Question Weighting
Questions can have different weights when aggregating scores (fromquestions/models.py:90):
- Weight = 1.0: Standard question (default)
- Weight > 1.0: More important question (higher contribution to leaderboards)
- Weight = 0.0: Question excluded from scoring entirely
- Weight < 1.0: Less important question
Unsuccessful Resolutions
Some questions cannot be resolved normally:- Ambiguous: Resolution criteria cannot be clearly applied
- Annulled: Question was flawed or should not have been asked
questions/constants.py:4-6:
Questions resolved as ambiguous or annulled are typically excluded from scoring and leaderboards.
Scoring Best Practices
Maximize Coverage
Maximize Coverage
Forecast early and update regularly to maximize your coverage. Even a simple initial forecast can boost your score potential.
Be Calibrated
Be Calibrated
Over many forecasts, events you predict at 70% should happen about 70% of the time. Track your calibration!
Be Bold When Warranted
Be Bold When Warranted
The scoring system rewards confident, accurate predictions. If you have strong evidence, don’t be afraid to make extreme forecasts.
Update on New Information
Update on New Information
Scores are weighted by time, so updating your forecast when you learn new information helps your score.
Understand the Baseline
Understand the Baseline
To beat the baseline, you need to do better than uniform/naive probabilities. Consider what “uninformed” would predict.
API Reference
Scores API
Explore the full Scores API documentation
Related Topics
Questions
Understand question types and structure
Forecasting
Learn how to make predictions
Leaderboards
See how scores aggregate into rankings
Tournaments
Compete for prizes using these scoring rules
