Skip to main content

Overview

The Scores endpoints provide detailed information about how forecasts are scored on Metaculus. Scores measure forecast accuracy using proper scoring rules and are the basis for leaderboards and performance tracking.
Scores are typically accessed through leaderboard endpoints and post data downloads. These endpoints provide raw scoring data for advanced analysis.

Understanding Metaculus Scoring

Score Types

Metaculus uses several scoring methods to evaluate forecasts:
Score Type
Peer Score: Measures performance relative to the community aggregate prediction. Rewards forecasters who beat the crowd.Calculated as: Your log score - Community log score
Score Type
Baseline Score: Measures performance relative to a simple baseline prior (e.g., uniform distribution for continuous questions, 50% for binary).More generous than peer score, useful for beginners.
Score Type
Spot Peer Score: Peer score evaluated at a specific time (CP reveal time) rather than continuously weighted.Used for tournament scoring to prevent gaming through frequent updates.
Score Type
Spot Baseline Score: Baseline score evaluated at CP reveal time.
Score Type
Legacy Relative Score: Historical scoring method from old Metaculus. Deprecated.

Scoring Mechanics

How Scores Are Calculated
  1. Log Score: Your forecast is scored using logarithmic scoring rules. For binary questions with outcome O and prediction p:
    • Score = log₂(p) if O = Yes
    • Score = log₂(1-p) if O = No
  2. Continuous Questions: CDF is converted to PMF, then log score is calculated based on the probability mass assigned to the actual outcome.
  3. Coverage: Your score is weighted by what fraction of scored questions you forecasted. Higher coverage = more reliable score.
  4. Aggregation: Scores across questions are averaged with question weights applied.

Download Score Data

Score data is primarily accessed through the data download endpoints:

Post-Level Scores

See the Posts endpoint documentation for full details on the download-data endpoint.

Project-Level Scores

See the Projects endpoint documentation for details.

Score Data Format

When you download score data, you receive a CSV with the following schema:

Score Data CSV Schema

integer
The question ID this score is for
integer
The user ID who earned this score
string
The username of the scorer
string
Type of score: peer, baseline, spot_peer, spot_baseline, relative_legacy, or manual
number
The score value. Higher is better. Can be negative.
number
The coverage value (0-1) representing what fraction of time the user had an active forecast

Accessing Scores in Aggregations

Scores for community aggregations are included in question data when using with_cp=true:

Example: Analyze User Performance


Example: Coverage Analysis


Example: Historical Score Tracking


Important Notes

Score Data AccessIndividual user scores are only available:
  • To the user themselves
  • To site administrators
  • In aggregate form (leaderboards)
You cannot download other users’ detailed score data for privacy reasons.
Score TimingScores are calculated:
  • When questions resolve
  • When leaderboards are updated (typically daily)
  • When explicitly recalculated by admins
There may be a delay between question resolution and score appearance.
Coverage MattersHigh coverage (forecasting many questions) makes scores more reliable and statistically meaningful. Users with low coverage may have high variance in their scores.