Understanding CDFs on Metaculus
Metaculus requires continuous forecasts as a 201-point CDF - a list of 201 probability values representing the cumulative probability at evenly-spaced points across the question’s range.Key Concepts
1
What is a CDF?
A CDF at point
x represents the probability that the outcome is less than or equal to x.- First value (index 0): Probability that outcome is below the lower bound
- Middle values (indices 1-200): Probabilities at evenly-spaced points within the range
- Last value (index 200): Should always be 1.0 (or close to it)
2
Question Scaling
Questions can have:
- Linear scaling: Points are evenly spaced in the actual scale (e.g., 0, 10, 20, 30…)
- Logarithmic scaling: Points are evenly spaced on a log scale (useful for wide ranges like 1 to 1,000,000)
- Open vs. closed bounds: Open bounds require probability mass outside the range
3
CDF Requirements
Your CDF must:
- Have exactly 201 values (or
inbound_outcome_count + 1for discrete questions) - Be strictly increasing by at least 0.00005 per step (1% / 200)
- Not increase by more than 0.2 at any single step
- Respect boundary conditions (open vs. closed)
Getting Question Scaling Information
First, retrieve the question’s scaling parameters:Complete CDF Generation Functions
Here’s production-ready code from the Metaculus OpenAPI specification:Converting Nominal Values to CDF Locations
This function converts a real-world value (e.g., “500 deaths” or “2025-06-15”) to the internal [0, 1] scale:Python
Generating CDF from Percentiles
This is the recommended approach - specify a few key percentiles and generate a full CDF:Python
Standardizing the CDF
This function ensures your CDF meets all Metaculus requirements:Python
Complete Example: Linear Scale, Closed Bounds
Here’s a complete workflow for a simple case:Complete Example: Open Bounds
For questions with open bounds, you must assign probability mass outside the range:Python
Date Questions
Date questions work the same way, but use ISO format timestamps:Python
CDF Validation Rules
Your CDF will be rejected if it violates these rules:Rule 1: Strictly Increasing
Rule 1: Strictly Increasing
The CDF must increase by at least 0.00005 (0.005%) at each step.
Rule 2: Maximum Step Size
Rule 2: Maximum Step Size
No step can increase by more than 0.2 (20%).
Rule 3: Boundary Conditions
Rule 3: Boundary Conditions
- Closed lower bound: First value must be 0.0
- Open lower bound: First value must be at least 0.001 (0.1%)
- Closed upper bound: Last value must be 1.0
- Open upper bound: Last value must be at most 0.999 (99.9%)
Rule 4: Length
Rule 4: Length
Must have exactly
inbound_outcome_count + 1 points (usually 201).Common Errors and Solutions
Error: 'Percentiles must encompass bounds of the question'
Error: 'Percentiles must encompass bounds of the question'
Problem: Your percentiles don’t cover the full range from 0 to 1.Solution: Either:
- Add extreme percentiles (e.g., 1st and 99th)
- Specify
below_lower_boundandabove_upper_boundparameters
Error: 'CDF not increasing fast enough'
Error: 'CDF not increasing fast enough'
Problem: Your distribution is too concentrated (too much probability in one place).Solution: Use
standardize_cdf() which adds a uniform component to ensure minimum increase rates.Error: 'Step size too large'
Error: 'Step size too large'
Problem: Your distribution has too sharp a spike.Solution: Spread out your percentiles more evenly, or use
standardize_cdf() which caps maximum step size.Tips for Better CDFs
- Start with percentiles: It’s much easier to think in terms of “I believe there’s a 50% chance the answer is below X” than to manually construct 201 probability values.
- Use more percentiles for complex beliefs: If you have a bimodal or unusual distribution, specify more percentiles (10th, 20th, 30th, etc.).
-
Always standardize: The
standardize_cdf()function ensures your CDF will be accepted and adds a small uniform component that actually improves forecasting performance. - Check your work: Print out key percentiles from your generated CDF to verify it matches your beliefs:
- Test with closed bounds first: Start by practicing with questions that have closed bounds - they’re simpler to work with.
