Grading a prediction, or grading the mechanism behind it
Metaculus is the serious version of public forecasting: proper scoring rules, calibration curves, years of resolved questions. The difference here is the unit being graded.
The short answer
On Metaculus the object is a question, and a forecaster's skill is the accuracy of their probabilities across many of them, scored against a baseline and against other forecasters. It works, and the calibration feedback is real.
Here the object is a model: premises, a mechanism, falsifiable consequences, and a deletion clause. Predictions are how the model gets tested, but they are not the point — the point is whether the mechanism holds. So a claim carries the mechanism that drove it, and a model that is repeatedly wrong is not just badly scored, it hits its own retirement condition.
Two consequences. First, being wrong is informative rather than shameful: refuted models stay listed with their records showing, because a mechanism that failed where it said it would is more useful than one that never committed. Second, models are portable — each is a git repo you can clone, edit and run in your own assistant, not an account on a platform.
Side by side
| Metaculus | Model Meets Reality |
|---|---|
| Grades a prediction | Grades the mechanism that produced it |
| Skill = calibration across questions | A model states its own retirement condition |
| Community forecasts aggregate | Each model is a separate falsifiable theory |
| Lives on the platform | Each model is a git repo you can clone and edit |
| Years of resolved questions | Opening small; no independent resolution yet |
What Metaculus does better
Almost everything measurable. Metaculus has a decade of resolved questions, proper scoring rules, per-user calibration curves, and since February 2026 the FutureEval benchmark tracking how AI forecasters compare to humans. It has the track record this does not. If you want to find out how well-calibrated you are, go there — nothing on this site can tell you that yet.
Where this stands today
https://github.com/someone/their-model Help me use this
Paste that into any assistant. It reads the model — premises, falsifiers, deletion clause — and reasons through it.
The difference that is not on the table
Share your understanding →