Model Meets Reality · compare

Grading a prediction, or grading the mechanism behind it

Metaculus is the serious version of public forecasting: proper scoring rules, calibration curves, years of resolved questions. The difference here is the unit being graded.

The short answer

Metaculus asks whether your number was right. This asks whether the mechanism that produced the number was right — and what would retire it if not.

On Metaculus the object is a question, and a forecaster's skill is the accuracy of their probabilities across many of them, scored against a baseline and against other forecasters. It works, and the calibration feedback is real.

Here the object is a model: premises, a mechanism, falsifiable consequences, and a deletion clause. Predictions are how the model gets tested, but they are not the point — the point is whether the mechanism holds. So a claim carries the mechanism that drove it, and a model that is repeatedly wrong is not just badly scored, it hits its own retirement condition.

Two consequences. First, being wrong is informative rather than shameful: refuted models stay listed with their records showing, because a mechanism that failed where it said it would is more useful than one that never committed. Second, models are portable — each is a git repo you can clone, edit and run in your own assistant, not an account on a platform.

Side by side

MetaculusModel Meets Reality
Grades a predictionGrades the mechanism that produced it
Skill = calibration across questionsA model states its own retirement condition
Community forecasts aggregateEach model is a separate falsifiable theory
Lives on the platformEach model is a git repo you can clone and edit
Years of resolved questionsOpening small; no independent resolution yet

What Metaculus does better

Almost everything measurable. Metaculus has a decade of resolved questions, proper scoring rules, per-user calibration curves, and since February 2026 the FutureEval benchmark tracking how AI forecasters compare to humans. It has the track record this does not. If you want to find out how well-calibrated you are, go there — nothing on this site can tell you that yet.

Where this stands today

Said plainly, because the whole point of this site is not overclaiming. The Model Garden is opening small — no models listed as of September 2026. Every record shown is self-graded: the author's own count of claims made and resolved, labelled as such on each card. Nothing here is ranked, and no independent resolution layer exists yet. Metaculus has things this does not, named above rather than omitted. Compare the designs, not the scoreboards — there is no scoreboard.
https://github.com/someone/their-model Help me use this

Paste that into any assistant. It reads the model — premises, falsifiers, deletion clause — and reasons through it.

The difference that is not on the table

Metaculus grades your prediction. Here what accrues a record is the reasoning behind it. Both let you register a claim before the outcome is known, and that overlap is real. The difference is the unit: there, a resolved question is a score; here, the mechanism you claimed would produce it is the thing that survives or fails, across every claim it makes. You can be right about an event for reasons that do not generalise — this is built to catch that.

Share your understanding →

Browse the Model Garden Publish your own