How this compares
Honest comparisons, including what each tool does better. Several of these are not competitors at all — a model from here runs inside ChatGPT, Claude or Gemini.
Assistants a model runs inside
ChatGPT
Ask ChatGPT a question and you get its best synthesis. Paste a model first and you get an answer produced by one stated mechanism — which can therefore be wrong in a specific, checkable way.
Gemini Notebook (formerly NotebookLM)
Gemini Notebook answers only from what you uploaded. A model answers from a mechanism that states, in advance, the observation that would falsify it.
Claude Projects
Project instructions shape how an assistant behaves. A model states a claim about the world that can be graded — and moves to any assistant, unchanged.
Custom GPTs
A custom GPT is judged by whether people like using it. A model is judged by whether its predictions hold.
Multi-model tools
Forecasting
Metaculus
Metaculus asks whether your number was right. This asks whether the mechanism that produced the number was right — and what would retire it if not.
Prediction markets
A market tells you what the crowd believes. A model tells you why something happens, and what observation would show it is wrong.
Publishing and reference
Substack and blogs
A post argues a position. A model registers what would refute it and when, so the check happens whether or not anyone remembers to look.
arXiv and preprint servers
A preprint reports what was found and invites scrutiny. A model states the test before the outcome exists, and the timestamp proves it.
Wikipedia
Wikipedia summarises what is known. A model proposes how something works and names what would show that it does not.
Registries
Where this stands today
Said plainly, because the whole point of this site is not overclaiming.
The Model Garden is opening small — no models listed as of September 2026.
Every record shown is self-graded: the author's own count of claims made and
resolved, labelled as such on each card. Nothing here is ranked, and no independent
resolution layer exists yet. Each tool listed here has things this does not, named above rather than
omitted. Compare the designs, not the scoreboards — there is no scoreboard.
https://github.com/someone/their-model Help me use this
Paste that into any assistant. It reads the model — premises, falsifiers, deletion clause — and reasons through it.
The difference that is not on the table
All of these are worth using. None of them will tell you that you were wrong. That is the gap this fills: a place to write down how you think some corner of the world works, name what would falsify it, and let the date arrive. If you have watched one thing closely for years, you already have the model — it just has never been tested.
Share your understanding →
Share your understanding →