The proof, one receipt at a time.
Each claim on its own page: the test we locked, the run, and how you rerun it. The wall is at /proof/. These are the receipts behind it, one link each.
- 76% faster Speed
The same model, 76% faster, same chip
No new chip, no new model. Just a leaner 4-bit mode the chip runs in hardware, and the same model typed 76% faster on a fraction of the memory. Rerun it yourself.
See the receipt → - 38 of 44 checked Reasoning
One number you can actually check
For math, the right shape of answer is a number you can verify, not prose. Kepler returns the number, a checker compares it to the known answer, and on a locked test it passed 38 of 44. Rerun it yourself.
See the receipt → - $1 guards $1,679 Cost
A $1 test that guards $1,679
Before booking a big cloud training run, we test the design on the desk first. The cheap test catches the bad designs, and one wrong booking it stops pays for the whole box. Rerun it yourself.
See the receipt → - 2 cents Cost
An AI lab that ran overnight for 2 cents
Fifty experiments, 73 minutes, about two cents of power, and nothing sent out. The same loop on rented cloud chips runs to dollars plus a fee for every word. Rerun it yourself.
See the receipt → - Show the work Reasoning
Reasoning you can read step by step
A right answer with a hidden path is hard to trust on the next question. This model shows each step, and its score climbs in a clear, checkable way as you hand it better source. Rerun it yourself.
See the receipt → - 18/21 · 86% · $0.00 local Trust
A 4B model on your desk out-trusts frontier cloud
Same 21 hard questions, three models. The small one you can own beat the big ones you rent, at zero cost and full privacy. Rerun it yourself.
See the receipt → - Cite or refuse · 18/21 Trust
Answers that name their exact source, or refuse
A good answer you cannot trace is still a guess. This receipt foregrounds one gate of the test: name the exact source, or refuse. Same run, same 21 questions.
See the receipt → - Refuse gate · 18/21 Trust
A model that refuses when its notes do not hold the answer
Honest scope: this is the refuse gate of the test, the model checking what it can answer from what it was given. It is not the model grading how good its own memory is. That gate does not exist here, so we do not claim it.
See the receipt →
Get the proof playbook.
The one-page guide to locking a test, running it on your own desk, and reading the receipts. Plus a note when we publish new proof.
One-page PDF + new-proof notes
By subscribing you agree to receive the AI For Everyone digest, one email a week, no more. You can unsubscribe any time. See our privacy policy.