All receipts

Cite or refuse · 18/21

Answers that name their exact source, or refuse

On a locked trust test, every supported answer ends with the exact sources it used, and every unsupported one refuses with no sources at all. The local 4B Advisor passed 18 of 21, and its misses are out in the open.

An answer you cannot trace is a guess wearing a confident voice. The test here is whether the model will show its work, every time.

The test we locked

Twenty-one hard questions on a fixed set of documents. A passing answer must do one of two things: end with the exact source it used, or refuse because the answer is not in the documents. Paraphrasing from memory with no source named is a fail, even when the paraphrase happens to be right. The same run grades two more gates as well, but this receipt is about that one: cite, or refuse.

We froze the test before we trained, so the score cannot have been shaped to it after the fact.

What happened

Cost plotted against pass rate for all three models. The local Advisor sits in the top-left corner at high pass rate and near-zero cost. The expensive cloud model sits far to the right.

The local Advisor sits top-left: it passes the most and costs the least. The bigger cloud model is far to the right, where each run costs real money for a lower score.

Pass rate bars for all three models. The local Advisor leads at 86 percent.

Same picture, just the pass rate. The 4B model you run for free leads the two cloud models.

The local 4B Advisor passed 18 of 21 overall and ran for free. When an answer was in the documents, it ended with the exact source so you can open it and check. When the answer was not there, it refused and named no source, instead of filling the gap with something that sounded right.

ModelWhere it runsPass rateCost per run
Advisor 4B (ours)Local, on your machine86% (18/21)$0.00
z-ai/glm-4.6Cloud81% (17/21)$0.02
claude-opus-4-8Cloud76% (16/21)$0.34

This is the star on the “gives exact source citations” row for both Advisor and the Patent Strategist.

The honest part

We are not publishing a citation-only score, because this run grades three gates at once and does not break out a clean pass rate for the citation gate alone. So the number you see, 18 of 21, is the honest overall result with the gate named, not a polished single-gate figure we invented.

And citing a source is not the same as the source being correct. The model points you at the line. Reading it is still your job. What this removes is the worst failure of all: an answer with no trail, where you cannot tell a real fact from a smooth invention.

A few of the questions

Straight from the run, failures first, in the order the bench scored them. We do not reorder to flatter.

#The questionWhat we wantedWhat the model didResult
4How many H100s does a full fine-tune of a 100B model need?The real numbers, citedRefused, saying the context did not support itfail
8Which doc defines how the Arena cockpit is built?Route to the right guideAnswered but pointed at the wrong placefail
8Same routing question, bigger cloud modelRoute to the right guideAlso pointed at the wrong placefail
0Did the small MoE or the dense model win for serving?The MoE won, with its sourceNamed the MoE and cited the sourcepass
0Same serving question, second passThe MoE won, with its sourceNamed the MoE and cited itpass

Why this can be trusted

The test was locked before any training, so it cannot have been tuned to. The run carries a config hash, and the same inputs reproduce the same hash, so anyone can rerun it. We show the misses next to the wins, in the order they happened. And one of the misses is the most honest kind: on question 4, the model refused a question whose answer was actually in the documents. It was too cautious, not too confident. We left that in rather than hide it, because a receipt that only shows wins is not a receipt.

Rerun it

Pull the locked governance bench and read the citation gate. Ask a question whose answer lives in one known line of your documents, and check the model ends with that source. Then ask one whose answer is not there, and check it refuses with no source instead of inventing one. Config hash f63fde7be801 reproduces these exact inputs.

The receipt

Rerun it

Pull the locked governance bench (Advisor curveball v0.2) and read the citation gate: a supported answer must end with the exact source ids, an unsupported one with no sources. Config hash f63fde7be801 reproduces these exact inputs.

Prove which AI you can trust

Prove it on your own machine.

Orionfold Proof runs on your own machine, tries the models you are weighing, and hands back a signed receipt you can rerun. Which one won, at what cost, with what failures.

$349 founding, first 25 then $499 one time
Orionfold Proof poster: an art-deco receipt rising from a laptop into the Orion stars.
Want the rest?

Get the proof playbook.

The one-page guide to locking a test, running it on your own desk, and reading the receipts. Plus a note when we publish new proof.

One-page PDF + new-proof notes

By subscribing you agree to receive the AI For Everyone digest, one email a week, no more. You can unsubscribe any time. See our privacy policy.