Product tour · Part 4 of 4
Know what AI did. Keep the record.
Every approved run records the model, location, cost, checks, evidence, and saved change. Benchmarks show how local models performed on your Mac before you choose.
The proof
One run card bound to the exact document before and after the change.
Receipts
One run. One durable record.
Approve a change and Flow records the model, provider, place, cost, checks, evidence, and exact before-and-after content. Your prompt and your API key are not recorded, and there is structurally nowhere in a receipt for them to go.
The check and the rule stay distinct.
A stored outcome says what happened then. A severity pill says what the rule would do now: Stops the change, Warns you, or Needs your acknowledgment. Flow keeps both facts instead of rewriting history when a rule changes.
The receipt is bound to the bytes.
Before and after fingerprints tie the record to the exact saved document. A validity mark makes that binding checkable later.
Evidence
A score must say what it rests on.
The Evidence gallery leads with the verdict, then checks, dimensions, coverage, uncertainty, and the recorded reason. A dimension with no measurement says Not scored.
Read the receipt's truth boundaries
- A score delta appears only after the baseline revision was verified in History, byte for byte. Otherwise Flow shows its reason instead of a number.
- An estimate is never recorded as a charge. A local run records no charge because none was owed.
- The real billed run in the measured facts recorded $0.00425 from 105 input and 149 output tokens, reproducible to the digit.
- When the same model writes and judges a change, the receipt says so.
The historical hosted-run capture is from 2026-08-26, installed build 1255 (v1.5), driven by the operator. The Evidence settings detail is Flow 1.6 (1899), captured 2026-09-06. The pictured run is real: 0.121375 USD billed for claude-opus-5 through Anthropic. The $0.00425 run in the truth boundaries is a separate measured run.
Benchmarks
Measure the models on the Mac you own.
Flow measures time to first word, reading speed, and memory fit for the local models already on your Mac. On one M3 Max, a 20.4 GB model reached its first word in 4.5 seconds while an 8.5 GB model took 10.2. File size alone would rank them backwards.
- A score ranks these models against each other, on this Mac. It is never about how good the answers are.
- Flow publishes no shared leaderboard and downloads no results.
- The reading window is worked out per model, per machine, and names the limit as Flow's own, not the model's.
- Unmeasured models receive no score. Uneven comparisons are called out instead of normalized away.
Measured 2026-08-18 on one Apple M3 Max, on the running build. Your Mac and models will produce different numbers.
Need the exact limits?
Requirements, formats, model states, plan boundaries, and qualified measurements live in Tech Specs.