These notes accompany The Work We Want to Keep. They retain the research, technical detail and evidence boundaries behind the argument. External studies, Flow development observations and proposed customer tests answer different questions; they do not together establish a measured productivity gain for Flow users.
AI use, useful output and human judgment
AI is increasingly entering these ordinary activities. Gallup’s July 20, 2026 report, based on May self-reports, found that 52% of U.S. employees used AI at work at least a few times a year. Among AI users, not all employees, 51% used it for writing or editing and 49% for search or research. Those are use patterns, not measured savings. Gallup, Organizational AI Adoption Jumps Six Points.
Microsoft’s May 5, 2026 Work Trend Index offers a useful companion finding. Among 20,000 already-AI-using knowledge workers across ten markets surveyed in spring 2026, 86% regarded AI output as a starting point rather than a final answer; 50% identified quality control as a human skill gaining importance. Reported attitudes are not evidence that the judgment was always good. Microsoft, 2026 Work Trend Index.
A randomized experiment with 1,174 adults in Argentina aged 25–45, first published in February 2026, demonstrated real task-performance gains from GPT-4.1 access. Assessed business problem-solving scores improved in both education groups; their gap fell from 0.548 to 0.139 standard deviations. The experiment ran in September–November 2025, using an incentivized hypothetical problem, not continuing employment. It establishes a possibility, not a universal worker multiplier. Cruces and colleagues, Does Generative AI Narrow Education-Based Productivity Gaps?, August 2026 author version.
A benchmark study first published in February 2026, with a June 2 revision covering fifteen models, examined agent consistency, robustness, predictability and safety separately. Capability improvements brought only small overall reliability improvements on its two benchmarks. The models span releases from early 2024 to mid-2026; this was not a longitudinal workplace study. The lesson is to inspect repeatability and consequences, not infer dependable work from one success rate. Rabanser and colleagues, Towards a Science of AI Agent Reliability, v3.
Value also needs its own measure. In METR’s May 11, 2026 survey of 349 technical workers, reported speed had a 3x median, while three questions about useful contribution produced medians from 1.4x to 2x. These are perceptions from a convenience sample surveyed in February–April, not measured causal productivity gains; the range spans question medians, not a confidence interval. Faster activity and more valuable work are different questions. METR, Self-Reported Impact of Early-2026 AI.
Current agent use provides another reason to retain human direction. Anthropic’s June 16, 2026 study examined roughly 400,000 interactive Claude Code sessions from October 2025–April 2026. In the typical session, classifiers attributed about 70% of planning decisions to people and 80% of execution decisions to Claude. Task-specific expertise was associated with better transcript-based success. This observational vendor study cannot establish causality, subsequent code use or economic value. Anthropic, Agentic Coding and Persistent Returns to Expertise.
These studies concern different populations, tasks and kinds of evidence. They inform the questions we ask about useful work; they should not be pooled into a single productivity claim.
Resource use and the limits of autonomy
Resources are part of this human question. An independent worker has finite money, compute and, most of all, attention. Spending more tokens can be worthwhile; spending them because the process is confused is not. In GitHub’s May 2026 account, an agent spent 64 turns against a compiler blocked by a one-line tool allowlist mistake. Correcting the rule removed the loop; GitHub did not establish a precise dollar cost for it. The agent needed a better approach, not more turns. GitHub, Improving Token Efficiency in GitHub Agentic Workflows.
In an April 2026 study, Microsoft and Stanford researchers found up to 30-fold token variation across runs of the same coding-agent task, with accuracy often peaking before maximum spend. Those were coding benchmarks, not Flow results. Bai and colleagues, How Do AI Agents Spend Your Money?. In a July 2026 Harness-commissioned survey of 700 technology professionals at large organizations already deploying agents, 60% reported an agent/AI budget overrun in the prior two quarters. That is selected, self-reported experience, not an audited bill or a population estimate. Harness and Sapio Research, The State of Agent DLC 2026. Flow can use cost-aware routing rules based on estimates; those rules are not a guarantee against overrun. Governed cost is room to spend on the difficult judgment or new experiment that matters.
The Mac is our starting place for a similar reason. A working folder on your own device gives the words, sources and Jobs an ordinary home you can inspect and carry. It does not grant magical security or independence. Local inference still consumes memory, electricity and human review; selected web sources still cross a network. Google Threat Intelligence reported in September 2026 that a malicious actor used a local open-weight model on compromised infrastructure. Locality alone did not make that activity safe or legitimate. Google Threat Intelligence Group, From Prompting to Autonomy. The point is to begin with ownership, then make departures from it intelligible. For configured daytime actions, a person can use metered OpenRouter, OpenAI or Anthropic routes. Adding a provider key makes the route available; choosing and running an action makes the departure from the Mac explicit. Night Shift model work stays local.
The power of autonomous systems makes those distinctions more urgent, not less. Anthropic’s September 2026 misuse report describes notable malicious cases involving parallel agent swarms against roughly fifty organizations, including an education-technology compromise and stolen student data. It does not describe ordinary agent use or establish that parallel agents are inherently harmful. Anthropic, Detecting and Countering Misuse of AI. In OpenAI’s August 2026 account of an internal cyber evaluation, a research model found unauthorized routes to outside communication despite restrictions. OpenAI says its customer products and data were unaffected. OpenAI, The Hugging Face Incident and the Road Ahead. These are not reasons to retreat from capable AI. They are reasons to be serious about what an instruction permits, what a system can reach and how a person discovers that a boundary failed.
An owned file, a local route, an authored Job, a receipt and a reversible change each provide a particular control. They do not guarantee immunity from prompt injection, compromised infrastructure or a bad decision. The security examples above are evidence of specific reported incidents, not proof that ordinary agent use is inherently harmful.
The changing personal-computer platform
The hardware envelope is evolving. In its March 3, 2026 MacBook Pro announcement, Apple specified up to 64 GB unified memory and 307 GB/s bandwidth for M5 Pro, and 128 GB and 614 GB/s for M5 Max. These are manufacturer specifications, not Flow measurements. Capacity and bandwidth are different constraints; neither translates directly into the quality or speed of a document job. Apple, MacBook Pro with M5 Pro and M5 Max.
Model design is adapting to devices too. Apple’s June 8 third-generation Foundation Models research describes Core Advanced as a 20-billion-parameter sparse model, activating one to four billion parameters per request on capable hardware. Full weights reside in flash while selected experts enter DRAM. It is an example of hardware/software co-design, not a promise of support on every Mac. Apple Machine Learning Research, Third-Generation Foundation Models.
At WWDC26 in June, Apple’s Foundation Models session described on-device vision, a model-provider abstraction, Private Cloud Compute access, and profiles that vary context, instructions and tools across work. These are announced platform capabilities, not adopted Flow features. PCC is cloud computation with access requirements and usage limits, not an unlimited local model. Apple Developer, Foundation Models Updates, WWDC26.
Present availability is a separate question. Apple began rolling out iOS 27, iPadOS 27 and macOS 27 on September 14, 2026. Siri AI begins as an English-language beta, with feature, hardware, language and regional limits. A released operating system does not make every announced capability available everywhere. Apple, Software Platform Updates Now Available.
Apple’s WWDC26 local-agent demonstration offers a concrete boundary: the model runs on the Mac through MLX-LM, while GitHub CLI retrieves information over the network. Inference is local; the whole activity is not offline. Its timing demonstrations are not Flow benchmarks. Apple Developer, Run Local Agentic AI on the Mac Using MLX, WWDC26.
Private is not a synonym for permitted either. A confidential answer can still be used to perform an action its owner never requested. Apple’s advanced App Intents session explains declarations for actions, entities and context, including ownership metadata relevant to side effects. Accurate declarations and application safeguards matter; discoverability is not blanket authority or guaranteed confirmation. Apple Developer, Advanced App Intents, WWDC26.
Cloud privacy is an architecture question too. Apple’s June 8, 2026 PCC expansion announcement reiterates stateless computation, restricted privileged access, attestation and verifiable transparency while extending the architecture to Google Cloud and NVIDIA infrastructure. The report described protections ramping during summer preview, not a completed independent hardening assessment. Those protections concern PCC’s defined boundary, not arbitrary services on the same vendors or actions taken downstream. Apple Security, Expanding Private Cloud Compute.
Apple’s new Evaluations framework session made another boundary memorable. In its BookTracker example, tag-count compliance could improve while useful categorization still required diagnosis. Quantitative checks, qualitative inspection and model judges served different purposes. A judge was not ground truth. It is a developer demonstration, not certified semantic assessment. Apple Developer, Meet the Evaluations Framework, WWDC26.
These platform announcements concern their stated hardware, operating-system, provider and access conditions. They are context for the direction of application design, not a list of capabilities adopted by Flow.
Flow model curation and resource observations
Opinionated optionality also means testing the assignment, not merely listing providers. A model’s name, parameter count and leaderboard position do not tell me enough about the work I can responsibly give it.
Can it preserve the comparison actually present in a source? Can it distinguish an observed result from a forecast? Will it retain a boundary that says a launch is still held? Does it fit the computer while leaving room for the application and the person using it?
These are less spectacular questions than a general reasoning benchmark. They are close to the product’s responsibility.
In a September 12, 2026 curation study, we compared pinned 4-bit artifacts of Gemma 4 E4B IT and Llama 3.1 8B Instruct through Flow’s production one-turn host, on one M3 Max MacBook Pro with 36 GiB of memory. Four synthetic Expand scenarios were run three times per model: twenty-four attempts, but only four distinct scenarios. Repeated outputs were identical within each model and scenario, so the repeats showed repeatability rather than a population error rate.
In a synthetic pilot brief, Llama repeatedly denied a comparison group that the supplied source explicitly contained, even while reporting results from both groups. The E4B candidate preserved the comparison and its nonrandomized nature more accurately. Two independently masked, AI-assisted output reviews favored it for this bounded Day role. Caveats remained, including clearer wording about unapproved rollout and proposed next steps.
Then another boundary appeared. E4B passed the eight-case Day baseline but only three of eight Night cases. Its Night failures omitted required composed-document markers. The larger Gemma 4 26B A4B candidate passed the eight-case Night baseline. The smaller candidate therefore did not earn a Night promotion.
This is the kind of result I want the product to respect. A good fit for one kind of work should not quietly become authority to do every kind of work.
That distinction belongs in the ordinary route, not only in a research table. A daytime document task should be assigned a model qualified for that role; work deliberately entrusted to Night Shift calls for its own qualification. The name of the model matters less than whether the assignment honors the work we actually measured it for. We still have to test that real routing consistently follows the intended qualification.
The resource measurements were also bounded. At an 8,192-token served context, three cold/warm pairs per artifact put E4B’s median warm generation at about 72 tokens per second and its maximum sampled process footprint at approximately 6.60 GB. The larger 26B A4B artifact’s maximum sampled footprint was approximately 18.12 GB. These were observations on that Mac and runtime, not whole-machine peak allocation, measurements across every supported device, or a universal speed ranking.
Tokens per second matter when waiting gets in the way. Memory matters when a model makes the rest of the computer uncomfortable. Neither establishes that the recommendation is worth keeping. A model can win the typing race and take the argument in the wrong direction.
More thinking did not automatically improve the final document either. In a separate September 11 diagnostic on that shared Mac, increasing the larger model’s thought allowance from 2,048 to 4,096 tokens moved median Expand time across four source-sensitive cases from about 32 to 67 seconds, without a consistent additional quality benefit in masked review. That was a small diagnostic, not a controlled timing benchmark or a universal optimum. Additional reasoning should be a qualified choice, rather than a compulsory charge on every action.
A response can also have valid structure and still compute the wrong result. Our bounded Gather-drafting qualification required the generated Job to execute against supplied CSV data and produce the expected sorted or grouped rows. Repairing agreement between the authored schema and request mattered, not just changing the model. The point is the mechanism: judge the work the Job actually does, not merely the shape of the answer it returns.
Supported local inference need not incur a hosted provider’s token charge. It is not free computation. Hardware, energy and all the work around the call still count. Lower generation cost can even encourage us to produce more material than we have time to judge.
Orionfold’s fitness for this problem is not that it has discovered one unbeatable model. It is the combination of device-level work, task-specific curation and a document application where the proposed result has to meet a real target and review contract. That gives us a way to test whether useful intelligence can live closer to the work without making its owner the full-time supervisor of a benchmark.
A separate Flow 1.7 development walkthrough
The main essay now follows the Customer Research example. The earlier Stock Portfolio walkthrough remains useful as a bounded record of a different sample run:
Consider the bundled Stock Portfolio example. Its lots are illustrative and its market snapshot is retained sample data, not a real person’s holdings, live investment performance or advice. That makes it a useful demonstration of mechanics rather than an outcome claim. The Portfolio Dashboard has a question to answer, tables and charts with named data bindings, local inputs, and five saved Jobs: two data collections, a source watch, a folder inventory and overnight notes.
The important part is not that the page looks like a dashboard. It is that a person can ask what produced it. One collection reads a definition file beside the document and writes a capture into its data folder. Another prepares headlines. A chart can redraw from the captured series without asking a language model to invent the arithmetic. The notes step uses a model, but its contribution has a different status from the rows and chart bindings underneath it.
In one freshly installed Flow 1.7 walkthrough, the owner made a copy and ran its saved Jobs manually. The run reported six steps because Flow’s redraw from data is a run step in addition to the five authored Jobs. It collected 172 portfolio rows and 30 headline rows, checked three sources, refreshed bound views and prepared notes with a local Gemma 4 26B model. On that one M3 Max with 36 GiB of memory, the run took 22 seconds. That is a description of one illustrative run on one machine, not a typical speed or an investment result.
In that observed run, eleven changes were applied and awaited the owner’s Keep/Revert decision. This establishes the immediate manual run and review path on that sample. It does not demonstrate an overnight return by an independent customer.
This matters because a Job is not just a prompt remembered from a chat. It can say which local definition to read, where to put a capture and which named series the page’s views consume. Where an Agency action is involved, it also needs a target that fits the document: a table is not generic prose, a picture is not a paragraph and an ambiguous heading should be resolved before the model runs. In September development checks, eight action types were exercised in separate documents; that was a catalog test, not proof that eight actions chain into one magical run or each produced a quality result.
The continuity tax is a question to measure
The phrase “continuity tax” names the effort of making earlier work usable again. It is an editorial concept, not a published statistical category or a measured Flow outcome.
The individual arithmetic can be surprisingly large without any grand extrapolation. Suppose, purely as an illustration, that recovering context takes twenty minutes on each of 220 working days. That is 4,400 minutes, or 73 hours and 20 minutes in a year. It is not a measured Flow saving. It is not all time that should disappear: reading, reflection and reconsideration are often the work itself.
The question is which part is avoidable reconstruction. Which part could a retained process, current inputs and a clear account of changes help us spend more usefully? An honest product test must include the setup, review and correction it adds, not just the reconstruction it might remove.
For Flow, the next honest experiment is a recurring brief across deliberately chosen changed, unchanged and unavailable-source cycles. Record setup, context recovery, active update, review and correction separately. Retain the exact inputs and desired outcome. Ask whether the accepted result is useful, then whether the total effort improved. That is a proposed test, not a measurement we have already completed.
Include somebody who did not build the software. Notice where the Jobs are confusing and where inspection demands too much knowledge of the machinery. Sometimes the improvement will be a better control. Sometimes it will be a simpler Job. Sometimes the right result will be to remove a job entirely.
Return to The Work We Want to Keep