Flow Philosophy
A support workflow reads a ticket, asks a billing service whether a charge settled, asks a risk service for a score, waits overnight for a person to approve a refund, and issues it. By morning the process that started the work is gone and the refund call has timed out with nothing returned. What happened has to be reassembled from service dashboards, an application log that recorded what the code intended rather than what it observed, and whoever was on call.
That is the ordinary shape of automated work that touches other systems and waits on people. What goes missing is not the outcome but the account: which questions the program asked, which answers came back, why it continued, and what is still unknown.
Flow exists to make that account part of the program rather than something reassembled afterward. It is a small language for closed, deterministic evaluation whose only way to touch the world is a typed Tool call: a call to a function whose body lives outside Flow. Because every outside observation and every meaning-bearing scheduling choice is explicit, a Flow run is recorded by default, can be replayed from its history without calling anything again, and can pause for days and resume in a new process from that same history.
Flow does not replace the application around it. It replaces the stretch of ordinary code between a decision and an effect, where the interesting crossings hide.
Two commitments
Tools make outside interaction explicit. A Tool is a function whose body
is outside Flow. A program declares it with its full signature,
@tool flow chat(prompt: Text) -> Result<Text, ChatError>, and calls it like
any other function; the host supplies the implementation. Every function that
can reach a Tool says so with !tool, so outside contact is visible in
signatures before anyone reads a body. Flow defines the typed conversation and
nothing about who answers it or where they run.
History carries what computation cannot recover. Flow retains the boundary calls, the outcomes admitted at each one and where those outcomes entered the run, and the runtime choices that deterministic evaluation cannot derive again. An admitted outcome may be a reply, a boundary failure, or no conclusive result at all. Recording is part of running, not logging bolted on beside it. Semantic history retains only what deterministic computation cannot recover, and where a run's whole state is a function and its arguments it can say so with a checkpoint, so long runs stay small. Given the same program, inputs, and a complete record, a run can be explained, replayed or resumed without repeating recorded work.
Replay re-executes the program and serves the recorded observations instead of invoking the tool implementations. That matters when a call can charge money or send a message: a fresh run may do it again, a replay may not. What is missing or withheld cannot be made up by calling out to the world again, and a record that contradicts the program cannot be quietly accepted. A partial history still explains what it holds, but cannot carry a replay.
History has to keep uncertainty too. When a call may have reached the world and no answer came back, the unresolved status stays unresolved and Flow does not retry on its own. Flow promises nothing about exactly-once effects, because the systems where effects land are not Flow's. An admitted reply is a claim a tool made about its own protocol, not evidence that the world matches it.
Why a language
A library can log calls, pause work, and re-run code. The hard part is making all three mean the same thing at once. Replay without repeated effects requires knowing every outside influence on a run, every outcome computation cannot reconstruct, and every point where the run can stop. Those facts have to be complete, not merely plentiful: one unrecorded clock read, thread handoff, or callback is enough to make a replay diverge or an effect repeat.
Inside a general-purpose language, completeness is a matter of discipline. A library can ask authors to route every observation through its API, but it cannot make the other routes absent: another dependency, a callback, a thread, a file read, an environment variable. Convention discourages; only a closed evaluation model excludes. Enforcing closure on a general-purpose host means restricting it to a profile, and that profile is a language design whether or not it is called one. Other languages could carry the same contract by making the same restrictions binding, and a restricted surface that does so is another implementation of Flow rather than an escape from it. Flow chooses to make the closed profile the whole language: small enough to specify completely, checked before it runs, and enforced by the evaluator rather than by review.
So Flow gives up power deliberately. Outside values arrive through declared inputs and typed tool calls, ordinary computation follows deterministic rules, and any other choice that can steer the run without being derivable from them becomes history. Replay and resume read the same record: replay checks it, and resume continues where it ends. One record serving both is what lets each make an honest claim.
What it costs
The cost is real, and it is the strongest argument against adopting Flow. Inside the boundary there is no host ecosystem to reach for, and every outside value the program needs has to arrive as declared input or through a typed tool. The trade pays only where the account matters as much as the answer. Start with one bounded piece of work, like the refund above, and move more in only where that makes it easier to understand.
Where Flow stops
A Tool declaration is a typed invocation boundary. It is not proof of authority, identity, safety, or confinement. A well-typed program can still make a terrible request, an implementation can hold power Flow never sees, and a well-typed reply proves only that Flow could interpret it. Types make the conversation legible; judgment, authorization, and isolation come from the systems that enforce them, and a claim about any of those names the system that made it good.
Flow's history covers one run: the boundary interaction Flow observed and the language-level choices that steered it. It does not reach the work hidden inside a tool, and it is neither general application observability nor the application's system of record. An application may link its own records to a run and its calls; that adds evidence beside the run, not meaning to it.
Applications stay applications. An agent product owns its model loop, sessions, subagents, prompts, and product state. It also owns policy, credentials, approvals, budgets, sandboxing, retries, lifecycle and recovery beyond the run, and system-wide telemetry. Flow runs inside or beside such a system for one bounded piece of work, and the host that embeds it supplies the implementations behind the run's tools. Encoding one product's policy as a language mechanism would force every other host to pretend to implement it.
What has to be trusted
Two pieces of trust are load-bearing. The compiler and runtime are trusted to execute the language and record it honestly, because no record can authenticate its own writer. A history can hold sensitive values, and it is the account itself, so retention, access, and integrity belong to the host that stores it. A record good enough to explain a run is also a record worth stealing, and a holder whose copy answers to nothing outside itself can rewrite it. Anchoring a history where later alteration becomes detectable narrows that; it does not make a record prove itself.
How we decide changes
Contributors use these rules to decide what belongs in Flow.
- Show the missing program. A proposed mechanism arrives with an orchestration that existing deterministic forms, typed tools, and history cannot express clearly. "It would be more convenient" argues for a library.
- Keep outside influence visible. Ask whether anything outside a run can steer it without arriving as declared input, through a Tool call, or as a recorded non-derivable choice. Clocks, entropy, configuration, and human input get no exemption; none arrives through an ambient builtin.
- Count the whole effort. A change is judged by the total effort to write, repair, review and explain correct programs, by people and by agents writing them at runtime. Honesty comes first; strictness is a means to it, and has to pay for itself.
- Prefer few general rules. The compiler knows mechanisms and categories, never particular names; anything it treats specially is declared in the standard library. A rule that exists for one type's name is a defect.
- Claim only what Flow can know. Ask whether it records what Flow did not observe, turns a provider's claim into a fact about the world, or settles an uncertain effect by trying again. Any of the three rules it out.
- Leave applications in charge. Ask who can enforce it. Provider choice, policy, isolation, and lifecycle belong to the application or host that enforces them. A mechanism specific to one provider or deployment belongs in a library or host.
- Change meaning explicitly. Once released, a published program or stored history never means something different tomorrow than it did when written; a change of meaning takes a new version, never a quiet reinterpretation. Flow is pre-release today and makes no compatibility promise yet; the first public release establishes the initial versioned baseline.
Flow code is allowed to get longer in order to stay honest. It is not allowed to get shorter by hiding what runs.
This document owns the why. Foundations states the invariants, the language architecture chooses the mechanisms and draws the application boundary, the specifications fix exact meaning, and the runtime architecture maps that meaning onto an implementation. The documentation index lists everything else.