Running a program
This specification fixes how a runtime runs a Flow program: choosing the entry, taking input and giving output, how a run ends and what the CLI reports, pausing and resuming, recovering from a crash, replay, forks, nested runs, limits, and the CLI and MCP surface.
It realizes the language contract and never redefines it. semantics.md owns entry functions and evaluation; errors.md owns faults and how a run ends at language level; history.md owns what history records, pauses, checkpoints, replay verdicts and forks; tools.md owns the language view of a Flow run as a Tool; modules.md owns resolution and pins.
Within the runtime track, abi.md owns the value encoding, the Tool protocol and binding; trace.md owns the history format; diagnostics.md owns the diagnostic shape.
Entry
Any top-level pub flow whose parameters and result are data types can start a
run (semantics.md). The language names no
default entry. The CLI runs main unless --entry names another; an embedding
host chooses its own.
pub flow main(ticket: Text, max_rounds: Int) -> Report !tool = ...Open in playground →flow run support.flow --input '{"ticket": "Login fails", "max_rounds": 10}'
flow run support.flow --entry triage --input '{"ticket": "Login fails"}'- Input is a JSON object keyed by parameter name, decoded with the decoding rules of abi.md. A field the entry does not declare refuses the run, as does any other decoding failure.
- Output is the entry's result in canonical form.
- Before the first fact, the runtime checks the program, binds every Tool it can reach (abi.md), and decodes the input. A failure at any of these refuses the run, and nothing is recorded.
Serves foundations: Visible outside influence.
How a run ends
| State | History | CLI exit code |
|---|---|---|
| Completed | ends in End Completed(value) |
0 |
| Faulted | ends in End Faulted(message, location) |
1 |
| Refused: a compile error, a missing or unfitting binding, bad input | none | 2 |
| Paused: every live task waits on a Tool call and the process exits | ends with its open calls | 3 |
Completed, and the entry's result is an Err |
ends in End Completed(Err(…)) |
4 |
| Halted by the host | no End |
130 |
| Cut short by a storage failure | ends at its last accepted fact, with no End and no host record |
1 |
- Refused is not a run. The CLI prints the diagnostics (diagnostics.md).
- Paused prints the run id and the id of every pending call.
- Exit code 4 lets a script tell an entry that returned its own
Errfrom one that succeeded; to the runtime both completed. An entry often returnsResultso it can use?and still reach checkpoints. - Cut short by a storage failure is not a Flow outcome: the store could not accept a fact, so the run went no further. The CLI prints the storage failure (diagnostics.md), and the run resumes like a crashed one.
- Halted is the host ending a run, for example on an interrupt. Halting is
not a Flow outcome: no
Endis recorded and the program observes nothing. The host records that it halted the run; history does not. A halted run can be resumed like a crashed one (Open calls after a crash). "Stopped" names only whattask.stopreturns (concurrency.md), never a run.
A fault is never caught and never becomes a value, except across a nested run (below).
Serves foundations: Honest uncertainty.
Live execution
A live run writes its history as it goes (trace.md), each fact
accepted before its consequence
(history.md).
An admitted reply is decoded against the call's declared types, and one that
does not decode, or exceeds the reply size limit (Limits), is admitted as
BadReply.
Live history is the progress feed. A host may stream facts to a viewer as
they are accepted. There is no other progress or logging mechanism; a program
that wants to report more makes an ordinary Tool call, such as
ui.status("Searching docs"), recorded like any other.
Serves foundations: Recorded before its consequence; Visible outside influence.
Pause and resume
A run pauses when every live task is waiting on a Tool call and the host ends the process rather than wait in it, for example while a person answers a question. Nothing is written for the pause: the history ends at its open calls, and each is named by its call id (history.md).
Resume continues a run in a new process from its history alone:
- read the history and check it (trace.md);
- re-form the program from the module sources it stores;
- bind every Tool the program can reach; resume may bind different implementations, each fitting its declaration (abi.md);
- start at the latest checkpoint, or at
Startif there is none, and re-execute, serving every recorded reply instead of calling; - admit any reply supplied with the resume for the call it names, recording it
as that call's
Reply; - settle each other open call (below), then continue live where the record ends.
A supplied reply is an outcome in the form abi.md gives for a
reply message, decoded against the call's declared types. A reply for a call
that is not open is refused.
flow resume run-7f3a --call 'c1:root#2' --reply '{"Ok": "Safari 17.4"}'Resume does not rebuild state any other way. A runtime's own snapshot of a run is a private cache (history.md), and a resume that uses it behaves exactly as one that does not.
Serves foundations: Replay and resume read one history; Crash-resumable execution.
Open calls after a crash
A process can die, or be halted, while calls are out. A call with a Call and
no Reply is open. On resume the host settles each open call, call by call:
- A call it left pending at a pause stays open; its reply arrives through its implementation or through a resume that names it.
- A call that was out when the process died is decided by the host, as
history.md
states: by default a
ReplyofUnknown, or the same call sent again (abi.md gives the protocol).
The runtime never decides this itself and never treats a missing reply as
success, failure or NotRun. Retries are
history.md's.
A torn last fact is truncation and is removed before anything is appended (trace.md).
Serves foundations: Honest uncertainty; Effects cannot always be repeated.
Replay
Replay re-executes a run against its history with no implementation bound and
nothing to dispatch to. It reads the program from the stored sources, starts at
Start or at a checkpoint, serves every recorded reply, and checks each call
the program makes against the record. It reports the verdict
history.md defines: matched, incomplete, or
diverged at a fact, with the reason.
A storage failure found while reading is reported as a storage failure, never as a verdict (trace.md).
flow test replays saved histories as regression tests and fails on any verdict
other than matched.
Serves foundations: Replay without dispatch.
Forks
A fork is a new run started from a checkpoint of another
(history.md). Its Start records the parent
run and the checkpoint (trace.md); the parent's history is never
changed.
- On the same program, the checkpoint's function and arguments are rebuilt as they are on resume.
- With changed arguments, each argument the forker supplies is decoded like run input; the others are rebuilt from the checkpoint.
- On a newer version of the program, the new program must have a top-level function of the same name in the same module, and every argument must fit its parameter types. Otherwise the fork is refused, with exit code 2 at the CLI.
The fork's parked tasks keep their calls. Each open call is settled as after a crash, under the fork's own run id.
Serves foundations: Crash-resumable execution; Stable meaning.
Nested runs
The host provides checking, running and resuming other Flow programs as a Tool,
declared with @tool like any other
(tools.md). Its
declarations are std's runner module
(stdlib.md), which the host
implements as it implements clock:
@tool pub type Run
pub type Pending { pub id: Text, pub tool: Text, pub arguments: Text }
pub type RunState { Completed(Text), Faulted(Text), Paused { run: Run, pending: List<Pending> } }
pub type RunError { pub diagnostics: List<Diagnostic> }
@tool pub flow check(source: Text) -> Result<Unit, List<Diagnostic>>
@tool pub flow run(source: Text, input: Text) -> Result<RunState, RunError>
@tool pub flow resume(run: Run, call: Text, reply: Text) -> Result<RunState, RunError>Open in playground →check returns a source's diagnostics or that it checks; run starts a run;
and resume resumes a paused one with a reply for a named call. Diagnostic
is the shape diagnostics.md gives, declared field for field
as a record, with a location's every field optional. The host realizes them as
follows.
- A nested run is a run. It has its own run id and its own history, whose
Startnames the outer run and the call that started it (trace.md). The outer history holds only the outer run's calls and replies. The reference host keeps a nested run's directory beside the outer run's, under the same root. - The source is submitted.
sourceis one submitted file (modules.md), namedmain.flow, and a run starts at itsmain. - Values cross as canonical text. The outer program cannot know the inner
program's types, so input, output and supplied replies are canonical JSON in
Text.inputis a JSON object keyed by parameter name, as run input is (Entry); a completed run's output is its entry's result; a supplied reply is an outcome as areplymessage carries one (Pause and resume). A pending call names its declaration asflow checklists one, and its arguments as a JSON object keyed by parameter name. - Bindings are the outer run's, unless the host overrides them. An override that names nothing the nested run can reach binds nothing there, and does not refuse it.
- Endings become values. A nested run that completes, faults or pauses
comes back as
Completed,FaultedorPaused, a paused run as a host value whose token is its run id. A nested fault never faults the outer run. A source that does not check comes back fromcheckas its diagnostics. A refused run is an error ofrun, and a refused resume an error ofresume, each with its diagnostics. A nested run cut short by a storage failure comes back asUnknown. - Resume names the call it answers by its history id.
- Stopping the call stops the run. When the outer run closes a call that is driving a nested run, the host halts the nested run and answers nothing.
- Halting the outer run halts the nested run. When the host halts a run, it halts every nested run the run is driving, and theirs in turn. Each records its own halt and resumes like a crashed run; the outer call stays open, and is settled on resume as any call out at a halt is.
Replaying the outer run serves the recorded replies of these calls; it does not run the nested program again.
Limits
The specification sets minimums that every implementation supports. A host may
allow more. A program within the minimums runs on every conforming
implementation; a program's own value exceeding an implementation's actual
limit is a fault (errors.md). A Tool reply
larger than the reply limit is not a fault: it is admitted as BadReply, which
the program can handle.
| Limit | Minimum |
|---|---|
Int size, and a Decimal's coefficient: the integer its significant digits make |
4,096 bits |
One Text or Bytes |
16 MiB |
| One Tool reply | 16 MiB |
| Nesting depth of a value | 256 |
| Non-tail call depth | 10,000 |
Tail calls do not count toward call depth. A Decimal's exponent, the power of
ten of its last significant digit, may be any 64-bit signed integer.
Testing
flow test runs two kinds of test.
@test functions. A test is a top-level function marked @test that
returns Result<Unit, Text>; an Err fails it. Its Tool calls go to fakes
written in Flow and marked @fake with the declaration they stand in for:
@fake(openai.chat) flow chat(prompt: Text) -> Result<Text, ToolProblem<ChatError>> = Ok("hi")
@test flow answers_greeting() -> Result<Unit, Text> !tool = {
guard greet("Ada") == "hi" else { return Err("unexpected greeting") }
Ok(())
}Open in playground →A fake returns Result<T, ToolProblem<E>> for a Tool declared with
Result<T, E>, so a test can produce NotRun, Unknown and BadReply as well
as Failed. A fake may also make test values of a @tool type, which no other
Flow code can build: it calls the type on the token the value carries, so
Page("tab-1") in a fake is a Page whose token is "tab-1"
(abi). A fake is a binding written in Flow: it is matched to the
declaration by the same identity a host binding uses
(abi), and like a run, a test that can reach a declaration no fake
stands in for does not start.
A fake may take one more parameter, an Int before its Tool's own: the call's
number among the test's calls of that Tool, counting from 1 in the order the
calls were made. A fake can so answer a retry differently from the first try:
@fake(payments.charge) flow charge(call: Int, amount: Int) -> Result<Text, ToolProblem<ChargeError>> =
if call == 1 { Err(Unknown("connection reset")) } else { Ok("charged ${amount}") }Open in playground →Time in a test. A test runs on a virtual clock that starts at
1970-01-01T00:00:00Z. Std's clock.now and clock.sleep
(stdlib) need no fake: unless
the program has a fake for one, matched by identity like any other, flow test
answers it from the virtual clock — now with the virtual instant, and
sleep(d) once d has passed, at once for a negative d. A fake is evaluated
beside the test from the moment its call is made, and what it returns is the
call's reply, due at the virtual instant it returns, so a fake that sleeps
before returning replies that much later. The clock moves only when every task
of the test and of its fakes waits: the reply due earliest is then handed over,
replies due at one instant in the order their calls were made, and the clock
moves to the instant it was due. task.first picks the candidate that finished
first in virtual time. A call that is closed stops the fake answering it. A
test takes no real time for the time it waits, and runs the same way every
time.
A test is bounded. A test that never ends would never fail, so a host
bounds the evaluation a test takes, counted across the test and its fakes, and
a test that reaches the bound fails, saying so. The bound is the host's: the
reference CLI allows 2^26 evaluation steps, and flow test --steps sets
another.
Recorded histories. A saved history is a regression test: flow test
replays it against the current program and fails on any verdict other than
Matched (history). The current program is
the root file the history's entry module names, found in the directory that
holds the history or the nearest directory above it that has one.
Given a directory, flow test searches it as flow fmt does (below) for root
files and for directories holding a history.
flow test exits 0 when every test passes and every history matches, 1 when
any fails, and 2 when a program does not check, a history or a directory cannot
be read, or its current program cannot be found.
The CLI and MCP surface
The reference CLI offers:
| Command | Does |
|---|---|
flow check |
checks a program, lists every Tool it can reach and where checkpoints can be taken, and each function of its own tree, local and anonymous ones included, with its effect and its type parameters' requirements |
flow run |
starts a run |
flow resume |
resumes a run, optionally with a reply for a named call |
flow replay |
replays a history and reports the verdict |
flow fork |
starts a run from a checkpoint of another |
flow history |
shows a history, re-running the program to show intermediate values, and points out tasks that stay alive across many checkpoints or rounds |
flow test |
runs @test functions and replays saved histories |
flow fmt |
formats source with the one canonical formatter |
flow doc |
generates documentation from /// comments |
flow tools |
lists every @tool declaration of a program, with what a toolkit generates from |
flow update |
accepts moved repository tags into the pin cache |
flow explain |
prints a diagnostic code's explanation, usual repair and example |
flow check lists a place a checkpoint can be taken by what the arguments'
types can hold there, since a function in the arguments prevents a checkpoint
(history.md). A place where an
argument's type always holds a function, such as a function type or a record
with a function field, is not listed. A place where an argument may hold one,
such as a List or an Option of functions, is listed and marked so: whether
a checkpoint is taken there is decided as the run goes. A type parameter holds
what the types the calls reaching the place give it can hold.
flow history re-runs the program as replay does, serving every recorded reply
and calling nothing. Its rounds are the places a checkpoint can be taken
(history.md), each shown with the
function called and its arguments, and it points out every task that two rounds
or more carry.
flow fmt lays source out as grammar.md
says. It rewrites each file it is given in place, searches each directory it is
given for .flow files, and with no path formats standard input to standard
output. A search skips entries whose names start with . and follows no
symbolic link, so it stays inside the directory it was given and ends. With
--check it writes nothing and names each file that is not in canonical form.
It exits 0 when it is done, 1 when --check found a file that is not
canonical, and 2 when a file or a directory it searches cannot be read, a file
does not parse, or a file cannot be written; the other files are still
formatted.
flow doc documents every module a program forms, std's and its
repositories' included, each named by its path at the version it resolved to.
It documents a function with its effect, its type parameters' requirements,
and the local and anonymous functions written in it with the effect inferred
for each, and a type together with the doc comments of its fields and
constructors, and of their named payloads' fields, in the order they are
written. With --out it writes one page per module, named by the module's path.
When modules of different roots have the same path, none of their pages is
written and it exits 2; the other pages are still written.
flow tools lists every @tool function of every module a program forms,
std's included, by identity and fingerprint. Its JSON document is the one input
a toolkit generates from, so no toolkit parses Flow:
{ "data": the data version,
"programs": [{ "root", "diagnostics",
"tools": [{ "tool", "fingerprint", "signature",
"generic", "params": [{"name", "type"}], "ok", "err" }],
"types": { key: declaration } }] }tool, fingerprint and signature are what describe lists for the
declaration (abi). Each type, ok and
err is a type term: a built-in type by name ("Int", "Decimal",
"Bool", "Text", "Bytes", "Duration", "Instant", "Unit", "Never",
"json.Value"); {"List": t}, {"Set": t}, {"Option": t}, {"Dict": [k, v]} or {"Tuple": [t, ...]}; a type parameter as {"Param": "T"}; and a
declared type as {"Named": key}. types describes each declared type the
terms reach, once per list of type arguments, keyed as a schema's $defs is
(abi): its name, identity, args and doc, and its kind — a
record with its fields, an enum with its constructors, each with a
payload that is null, {"positional": [t, ...]} or {"named": fields}, or
a host type. A field is {"name", "type", "doc"}, and every doc is the doc
comment or null. It exits 0 when every program checks and 2 otherwise.
The MCP server offers check, run, resume and replay as tools, with the
same meaning. It answers requests concurrently, so a long run holds up no other
request. A client's cancellation of a call halts the run it drives as an
interrupt does (How a run ends), and the call is answered by nothing. A run's
progress is its history (Live execution).
Every command's result is structured data, and its human-readable text is a rendering of that data. A viewer's intermediate values come from re-running, not from history, and are views.
Running a repository's implementations
Whether a host runs an implementation that an imported repository's
flow-tool.toml names is host policy. The reference hosts decide this way:
- The CLI asks the first time it would run an implementation from a given
repository and commit, before the run starts.
--allowgives that consent for non-interactive use, such as CI. A declined request refuses the run. - The MCP server never asks and never runs an implementation that has not been approved; the run is refused.
- A nested run never asks either (Nested runs): it runs an
implementation that
--allowor an earlier approval covers, and is refused otherwise.
This is the reference hosts' policy, not a language rule. Another host may sandbox, allowlist or forbid implementations as it sees fit.
Serves foundations: Applications remain applications.
What the host keeps
A history holds every value the run observed: inputs, arguments and replies. It may be more sensitive than an application log. Where it is stored, who can read it, its encryption, retention and deletion, and which values become holes belong to the host (trace.md). So do the host's own records: bindings used, retries, halts, approvals, and decisions about open calls.
Serves foundations: Bounded trust and disclosure; Applications remain applications.