History
This specification owns the semantic history of one Flow run: which facts it records, how calls and tasks are named, what a pause is, checkpoints, replay and its verdicts, closed evaluation, a run staying on its program, forks, holes, integrity, and program identity. It serves foundations: semantic execution history and the requirements around it.
It does not own the byte encoding of any of this. Fact encoding, the hash chain, content-hashed values and module sources, checkpoint encoding, the encoding of holes, and the data version belong to the runtime's trace format. Running, pausing, resuming and crash handling as host operations belong to execution.
What history records
History records only what the program cannot compute itself. Replay re-runs the program and reads each recorded answer in place of calling a Tool; anything the program can compute is computed again. There is one source of truth, and a record that disagrees with the program shows up as a divergence rather than as a second opinion.
Spawning a task, handing a task to a caller, branches, let bindings, function
results and every other pure step are not recorded. A viewer that shows
intermediate values gets them by re-running the program; those values are a
view, not history.
Serves foundations: history carries non-derivable meaning.
Facts
A history is a sequence of seven kinds of fact.
| Fact | Holds | Written |
|---|---|---|
| Start | program identity, language version, data version, the run's input | before anything else |
| Call | call id, the @tool declaration, the arguments; for a generic Tool, the concrete type |
before the call is sent |
| Reply | call id, and Ok(value) or the ToolProblem the program will see |
when the reply is accepted, before the program sees it |
| Choice | which task task.first picked |
when the choice is made |
| Stop | the stopped task, whether it was found finished (with its value) or running, then how each of its open calls was closed | when the task is stopped |
| Checkpoint | a top-level function, its type arguments for a generic one, and its arguments | when the runtime chooses |
| End | Completed(value) or Faulted(message, location) |
at the end |
- Start names the program by its identity. The language version is the program's std version (modules); the data version covers the value encoding, the Tool protocol and the history format together (trace format).
- Call names the Tool by its declaration's identity (tools). It never names the implementation that served the call: which implementation a host bound is the host's record, not history.
- Reply holds the outcome as the program receives it. The four
ToolProblemcases and what each claims are owned by errors. A@tool typevalue in a reply is recorded as the host's token (tools). - Choice records the outcome of timing. Which task finishes first is not
derivable, so every pick by
task.firstis recorded, including one made inside a library function such astask.race. When these operations run is owned by concurrency. - Stop is recorded whenever a running task is stopped: by
task.stop, bytask.racestopping the losers, or by its owning call returning or faulting (concurrency). It records whether the task was found already finished, with its value, or still running, which is timing and cannot be recomputed. It records the stop first, then closes each open call of the stopped task: a reply that had already arrived is recorded as that call's Reply, and a call with no reply is closed as having none. A reply recorded during a stop reaches no program value. This keeps apart a reply that arrived but was never admitted, a reply admitted and waiting for its continuation, and a task that had already returned; a value the task bound before parking is recomputed. - End is written once. The fault cases are owned by errors. A run the host halts has no End; how a run ends and what each ending leaves behind is owned by execution.
Results nobody reads are dropped when their owning call returns, but their Calls and Replies stay in history. A losing task's calls and replies are recorded like any other; its pure work is recomputed.
Recorded before its consequence
A Call is written before the call is sent, and a Reply is written before the program sees it. Choice and Stop are written before the program continues past them. How durably a runtime writes a fact is owned by the trace format and execution; the order is the language rule.
Serves foundations: recorded before its consequence.
Identity
Ids come from position in the program, never from randomness, clocks, or the host.
- A task is its parent's id plus a counter of the parent's spawns:
root,root/1,root/1/2. - A call is its task plus a counter of that task's Tool calls:
root/1#4. - After a checkpoint counters restart under it. The run's
nth checkpoint prefixes the ids that follow it withcn:, so the first call after the first checkpoint isc1:root#1. A call carried across a checkpoint by a parked task keeps the id it was made with.
Re-running the same program over the same history makes the same ids, so every recorded reply finds its call. Facts are written in the order the run admitted them, and replay admits them in that same order: ids match a fact to the program, and the order says how far each task had got when it was stopped, which history does not otherwise record.
A pause is an open call
A pause is a call with no reply yet. A run waiting on a person, a service or a timer is waiting on a Tool call, and that call's id names what it waits for. Resuming delivers that call's Reply: the process that resumes re-runs the program from the latest checkpoint, serves every recorded fact, records the delivered reply, and continues live. The host operations that pause and resume a run are owned by execution.
Checkpoints
Values are immutable and there is no global state, so at a tail call made by a task's outermost function the whole future of that task is the called function, its type arguments if it is generic, and its arguments. A checkpoint records exactly that.
- Only where it is complete. A checkpoint is taken in the root task, at a tail call to a top-level function. At that moment the stack is empty and nothing waits for a result, so the called function and its arguments are the whole state of the run. A local or anonymous function does not qualify: what it captured is not in its arguments.
- What the arguments may hold. Data,
@tool typevalues (the host's token, already recorded), and tasks parked on their only Tool call. A parked task's work is one Tool call, optionally wrapped in a constructor that may also carry data, such astask.spawn(flow() = Received(feed.next()))orArrived { index, reply: research.search(item) }. It counts whether or not its reply has arrived. It is recorded as its call id, the constructor and the data the constructor carries; on resume it waits on that call again, or is finished if the reply is recorded. Data the constructor would compute after the call is computed when the checkpoint is taken: a task whose data faults, whatever its reply, has none to record and prevents a checkpoint, and a runtime may decline a checkpoint whose data it has not finished computing. - Anything else prevents a checkpoint at that tail call. Any other live task does, whether it is in the arguments or not, such as a background task in the middle of its own loop.
- Derived, and checkable. A checkpoint holds nothing the program did not compute. Replay that re-runs up to a checkpoint compares what it computes with what is recorded, and a difference is a divergence.
- Written when the runtime chooses, for example every so many rounds or when history grows. Where a checkpoint is written does not change what the program means.
- Resume and replay start at the latest checkpoint. Facts before it are no longer needed for either and may be archived.
Restoring a checkpoint's arguments rebuilds the values the program made, opaque values and host tokens included; it is not decoding (tools).
A loop that passes a pending read from round to round checkpoints normally,
because the read is parked on its only call. flow check reports the tail calls
where a checkpoint can be taken (execution).
Serves foundations: crash-resumable execution.
A runtime's own snapshot
A runtime may save its own internal state, such as a stack, to resume faster. That is a private cache, never history: it has no meaning, may be thrown away at any time, and can always be rebuilt from history.
Keeping history small
- Record only what cannot be recomputed.
- Store a large value once, by content hash, and refer to it from facts (trace format).
- Take checkpoints, so earlier facts can be archived.
Retention, archiving and deletion are the host's.
Example
use "std@v1/task"
use "./planner"
use "./web"
use "./docs"
use "./human"
pub type Report { Resolved(Text), GaveUp(List<Text>), Unanswered(Text) }
pub flow main(ticket: Text) -> Report !tool = work(ticket, [], 1)
flow work(ticket: Text, notes: List<Text>, round: Int) -> Report !tool = {
guard round <= 10 else { return GaveUp(notes) }
match planner.next(ticket, notes) {
Ok(Search(query)) => match task.race([
flow() = web.search(query),
flow() = docs.search(query),
]) {
Ok(hit) => work(ticket, notes ++ [hit], round + 1),
Err(_) => work(ticket, notes, round + 1),
},
Ok(Ask(question)) => match human.ask(question) {
Ok(answer) => work(ticket, notes ++ [answer], round + 1),
Err(_) => Unanswered(question),
},
Ok(Done(summary)) => Resolved(summary),
Err(_) => GaveUp(notes),
}
}Open in playground →This notation is illustrative, not the trace format:
1 Start program=support@3f8a std@v1 input={ ticket: "Login fails on Safari" }
2 root#1 Call planner.next("Login fails on Safari", [])
3 root#1 Reply Ok(Search("safari login cookie"))
4 root/1#1 Call web.search("safari login cookie")
5 root/2#1 Call docs.search("safari login cookie")
6 root/2#1 Reply Ok("SameSite=None requires Secure ...")
7 root Choice race picked root/2
8 root/1 Stop closed root/1#1 with no reply
9 Checkpoint work("Login fails on Safari", ["SameSite=None requires Secure ..."], 2)
10 c1:root#1 Call planner.next("Login fails on Safari", ["SameSite=None ..."])
11 c1:root#1 Reply Ok(Ask("Which Safari version does the customer use?"))
12 c1:root#2 Call human.ask("Which Safari version does the customer use?")
← paused; facts 1–8 may be archived
13 c1:root#2 Reply Ok("Safari 17.4") ← delivered hours later; resume from fact 9
14 c1:root#3 Call planner.next("Login fails on Safari", ["SameSite=None ...", "Safari 17.4"])
15 c1:root#3 Reply Ok(Done("Enable Secure on the session cookie; ..."))
16 End Completed(Resolved("Enable Secure on the session cookie; ..."))The race's spawns, the branch taken, notes ++ [hit] and round + 1 appear
nowhere: re-running computes them. Fact 7 is recorded because which search
finished first is not derivable. Fact 9 is possible because the race stopped
its loser, leaving no live task, and work is a top-level function called in
tail position from the root task.
Replay
Replay re-runs the exact program from the start, or from a checkpoint, and reads each recorded fact in place of acting. It holds no binding and calls no Tool implementation, so a complete history replays with nothing installed behind any Tool.
At each step replay checks that what the program does matches the record: each Call names the same Tool with the same arguments, and the same concrete type for a generic Tool; each Choice picks a task the program can pick there; each Checkpoint holds what the program computes; End holds the result the program produces. It then serves the recorded Reply, Choice or Stop outcome, in the recorded order. Replay reports one verdict:
- Matched: the program reached the recorded End and used every fact.
- Incomplete: history ends before the program does, as a paused run's does, or the program needs a fact that is a hole.
- Diverged at fact N: the program does something other than what fact N records, or leaves fact N unused, with the reason.
No verdict falls back to calling a Tool, including when the missing fact is exactly the one the program needs next.
Serves foundations: replay without dispatch and foundations: deterministic reconstruction.
Rules around a run
Closed evaluation
There is no ambient clock, randomness, filesystem, network, environment or
configuration. Each is a Tool call. std computes nothing from a clock: its
clock module is @tool declarations the host implements, so reading the time
or sleeping is a recorded Call and Reply (stdlib). Every other
value that reaches the program arrives as the run's input.
Serves foundations: visible outside influence.
A program never reads its own history
No expression observes the history, its length, a fact, an id, whether it is running live or in replay, or which implementation serves a Tool. Reading them would make the record an input to the program it records. Viewers and hosts read history; another run may receive one as ordinary data through its input or a Tool call.
A crash while a call is out
When a process dies after a Call is written and before its Reply, the host
decides what that call becomes (execution).
By default it records the Reply Unknown. Sending the call again is allowed
only under the same call id, so it remains one call with one Reply; it never
becomes a new call. Flow never re-sends a call on its own.
Serves foundations: honest uncertainty.
Retries inside one call are invisible
One call has one Call and at most one Reply. Attempts a host makes inside one call are in the host's records, not history. A retry the program writes is a new call with its own id.
Program identity
A program's identity is a hash of every module's parsed program, without comments or layout, together with the resolved commit of every imported repository and the std version (modules). Start records it. Module sources are stored in history by content hash, so replay and resume never need the network to find the program.
A run stays on its program
A run always continues on the exact program it started with, read from the sources history stores. A new version of a program is used only by new runs, so a pending call's id never shifts under it.
Forks
A new run may start from any checkpoint of another run. It may run the same program or a newer version of it, and it may replace the checkpoint's arguments with new input. Its Start records the parent run and the checkpoint. The checkpoint's arguments are rebuilt as the program made them; they, or the replacement input, must fit the fork's parameter types, or the fork is refused. Replacement input is decoded like any run input (tools).
A fork is how a long-running program moves to a new version and how a run is
rewound: the parent run is not changed and stays where it is. A host token in a
forked checkpoint may no longer be known to its host, which then answers calls
using it with NotRun (tools).
Nested runs
A Flow run started through a Tool (tools) has its own history. The outer history records one Call and one Reply for each operation on it; the two histories are linked by the host, not merged.
Holes
A host may replace a stored value with a hole that keeps only the value's hash, for example to remove personal data. Inspection still shows everything else. Replay that needs the value reports Incomplete; a hash is not a value and cannot drive re-execution.
A hole carries no reason and no label. Why a value is missing — redaction, expiry, loss — does not change what the record says, and the language defines no category for it. A record that substitutes another value for a removed one and still calls itself complete is not a conforming history.
Serves foundations: honest withholding.
Integrity
Every fact carries the hash of the fact before it. Changing, removing or reordering a fact is detectable against a head hash held somewhere the holder of the record does not control. Nothing more is claimed: a chain whose content and head one holder can both rewrite proves nothing against that holder, and no hash authenticates the writer. The chain's encoding is owned by the trace format.
What history does not claim
History records what an honest compiler and runtime observed at Flow's boundary.
- A Reply is evidence that a reply arrived and was admitted, not that it is true.
- A Call says what Flow asked for, not what the implementation did to answer.
- History says nothing about the implementation behind a Tool, the host that bound it, or the application around the run.
- A runtime that fabricates facts can fabricate history; no record can make its own writer honest.
Applications may link their own records, such as approvals, provider logs or session records, to a run or a call id. Those records sit beside history and add nothing to its meaning.
Serves foundations: bounded trust and disclosure.
Versions
Every history's Start records its data version, which also covers the value encoding and the Tool protocol, and the language version of its program. A reader either reads a known data version or refuses the history whole; stored facts are never reinterpreted under a newer meaning. Flow is pre-release and makes no compatibility promise yet (runtime architecture).
Serves foundations: stable meaning.