Getting started with Flow
Flow is a small, typed language for orchestrating models, tools and people. A program computes deterministically, reaches the outside world only through typed Tool calls, and keeps a history of everything it could not compute itself. That history explains a run, replays it without calling anything, and resumes it in a new process.
How to read this page. It describes the target language that language/spec/ defines. The code is written in the decided syntax to illustrate the rules; each section links the specification that owns them. Flow is pre-release, and runtime/architecture.md is the one document that records where the implementation stands.
This page follows one program, a support desk that drafts a reply to a ticket, asks a person to review it, and sends it. Along the way it shows:
- Tools are functions declared in modules and imported like any other;
- the host binds an implementation to every Tool before the run starts;
- Tool calls return typed results, taken apart with
match,?andguard; - a pause is a Tool call that has no reply yet;
- history records only what the program cannot compute;
- replay reads that history without calling anything, and resume continues from it;
- a checkpoint lets both start partway through.
The Tools
A Tool is a function whose body is outside Flow. It is declared with @tool,
has no body, and returns a Result, because any call to the outside can fail.
These three modules live in one repository, github.com/example/support:
// desk.flow
pub type Ticket { pub id: Text, pub customer: Text, pub body: Text }
pub type Receipt { pub id: Text }
pub type DeskError { NoSuchTicket(Text), Closed(Text) }
@tool pub flow fetch(id: Text) -> Result<Ticket, DeskError>
@tool pub flow history(customer: Text) -> Result<List<Ticket>, DeskError>
@tool pub flow send(ticket: Text, subject: Text, body: Text) -> Result<Receipt, DeskError>
// model.flow
pub type ModelError { pub detail: Text }
@tool pub flow extract<T>(prompt: Text) -> Result<T, ModelError>
// person.flow
pub type Review { Approve, Revise(Text), Decline }
pub type AskError { pub detail: Text }
@tool pub flow review(ticket: Text, draft: Text) -> Result<Review, AskError>Open in playground →A declaration says what Flow sends and what it accepts back, and nothing else.
It names no process, service, credential or permission, and Flow infers nothing
from a Tool's name: person.review is not special because a person answers it.
model.extract is generic, so one declaration serves every structured output;
the caller chooses the type
(tools.md).
The program
The program is one file, reply.flow, in the application's own directory:
use "github.com/example/support@v1/desk"
use "github.com/example/support@v1/model"
use "github.com/example/support@v1/person"
pub type Case { pub ticket: desk.Ticket, pub previous: List<desk.Ticket> }
pub type Draft { pub subject: Text, pub body: Text }
pub type Problem {
Fetching(ToolProblem<desk.DeskError>),
Drafting(ToolProblem<model.ModelError>),
Reviewing(ToolProblem<person.AskError>),
Sending(ToolProblem<desk.DeskError>),
}
pub type Outcome {
Sent(desk.Receipt),
Declined,
OutOfDrafts(Int),
Trouble(Problem),
}
pub flow main(ticket_id: Text, max_drafts: Int) -> Outcome !tool =
match gather(ticket_id) {
Ok(case) => draft(case, [], 1, max_drafts),
Err(problem) => Trouble(problem),
}
flow gather(id: Text) -> Result<Case, Problem> !tool = {
let ticket = desk.fetch(id) ? Fetching
let previous = desk.history(ticket.customer) ? Fetching
Ok(Case { ticket, previous })
}
flow draft(case: Case, notes: List<Text>, round: Int, max_drafts: Int) -> Outcome !tool = {
guard round <= max_drafts else { return OutOfDrafts(max_drafts) }
let prompt = """
Draft a reply to this support ticket.
Ticket: ${case.ticket.body}
Earlier tickets from this customer: ${case.previous}
Reviewer notes so far: ${notes}
"""
let proposal: Draft = match model.extract(prompt) {
Ok(proposal) => proposal,
Err(problem) => return Trouble(Drafting(problem)),
}
match person.review(case.ticket.id, proposal.body) {
Ok(Approve) => send(case, proposal),
Ok(Revise(note)) => draft(case, notes ++ [note], round + 1, max_drafts),
Ok(Decline) => Declined,
Err(problem) => Trouble(Reviewing(problem)),
}
}
flow send(case: Case, proposal: Draft) -> Outcome !tool =
match desk.send(case.ticket.id, proposal.subject, proposal.body) {
Ok(receipt) => Sent(receipt),
Err(problem) => Trouble(Sending(problem)),
}Open in playground →A file is a module, and this one needs nothing beside it: its imports carry all
of its dependency information. A top-level let names a constant,
let max_rounds = 6, written in lower case like every value. Statements end
at the line end, text
interpolates with ${...}, and """ opens a multi-line text block whose
margin is removed (grammar.md). Loops are
recursion: draft calls itself for the next round
(semantics.md).
Tools are imported, and the whole surface is known
use "github.com/example/support@v1/desk"Open in playground →use binds the last segment of the path, and every imported name is written
qualified: desk.fetch, desk.Ticket. A constructor is written bare where
its type is known from context, as in Sent(receipt) above, and otherwise
qualified by its type, desk.DeskError.Closed, never desk.Closed
(modules.md). A remote path always carries the
repository's version after @. The first time a tag is resolved its commit is
pinned, and a tag that later moves refuses the program until the move is
accepted (modules.md).
There is no separate import form for Tools. Importing the module brings its
Tools with it, and the compiler lists every @tool function a program can
reach. Before anything runs, flow check shows that this program can call
exactly desk.fetch, desk.history, desk.send, model.extract and
person.review, and nothing else: there is no ambient clock, file, network,
environment or person. Reading the time is a Tool call too, through std's
clock module (stdlib.md).
A function that calls a Tool, directly or through anything it calls, is marked
!tool, and a function without the mark is checked to make no Tool call
(effects.md). A helper that only formats a prompt
cannot quietly call a model.
The host binds every Tool before the run
The declarations say what the program needs. The host — the CLI, or an application that embeds Flow — decides what answers each call. Before the first fact is recorded it binds every reachable Tool to an implementation, keyed by the declaration's identity: its repository, resolved version, file and name. An implementation fits a declaration when their signature fingerprints match. A Tool with no binding, or one that doesn't fit, refuses the run before anything happens (abi.md).
The repository that declares Tools may say how to run them, in a
flow-tool.toml at its root:
[tools."desk.flow"]
command = "support-tools desk" # a process speaking the Tool protocol
[tools."model.flow"]
mcp = "npx @example/model-mcp" # an MCP server, through the adapter
[tools."person.flow"]
command = "support-tools review-queue"The host may override any of these: a test binds fakes, a deployment routes a
Tool through a proxy, an application registers implementations in code. The
reference CLI asks before it first runs an implementation a repository names,
and --allow gives that consent for non-interactive use
(execution.md).
None of this is part of the program. The same file means the same thing whether
person.review is backed by a review queue, a chat message, or a fake in a
test. Credentials, approvals, sandboxing, retries and budgets belong to the host
and the application around the run.
Typed calls
Calling desk.fetch, declared to return Result<Ticket, DeskError>, gives a
Result<Ticket, ToolProblem<DeskError>>. ToolProblem has four cases, each a
different claim (errors.md):
| Case | Means |
|---|---|
Failed(e) |
the Tool replied with its declared error |
BadReply(reason) |
a reply arrived and didn't match the declared type |
NotRun(reason) |
the call definitely didn't happen |
Unknown(reason) |
the call may or may not have happened |
A reply is checked against the declared type before the program sees it. When a
model answers model.extract with something that isn't a Draft, the program
gets BadReply, a value it handles, not a crash.
An Err is ordinary data: it never leaves a function on its own. The program
takes results apart in three ways.
matchcovers every case.draftmatches on the person'sReviewand on the problem, and the checker refuses amatchthat misses one.?returns early, visibly, where it is written. Ingather,desk.fetch(id) ? Fetchinggives the ticket onOk, and onErr(e)returnsErr(Fetching(e))fromgather. The converter after?is any function, and a constructor is one.?works only inside a function that returnsResult;mainreturns its ownOutcome, so it matches once on whatgatherreturns.guardstates a precondition for the rest of the block. Itselseblock must leave, so afterguard round <= max_draftsthe rest ofdraftcan rely on it (semantics.md).
let proposal: Draft = match model.extract(prompt) { ... } chooses the generic
Tool's type through the result it expects. The call sends the schema of Draft
with the request, and the reply is decoded into a Draft or becomes
BadReply. A decoded value proves only that Flow can interpret it, not that it
is true (tools.md).
Flow never repeats a call on its own. NotRun says a retry can't repeat an
effect; Unknown says it might, so a program that must not send twice asks
another question instead of retrying. examples.md
shows one.
A pause is an open call
person.review may take hours. Nothing in the program says so, and nothing has
to: a pause is a Tool call with no reply yet.
When every live task is waiting on a Tool call, the host may end the process rather than wait in it. The run is then paused. Nothing is written for the pause; history simply ends at the open call, and that call's id names what the run waits for (history.md). The CLI exits with code 3 and prints the run id and the pending call id.
Resuming delivers that call's reply, decoded against its declared type like any other:
flow resume run-7f3a --call 'root#4' --reply '{"Ok": {"Revise": "Mention the 5-day refund window"}}'There is no separate wait, signal or approval form. A person, a timer, a webhook and a slow service are all Tool calls, and they pause the same way (execution.md).
What history records
History records only what the program cannot compute itself: the input,
each Tool call and its reply, the choices timing makes between tasks, and how
the run ended. Branches, let bindings and function results are computed again
whenever they are needed (history.md).
One run of the program, in illustrative notation rather than the trace format:
1 Start program=reply@9c1e std@v1 input={ ticket_id: "T-1042", max_drafts: 3 }
2 root#1 Call desk.fetch("T-1042")
3 root#1 Reply Ok(Ticket { id: "T-1042", customer: "c-77", body: "My refund hasn't arrived ..." })
4 root#2 Call desk.history("c-77")
5 root#2 Reply Ok([])
6 root#3 Call model.extract<Draft>("Draft a reply to this support ticket. ...")
7 root#3 Reply Ok(Draft { subject: "Your refund", body: "..." })
8 root#4 Call person.review("T-1042", "...")
← paused: exit code 3, pending root#4
9 root#4 Reply Ok(Revise("Mention the 5-day refund window")) ← supplied by resume
10 Checkpoint draft(Case { ... }, ["Mention the 5-day refund window"], 2, 3)
11 c1:root#1 Call model.extract<Draft>("... Reviewer notes so far: [\"Mention the 5-day ...\"]")
12 c1:root#1 Reply Ok(Draft { subject: "Your refund", body: "..." })
13 c1:root#2 Call person.review("T-1042", "...")
14 c1:root#2 Reply Ok(Approve)
15 c1:root#3 Call desk.send("T-1042", "Your refund", "...")
16 c1:root#3 Reply Ok(Receipt { id: "m-5512" })
17 End Completed(Sent(Receipt { id: "m-5512" }))- Ids come from position in the program, never from a clock or the host:
root#4is the root task's fourth Tool call. Re-running the same program over the same history makes the same ids, so every reply finds its call. - Each fact is written before its consequence. A Call is recorded before the call is sent, and a Reply before the program sees it (history.md).
- A generic call records its concrete type, here
Draft. Startnames the program by identity: a hash of every module's parsed program with the resolved commit of every import. The module sources are stored by content hash, so a run never needs the network to find its program again.- History says what Flow observed at its boundary, not what is true. A Reply shows that an answer arrived and had the declared type; it says nothing about which implementation produced it.
A history holds every value the run observed, so it can be more sensitive than an application log. Where it is stored, who reads it, and which values the host replaces with hash-only holes are the host's decisions (execution.md).
Checkpoints
Values are immutable and there is no global state. So when the root task makes a tail call to a top-level function, the called function and its arguments are the whole future of the run. A checkpoint records exactly that (history.md).
In this program, main's call to draft, draft's call to itself for the
next round, and draft's call to send are all such tail calls; flow check
lists them. The runtime decides when to write a checkpoint, and where it does
changes nothing about what the program means. Fact 10 above is one. After it,
ids restart under the checkpoint's prefix, c1:, and facts 1–9 are no longer
needed to resume or replay; the host may archive them.
A checkpoint holds nothing the program did not compute, so replay checks it: re-running up to fact 10 must produce exactly that function and those arguments. A loop shaped as tail calls is what lets a long-running program keep its history small; checkpoints.md shows the shapes that checkpoint and the ones that don't.
Replay is not resume
Both read the same history, and both re-run the same program from Start or
from the latest checkpoint, serving each recorded reply instead of calling.
They differ in what happens at the end of the record.
Replay binds nothing. It has no implementation to call, so it cannot call one. At each step it checks that the program does what the record says — the same Tool, the same arguments, the same choices — and reports one verdict (history.md):
- Matched: the program reached the recorded
Endand used every fact. The completed run above replays to Matched with nothing installed behind any Tool. - Incomplete: the history ends before the program does. The paused history,
facts 1–8, replays to Incomplete: the program needs the reply to
root#4, and replay does not ask anyone for it. - Diverged at fact N: the program did something other than fact N records,
for example because someone edited the prompt in
draftand replayed the old history against the new source.
No verdict falls back to calling a Tool. flow test replays saved histories
this way as regression tests.
Resume continues live. It re-forms the program from the sources the
history stores, binds every Tool again, re-executes to the end of the record
serving every recorded reply, admits the reply it was given for the open call,
and then carries on calling Tools for real. In the run above, resuming with the
reviewer's note re-runs gather and the first draft from facts 2–7 without
calling desk or the model again, records fact 9, and continues
(execution.md).
A run always continues on the exact program it started with. A new version of
reply.flow is used only by new runs, or by a fork: a new run started from
a checkpoint of an old one, on the same or a newer program, whose arguments must
fit the new function's parameters
(history.md).
When a process dies while a call is out, the host decides what that call
becomes. By default it is recorded as Unknown, which is exactly what the
program would have to handle anyway; Flow never treats a missing reply as
success and never re-sends a call under a new id
(history.md).
Run it
The reference CLI drives the whole cycle. These commands illustrate the surface execution.md defines:
flow check reply.flow
flow run reply.flow --input '{"ticket_id": "T-1042", "max_drafts": 3}'
flow resume run-7f3a --call 'root#4' --reply '{"Ok": "Approve"}'
flow replay run-7f3a
flow history run-7f3aflow check lists the reachable Tools and the checkpoint sites. flow run
starts the run with input keyed by parameter name; it runs main unless
--entry names another pub flow. The exit code says how it ended: 0
completed, 1 faulted, 2 refused before starting, 3 paused, 4 completed with an
entry result that is an Err, and 130 halted by the host. flow history shows
the record, re-running the program to show intermediate values as a view, and
flow fork starts a new run from a checkpoint of another.
flow test runs @test functions against fakes written in Flow and replays
saved histories (execution.md).
Where the application stays in charge
Flow owns the meaning of this program: what it computes, which Tool calls it
makes, what each reply means, and what the run records. The application around
it owns everything else — which model answers model.extract, how the review
reaches a person and who that is, whether desk.send needs an approval, what a
run may spend, where its history is kept, and what happens to the work after the
run ends (language/architecture.md).
Read next
- examples.md — small programs for the common patterns.
- philosophy.md — why Flow is a language, and the case against adopting it.
- foundations.md — the invariants every conforming implementation satisfies.
- language/spec/ — the exact rules, one surface per document.