Browse docs

Start here

examplesGetting started with Flowdocumentation

Design

FoundationsLanguage architecturePhilosophy

Language specification

Program checkingConcurrencyDataEffectsResults, Tool problems, and faultsGrammarHistoryModules and importsLanguage specificationEvaluationStandard libraryToolsTypes

Runtime

Runtime architectureThe host boundaryDiagnosticsRunning a programThe history format

Guides

Writing programs that reach checkpointsImplementing Tools with a toolkitLoops that never returnRecursive delegationSharing types between Tool modulesHarnesses over tool registries

Getting started with Flow

Flow is a small, typed language for orchestrating models, tools and people. A program computes deterministically, reaches the outside world only through typed Tool calls, and keeps a history of everything it could not compute itself. That history explains a run, replays it without calling anything, and resumes it in a new process.

How to read this page. It describes the target language that language/spec/ defines. The code is written in the decided syntax to illustrate the rules; each section links the specification that owns them. Flow is pre-release, and runtime/architecture.md is the one document that records where the implementation stands.

This page follows one program, a support desk that drafts a reply to a ticket, asks a person to review it, and sends it. Along the way it shows:

  1. Tools are functions declared in modules and imported like any other;
  2. the host binds an implementation to every Tool before the run starts;
  3. Tool calls return typed results, taken apart with match, ? and guard;
  4. a pause is a Tool call that has no reply yet;
  5. history records only what the program cannot compute;
  6. replay reads that history without calling anything, and resume continues from it;
  7. a checkpoint lets both start partway through.

The Tools

A Tool is a function whose body is outside Flow. It is declared with @tool, has no body, and returns a Result, because any call to the outside can fail. These three modules live in one repository, github.com/example/support:

// desk.flow
pub type Ticket { pub id: Text, pub customer: Text, pub body: Text }
pub type Receipt { pub id: Text }
pub type DeskError { NoSuchTicket(Text), Closed(Text) }

@tool pub flow fetch(id: Text) -> Result<Ticket, DeskError>
@tool pub flow history(customer: Text) -> Result<List<Ticket>, DeskError>
@tool pub flow send(ticket: Text, subject: Text, body: Text) -> Result<Receipt, DeskError>

// model.flow
pub type ModelError { pub detail: Text }

@tool pub flow extract<T>(prompt: Text) -> Result<T, ModelError>

// person.flow
pub type Review { Approve, Revise(Text), Decline }
pub type AskError { pub detail: Text }

@tool pub flow review(ticket: Text, draft: Text) -> Result<Review, AskError>
Open in playground →

A declaration says what Flow sends and what it accepts back, and nothing else. It names no process, service, credential or permission, and Flow infers nothing from a Tool's name: person.review is not special because a person answers it. model.extract is generic, so one declaration serves every structured output; the caller chooses the type (tools.md).

The program

The program is one file, reply.flow, in the application's own directory:

use "github.com/example/support@v1/desk"
use "github.com/example/support@v1/model"
use "github.com/example/support@v1/person"

pub type Case { pub ticket: desk.Ticket, pub previous: List<desk.Ticket> }
pub type Draft { pub subject: Text, pub body: Text }

pub type Problem {
    Fetching(ToolProblem<desk.DeskError>),
    Drafting(ToolProblem<model.ModelError>),
    Reviewing(ToolProblem<person.AskError>),
    Sending(ToolProblem<desk.DeskError>),
}

pub type Outcome {
    Sent(desk.Receipt),
    Declined,
    OutOfDrafts(Int),
    Trouble(Problem),
}

pub flow main(ticket_id: Text, max_drafts: Int) -> Outcome !tool =
    match gather(ticket_id) {
        Ok(case) => draft(case, [], 1, max_drafts),
        Err(problem) => Trouble(problem),
    }

flow gather(id: Text) -> Result<Case, Problem> !tool = {
    let ticket = desk.fetch(id) ? Fetching
    let previous = desk.history(ticket.customer) ? Fetching
    Ok(Case { ticket, previous })
}

flow draft(case: Case, notes: List<Text>, round: Int, max_drafts: Int) -> Outcome !tool = {
    guard round <= max_drafts else { return OutOfDrafts(max_drafts) }
    let prompt = """
        Draft a reply to this support ticket.
        Ticket: ${case.ticket.body}
        Earlier tickets from this customer: ${case.previous}
        Reviewer notes so far: ${notes}
        """
    let proposal: Draft = match model.extract(prompt) {
        Ok(proposal) => proposal,
        Err(problem) => return Trouble(Drafting(problem)),
    }
    match person.review(case.ticket.id, proposal.body) {
        Ok(Approve) => send(case, proposal),
        Ok(Revise(note)) => draft(case, notes ++ [note], round + 1, max_drafts),
        Ok(Decline) => Declined,
        Err(problem) => Trouble(Reviewing(problem)),
    }
}

flow send(case: Case, proposal: Draft) -> Outcome !tool =
    match desk.send(case.ticket.id, proposal.subject, proposal.body) {
        Ok(receipt) => Sent(receipt),
        Err(problem) => Trouble(Sending(problem)),
    }
Open in playground →

A file is a module, and this one needs nothing beside it: its imports carry all of its dependency information. A top-level let names a constant, let max_rounds = 6, written in lower case like every value. Statements end at the line end, text interpolates with ${...}, and """ opens a multi-line text block whose margin is removed (grammar.md). Loops are recursion: draft calls itself for the next round (semantics.md).

Tools are imported, and the whole surface is known

use "github.com/example/support@v1/desk"
Open in playground →

use binds the last segment of the path, and every imported name is written qualified: desk.fetch, desk.Ticket. A constructor is written bare where its type is known from context, as in Sent(receipt) above, and otherwise qualified by its type, desk.DeskError.Closed, never desk.Closed (modules.md). A remote path always carries the repository's version after @. The first time a tag is resolved its commit is pinned, and a tag that later moves refuses the program until the move is accepted (modules.md).

There is no separate import form for Tools. Importing the module brings its Tools with it, and the compiler lists every @tool function a program can reach. Before anything runs, flow check shows that this program can call exactly desk.fetch, desk.history, desk.send, model.extract and person.review, and nothing else: there is no ambient clock, file, network, environment or person. Reading the time is a Tool call too, through std's clock module (stdlib.md).

A function that calls a Tool, directly or through anything it calls, is marked !tool, and a function without the mark is checked to make no Tool call (effects.md). A helper that only formats a prompt cannot quietly call a model.

The host binds every Tool before the run

The declarations say what the program needs. The host — the CLI, or an application that embeds Flow — decides what answers each call. Before the first fact is recorded it binds every reachable Tool to an implementation, keyed by the declaration's identity: its repository, resolved version, file and name. An implementation fits a declaration when their signature fingerprints match. A Tool with no binding, or one that doesn't fit, refuses the run before anything happens (abi.md).

The repository that declares Tools may say how to run them, in a flow-tool.toml at its root:

[tools."desk.flow"]
command = "support-tools desk"        # a process speaking the Tool protocol

[tools."model.flow"]
mcp = "npx @example/model-mcp"        # an MCP server, through the adapter

[tools."person.flow"]
command = "support-tools review-queue"

The host may override any of these: a test binds fakes, a deployment routes a Tool through a proxy, an application registers implementations in code. The reference CLI asks before it first runs an implementation a repository names, and --allow gives that consent for non-interactive use (execution.md).

None of this is part of the program. The same file means the same thing whether person.review is backed by a review queue, a chat message, or a fake in a test. Credentials, approvals, sandboxing, retries and budgets belong to the host and the application around the run.

Typed calls

Calling desk.fetch, declared to return Result<Ticket, DeskError>, gives a Result<Ticket, ToolProblem<DeskError>>. ToolProblem has four cases, each a different claim (errors.md):

Case Means
Failed(e) the Tool replied with its declared error
BadReply(reason) a reply arrived and didn't match the declared type
NotRun(reason) the call definitely didn't happen
Unknown(reason) the call may or may not have happened

A reply is checked against the declared type before the program sees it. When a model answers model.extract with something that isn't a Draft, the program gets BadReply, a value it handles, not a crash.

An Err is ordinary data: it never leaves a function on its own. The program takes results apart in three ways.

  • match covers every case. draft matches on the person's Review and on the problem, and the checker refuses a match that misses one.
  • ? returns early, visibly, where it is written. In gather, desk.fetch(id) ? Fetching gives the ticket on Ok, and on Err(e) returns Err(Fetching(e)) from gather. The converter after ? is any function, and a constructor is one. ? works only inside a function that returns Result; main returns its own Outcome, so it matches once on what gather returns.
  • guard states a precondition for the rest of the block. Its else block must leave, so after guard round <= max_drafts the rest of draft can rely on it (semantics.md).

let proposal: Draft = match model.extract(prompt) { ... } chooses the generic Tool's type through the result it expects. The call sends the schema of Draft with the request, and the reply is decoded into a Draft or becomes BadReply. A decoded value proves only that Flow can interpret it, not that it is true (tools.md).

Flow never repeats a call on its own. NotRun says a retry can't repeat an effect; Unknown says it might, so a program that must not send twice asks another question instead of retrying. examples.md shows one.

A pause is an open call

person.review may take hours. Nothing in the program says so, and nothing has to: a pause is a Tool call with no reply yet.

When every live task is waiting on a Tool call, the host may end the process rather than wait in it. The run is then paused. Nothing is written for the pause; history simply ends at the open call, and that call's id names what the run waits for (history.md). The CLI exits with code 3 and prints the run id and the pending call id.

Resuming delivers that call's reply, decoded against its declared type like any other:

flow resume run-7f3a --call 'root#4' --reply '{"Ok": {"Revise": "Mention the 5-day refund window"}}'

There is no separate wait, signal or approval form. A person, a timer, a webhook and a slow service are all Tool calls, and they pause the same way (execution.md).

What history records

History records only what the program cannot compute itself: the input, each Tool call and its reply, the choices timing makes between tasks, and how the run ended. Branches, let bindings and function results are computed again whenever they are needed (history.md).

One run of the program, in illustrative notation rather than the trace format:

 1             Start       program=reply@9c1e  std@v1  input={ ticket_id: "T-1042", max_drafts: 3 }
 2  root#1     Call        desk.fetch("T-1042")
 3  root#1     Reply       Ok(Ticket { id: "T-1042", customer: "c-77", body: "My refund hasn't arrived ..." })
 4  root#2     Call        desk.history("c-77")
 5  root#2     Reply       Ok([])
 6  root#3     Call        model.extract<Draft>("Draft a reply to this support ticket. ...")
 7  root#3     Reply       Ok(Draft { subject: "Your refund", body: "..." })
 8  root#4     Call        person.review("T-1042", "...")
                           ← paused: exit code 3, pending root#4
 9  root#4     Reply       Ok(Revise("Mention the 5-day refund window"))   ← supplied by resume
10             Checkpoint  draft(Case { ... }, ["Mention the 5-day refund window"], 2, 3)
11  c1:root#1  Call        model.extract<Draft>("... Reviewer notes so far: [\"Mention the 5-day ...\"]")
12  c1:root#1  Reply       Ok(Draft { subject: "Your refund", body: "..." })
13  c1:root#2  Call        person.review("T-1042", "...")
14  c1:root#2  Reply       Ok(Approve)
15  c1:root#3  Call        desk.send("T-1042", "Your refund", "...")
16  c1:root#3  Reply       Ok(Receipt { id: "m-5512" })
17             End         Completed(Sent(Receipt { id: "m-5512" }))
  • Ids come from position in the program, never from a clock or the host: root#4 is the root task's fourth Tool call. Re-running the same program over the same history makes the same ids, so every reply finds its call.
  • Each fact is written before its consequence. A Call is recorded before the call is sent, and a Reply before the program sees it (history.md).
  • A generic call records its concrete type, here Draft.
  • Start names the program by identity: a hash of every module's parsed program with the resolved commit of every import. The module sources are stored by content hash, so a run never needs the network to find its program again.
  • History says what Flow observed at its boundary, not what is true. A Reply shows that an answer arrived and had the declared type; it says nothing about which implementation produced it.

A history holds every value the run observed, so it can be more sensitive than an application log. Where it is stored, who reads it, and which values the host replaces with hash-only holes are the host's decisions (execution.md).

Checkpoints

Values are immutable and there is no global state. So when the root task makes a tail call to a top-level function, the called function and its arguments are the whole future of the run. A checkpoint records exactly that (history.md).

In this program, main's call to draft, draft's call to itself for the next round, and draft's call to send are all such tail calls; flow check lists them. The runtime decides when to write a checkpoint, and where it does changes nothing about what the program means. Fact 10 above is one. After it, ids restart under the checkpoint's prefix, c1:, and facts 1–9 are no longer needed to resume or replay; the host may archive them.

A checkpoint holds nothing the program did not compute, so replay checks it: re-running up to fact 10 must produce exactly that function and those arguments. A loop shaped as tail calls is what lets a long-running program keep its history small; checkpoints.md shows the shapes that checkpoint and the ones that don't.

Replay is not resume

Both read the same history, and both re-run the same program from Start or from the latest checkpoint, serving each recorded reply instead of calling. They differ in what happens at the end of the record.

Replay binds nothing. It has no implementation to call, so it cannot call one. At each step it checks that the program does what the record says — the same Tool, the same arguments, the same choices — and reports one verdict (history.md):

  • Matched: the program reached the recorded End and used every fact. The completed run above replays to Matched with nothing installed behind any Tool.
  • Incomplete: the history ends before the program does. The paused history, facts 1–8, replays to Incomplete: the program needs the reply to root#4, and replay does not ask anyone for it.
  • Diverged at fact N: the program did something other than fact N records, for example because someone edited the prompt in draft and replayed the old history against the new source.

No verdict falls back to calling a Tool. flow test replays saved histories this way as regression tests.

Resume continues live. It re-forms the program from the sources the history stores, binds every Tool again, re-executes to the end of the record serving every recorded reply, admits the reply it was given for the open call, and then carries on calling Tools for real. In the run above, resuming with the reviewer's note re-runs gather and the first draft from facts 2–7 without calling desk or the model again, records fact 9, and continues (execution.md).

A run always continues on the exact program it started with. A new version of reply.flow is used only by new runs, or by a fork: a new run started from a checkpoint of an old one, on the same or a newer program, whose arguments must fit the new function's parameters (history.md).

When a process dies while a call is out, the host decides what that call becomes. By default it is recorded as Unknown, which is exactly what the program would have to handle anyway; Flow never treats a missing reply as success and never re-sends a call under a new id (history.md).

Run it

The reference CLI drives the whole cycle. These commands illustrate the surface execution.md defines:

flow check reply.flow
flow run reply.flow --input '{"ticket_id": "T-1042", "max_drafts": 3}'
flow resume run-7f3a --call 'root#4' --reply '{"Ok": "Approve"}'
flow replay run-7f3a
flow history run-7f3a

flow check lists the reachable Tools and the checkpoint sites. flow run starts the run with input keyed by parameter name; it runs main unless --entry names another pub flow. The exit code says how it ended: 0 completed, 1 faulted, 2 refused before starting, 3 paused, 4 completed with an entry result that is an Err, and 130 halted by the host. flow history shows the record, re-running the program to show intermediate values as a view, and flow fork starts a new run from a checkpoint of another. flow test runs @test functions against fakes written in Flow and replays saved histories (execution.md).

Where the application stays in charge

Flow owns the meaning of this program: what it computes, which Tool calls it makes, what each reply means, and what the run records. The application around it owns everything else — which model answers model.extract, how the review reaches a person and who that is, whether desk.send needs an approval, what a run may spend, where its history is kept, and what happens to the work after the run ends (language/architecture.md).

  • examples.md — small programs for the common patterns.
  • philosophy.md — why Flow is a language, and the case against adopting it.
  • foundations.md — the invariants every conforming implementation satisfies.
  • language/spec/ — the exact rules, one surface per document.