Writing programs that reach checkpoints
This guide is non-normative. It shows program shapes; it creates no semantics and states no guarantee. History owns what a checkpoint is and where one may be taken, evaluation owns tail calls, and execution owns resume, forks and exit codes.
Why the shape matters
A checkpoint lets resume and replay start partway through a run instead of at
Start, and lets the facts before it be archived. The runtime may take one
only where the called function and its arguments are the whole state of the
run: in the root task, at a tail call to a top-level function, with every other
live task parked on its only Tool call
(history).
A program that never reaches such a point still runs, pauses and resumes
correctly. Its history simply grows from Start for as long as it runs, and
every resume re-executes all of it. For a short run that is fine. For a loop
that runs for days it is not.
flow check lists the tail calls where a checkpoint can be taken
(execution). Read that
list after writing a long-running program; an empty list is the usual sign that
one of the shapes below is missing.
The loop shape
Write each round as a top-level function that ends in a tail call to the next round's top-level function. Everything the next round needs travels as arguments. Tasks that must outlive the round travel too, parked on their only Tool call.
type Event {
Generated(Generated),
Received(Result<Text, ToolProblem<conversation.Error>>),
}
pub flow main(opening: Text) -> Result<Summary, Problem> !tool =
converse([User(opening)], listen())
flow converse(history: List<Turn>, incoming: task.Task<Event>) -> Result<Summary, Problem> !tool = {
let generation = task.spawn(flow() = Generated(generate(history)))
match task.first([generation, incoming]) {
Generated(generated) => converse(history ++ [Assistant(generated.text)], incoming),
Received(Ok(text)) => {
let _ = task.stop(generation)
converse(history ++ [User(text)], listen())
},
Received(Err(problem)) => Err(Listening(problem)),
}
}
flow listen() -> task.Task<Event> !tool =
task.spawn(flow() = Received(conversation.next_message()))Open in playground →Each converse(...) in tail position is a candidate checkpoint. incoming is
parked on its only call, so it may be held. generation is not: it either won
the race or was stopped, so no other task is live at the tail call.
Three things quietly prevent checkpoints:
- The state is captured, not passed. A local or anonymous function's captured bindings are not in its arguments, so a tail call to one is never a checkpoint. Loop through top-level functions.
- A task is still running. A task in the middle of its own work, or a race loser left alive, blocks the checkpoint whether or not it appears in the arguments.
- The round is not in tail position.
let next = round(state)followed by more work keeps the caller's frame alive.
Restructuring direct-style phases
A program written as a sequence of phases reads naturally, but none of its phase calls is a tail call to a top-level function, so it reaches no checkpoint until it is nearly finished:
pub flow main(goal: Text) -> Result<Report, Problem> !tool = {
let plan = planner.plan(goal) ? Planning
let drafts = draft_all(plan.sections)?
let review = reviewer.review(drafts) ? Reviewing
Ok(Report { plan, drafts, review })
}Open in playground →Turn each phase into a top-level function that ends by tail-calling the next one, passing what the later phases need:
pub flow main(goal: Text) -> Result<Report, Problem> !tool = {
let plan = planner.plan(goal) ? Planning
drafting(plan, [], plan.sections)
}
flow drafting(plan: Plan, drafts: List<Draft>, left: List<Section>) -> Result<Report, Problem> !tool =
match left {
[] => reviewing(plan, drafts),
[section, ..rest] => {
let draft = writer.draft(section) ? Drafting
drafting(plan, drafts ++ [draft], rest)
},
}
flow reviewing(plan: Plan, drafts: List<Draft>) -> Result<Report, Problem> !tool = {
let review = reviewer.review(drafts) ? Reviewing
Ok(Report { plan, drafts, review })
}Open in playground →Now a checkpoint can be taken after planning, after every draft, and before review. The program means the same thing; only where the runtime may record its state has changed.
One at a time, or a bounded batch
drafting above handles one section per round, so checkpoints fall between
items. task.map_limit runs the same work with bounded parallelism
(concurrency):
flow drafting(plan: Plan) -> Result<Report, Problem> !tool = {
let drafts = task.map_limit(plan.sections, 4, flow(section) = writer.draft(section))
reviewing(plan, drafts)
}Open in playground →The batch is one call that is not in tail position, with tasks running inside it, so checkpoints fall only around the whole batch. If the process dies halfway through, resume re-executes the batch from the checkpoint before it, serving every recorded reply; the calls that already replied are not sent again, but the history since that checkpoint is all kept.
Choose by the size of the batch and the cost of re-executing it. A few dozen
quick calls fit comfortably in a batch. Hundreds of slow calls, or calls that
wait on people, are better handled one at a time, or in chunks of map_limit
with a tail call between chunks.
Return Result from the entry
? needs the enclosing function to return Result
(errors). A program that wants ? in its rounds
and also wants a friendly outcome enum is tempted to wrap the loop:
pub flow main(goal: Text) -> Outcome !tool = match drafting(goal, [], 1) {
Ok(report) => Done(report),
Err(problem) => GaveUp(problem),
}Open in playground →That match keeps main's frame waiting on the loop's result, so none of the
loop's tail calls is a checkpoint. Let the entry return the Result itself and
tail-call the loop. A run whose entry returns Err still completes, and the
CLI exits with code 4 so a script can tell it from success
(execution).
Moving a long-lived run onto a new version
A run always continues on the program it started with (history). To move a long-lived run onto a newer program, fork it from its latest checkpoint onto the new version (execution). Habits that keep that possible:
- Keep the loop's top-level function names and modules stable across versions. A fork onto a version without a function of the same name in the same module is refused.
- Keep the loop's state in a few record types and change them deliberately. Every checkpointed argument must fit the new parameter types; where one no longer fits, the forker supplies it as replacement input.
- Handle
UnknownandNotRunon a parked receive by starting a new receive. The fork's open calls are settled as after a crash, and the default isUnknown(history).
The parent run is left where it was; whether to halt it, and when, is the host's decision.