Save and resume a workflow
Use checkpoints to resume a workflow and explicitly retry interrupted tasks.
Save progress
Section titled “Save progress”Pass a checkpoint to the workflow’s start() method when you need to continue in a later process. Outpost saves task transitions and results under the checkpoint’s runId.
It prints { files: 12 } and saves the checkpoint under .outpost/storage. Run it again: scan does not run, its value comes from the checkpoint.
- Task recordsStatus, attempts, errors, gate requests and decisions of every task.
- OutputsThe value of each
donetask, restored instead of running it again. - UsageCumulative attempts and tokens, so a budget spans every resume.
A resumed run keeps its executionId and each task’s context.idempotencyKey. To resume, call start() on the same workflow definition. The checkpoint stores state and outputs; it does not save sandbox instances, their files or task code.
Return JSON outputs
Section titled “Return JSON outputs”A checkpointed task must return undefined or a value that survives JSON serialization without losing information. Otherwise the attempt fails. Convert dates to strings and return only the fields you need.
A dispatch result carries methods such as resume(). Project it in a defineTask, as shown below:
A checkpoint is limited to 16 MiB. Store large payloads as artifacts and return their reference.
Keep the checkpoint identity
Section titled “Keep the checkpoint identity”A checkpoint only resumes the workflow that wrote it. start() rejects a checkpoint whose identity differs.
| Part of the identity | Where you set it |
|---|---|
| Workflow name | defineWorkflow(name, tasks) |
| Version | checkpoint.version |
| Graph | Task keys and their after dependencies |
| Execution settings | timeoutMs, retry settings, whether condition or retry.accepts is set |
| Gates | Kind, prompt, actors and authentication of each approval or pause |
| Loop tasks | maxRounds |
| Interactive tasks | actors, agent, model, brief, repository, maxTurns and sandbox provider |
Other task code, briefs and workflow inputs are not part of it: change version when you change them. A saved run cannot move to another identity, including a new version: start it again under a new runId.
The workflow budget is not part of the identity either.
To reuse results across different runs, use the result cache instead.
Resume incomplete work
Section titled “Resume incomplete work”A run that ended with a failed, cancelled or interrupted task resumes only with resume: "retry-incomplete". This authorizes running those tasks again, with their side effects.
It prints failed, then done. Without resume, the second start() rejects.
| Saved state | Without resume | With resume: "retry-incomplete" |
|---|---|---|
Every task done or skipped | Returns the saved result, runs nothing | Same |
| Paused at a gate or waiting for an answer | Continues with your decisions or answers | Same |
| Paused by a quota | Runs the task again after its reset | Same |
A task failed, cancelled or interrupted | start() rejects | Reruns it and the tasks it skipped |
Rerun tasks start a new series of retry attempts. done tasks never run again. A run stopped by its budget resumes the same way; pass a larger budget, since usage keeps adding up.
Recover a run after a crash
Section titled “Recover a run after a crash”While start() is running, the process owns the checkpoint. A normal return releases that ownership. If the process dies, the ownership record remains and prevents another run from starting under the same runId until you recover it explicitly.
The checkpoint’s key is checkpoints/ followed by the SHA-256 of the runId. Progress stays intact; start the run again with resume: "retry-incomplete".
Resume a run from a queue job
Section titled “Resume a run from a queue job”defineWorkflowJob() runs each job under its runId. A queue returns the existing job for an ID it already has, so a finished job never runs again.
To continue the run, enqueue a new job ID with the same runId and the same input:
A different input changes the checkpoint version and the job fails. To replay failed or interrupted tasks, the handler needs checkpoint: { store, version, resume: "retry-incomplete" }.
Store checkpoints remotely
Section titled “Store checkpoints remotely”The store accepts any Transport. Use an S3 or R2 transport so that workers on several machines share runs: see Where data lives.
Limits
Section titled “Limits”- One
start()at a time perrunId. A second one rejects while the first runs. - Ownership never expires on its own. Clear it with
recoverWorkflowCheckpoint()after stopping the old runner. - Replay repeats side effects that an interrupted task already made. Deduplicate them with
context.idempotencyKey: see Job queues and workers.
API: createWorkflowCheckpointStore · WorkflowCheckpointOptions · recoverWorkflowCheckpoint · WorkflowCheckpoint · defineWorkflowJob.