Skip to content
Français

Save and resume a workflow

Use checkpoints to resume a workflow and explicitly retry interrupted tasks.

Pass a checkpoint to the workflow’s start() method when you need to continue in a later process. Outpost saves task transitions and results under the checkpoint’s runId.

import { reportValue } from "./reporter.ts";
import {
  createLocalTransport,
  defineTask,
  defineWorkflow,
  createWorkflowCheckpointStore,
} from "@elie-laloum/outpost";

const store = createWorkflowCheckpointStore({
  transporter: createLocalTransport({ directory: ".outpost/storage" }),
});
const scan = defineTask({ key: "scan", perform: () => ({ files: 12 }) });
const result = await defineWorkflow("scan", [scan]).start({
  checkpoint: { store, runId: "scan-2026-09", version: "1" },
});
result.unwrap();
reportValue(result.value(scan));
// Example output: { files: 12 }

It prints { files: 12 } and saves the checkpoint under .outpost/storage. Run it again: scan does not run, its value comes from the checkpoint.

  • Task recordsStatus, attempts, errors, gate requests and decisions of every task.
  • OutputsThe value of each done task, restored instead of running it again.
  • UsageCumulative attempts and tokens, so a budget spans every resume.

A resumed run keeps its executionId and each task’s context.idempotencyKey. To resume, call start() on the same workflow definition. The checkpoint stores state and outputs; it does not save sandbox instances, their files or task code.

A checkpointed task must return undefined or a value that survives JSON serialization without losing information. Otherwise the attempt fails. Convert dates to strings and return only the fields you need.

A dispatch result carries methods such as resume(). Project it in a defineTask, as shown below:

import { defineIsolatedTask, defineTask } from "@elie-laloum/outpost";
import { coder, repository, sandboxProvider } from "./outpost.config.ts";

const agent = defineIsolatedTask({
  key: "fix-agent",
  request: () => ({
    repository,
    sandboxProvider,
    agent: coder,
    brief: { text: "Fix the failing tests and commit the fix." },
  }),
});
export const fix = defineTask({
  key: "fix",
  perform: async (context) => {
    const { branch, commits } = await agent.perform(context);
    return { branch, commits, finishedAt: new Date().toISOString() };
  },
});

A checkpoint is limited to 16 MiB. Store large payloads as artifacts and return their reference.

A checkpoint only resumes the workflow that wrote it. start() rejects a checkpoint whose identity differs.

Part of the identityWhere you set it
Workflow namedefineWorkflow(name, tasks)
Versioncheckpoint.version
GraphTask keys and their after dependencies
Execution settingstimeoutMs, retry settings, whether condition or retry.accepts is set
GatesKind, prompt, actors and authentication of each approval or pause
Loop tasksmaxRounds
Interactive tasksactors, agent, model, brief, repository, maxTurns and sandbox provider

Other task code, briefs and workflow inputs are not part of it: change version when you change them. A saved run cannot move to another identity, including a new version: start it again under a new runId.

The workflow budget is not part of the identity either.

To reuse results across different runs, use the result cache instead.

A run that ended with a failed, cancelled or interrupted task resumes only with resume: "retry-incomplete". This authorizes running those tasks again, with their side effects.

import {
  createWorkflowCheckpointStore,
  createLocalTransport,
} from "@elie-laloum/outpost";

export const store = createWorkflowCheckpointStore({
  transporter: createLocalTransport({ directory: ".outpost/storage" }),
});
import { defineTask } from "@elie-laloum/outpost";

export let calls = 0;
export const upload = defineTask({
  key: "upload",
  perform: () => {
    calls += 1;
    if (calls === 1) throw new Error("Network unavailable");
    return { uploaded: true };
  },
});
export function uploadCount() {
  return calls;
}
import { reportValue } from "./reporter.ts";
import { defineWorkflow } from "@elie-laloum/outpost";
import { upload } from "./upload.ts";
import { store } from "./upload-store.ts";

export const workflow = defineWorkflow("upload", [upload]);
export const checkpoint = { store, runId: "upload-1", version: "1" };
reportValue((await workflow.start({ checkpoint })).status);
// Example output: failed
export const resumed = await workflow.start({
  checkpoint: { ...checkpoint, resume: "retry-incomplete" },
});
reportValue(resumed.status);
// Example output: done

It prints failed, then done. Without resume, the second start() rejects.

Saved stateWithout resumeWith resume: "retry-incomplete"
Every task done or skippedReturns the saved result, runs nothingSame
Paused at a gate or waiting for an answerContinues with your decisions or answersSame
Paused by a quotaRuns the task again after its resetSame
A task failed, cancelled or interruptedstart() rejectsReruns it and the tasks it skipped

Rerun tasks start a new series of retry attempts. done tasks never run again. A run stopped by its budget resumes the same way; pass a larger budget, since usage keeps adding up.

While start() is running, the process owns the checkpoint. A normal return releases that ownership. If the process dies, the ownership record remains and prevents another run from starting under the same runId until you recover it explicitly.

import { createHash } from "node:crypto";
import {
  createLocalTransport,
  recoverWorkflowCheckpoint,
} from "@elie-laloum/outpost";

const transporter = createLocalTransport({ directory: ".outpost/storage" });
const runId = "scan-2026-09";
const digest = createHash("sha256").update(runId).digest("hex");
const saved = await transporter.read(`checkpoints/${digest}.json`);
if (saved)
  await recoverWorkflowCheckpoint({
    transporter,
    runId,
    revision: saved.revision,
  });
Drag to move · Ctrl + scroll to zoom
100 %
  • StopMake sure the old runner no longer writes.
    1. Stop the processConfirm it exited. A PID does not prove that a remote runner stopped.
    (Steps)
    • → Unlock : then
  • UnlockClear the owner, keep the progress.
    1. Read the revisionRead the run’s checkpoint object from the transport. Transport
    2. Clear the ownerIt rejects if the object changed since you read it. recoverWorkflowCheckpoint()
    (Steps)
    • → Resume : then
  • ResumeStart the same workflow with the same checkpoint.
    1. Authorize the replayThe interrupted task reruns with resume: "retry-incomplete". start()
    (Steps)

The checkpoint’s key is checkpoints/ followed by the SHA-256 of the runId. Progress stays intact; start the run again with resume: "retry-incomplete".

defineWorkflowJob() runs each job under its runId. A queue returns the existing job for an ID it already has, so a finished job never runs again.

To continue the run, enqueue a new job ID with the same runId and the same input:

import { createSqliteTaskQueue } from "@elie-laloum/outpost";

const queue = await createSqliteTaskQueue(".outpost/jobs.sqlite");
try {
  await queue.enqueue({
    id: "fix-42-resume-1",
    handler: "fix",
    input: { runId: "fix-42", input: { issue: 42 } },
  });
} finally {
  queue.close();
}

A different input changes the checkpoint version and the job fails. To replay failed or interrupted tasks, the handler needs checkpoint: { store, version, resume: "retry-incomplete" }.

The store accepts any Transport. Use an S3 or R2 transport so that workers on several machines share runs: see Where data lives.

  • One start() at a time per runId. A second one rejects while the first runs.
  • Ownership never expires on its own. Clear it with recoverWorkflowCheckpoint() after stopping the old runner.
  • Replay repeats side effects that an interrupted task already made. Deduplicate them with context.idempotencyKey: see Job queues and workers.

API: createWorkflowCheckpointStore · WorkflowCheckpointOptions · recoverWorkflowCheckpoint · WorkflowCheckpoint · defineWorkflowJob.