Limit attempts and token usage
Set a workflow budget and understand how usage is counted across retries and resumes.
Set a workflow budget
Section titled “Set a workflow budget”Pass a budget to the workflow’s start() method to limit attempts, reported tokens or both. These limits apply across the workflow’s tasks, rather than to each task separately.
It prints { attempts: 1, tokens: { input: 10, cached: 0, output: 5 } }.
API reference: WorkflowBudget.
speculate() requires the same budget, shared by its candidates: see Competing candidates.
Understand attempt counts
Section titled “Understand attempt counts”Each attempt is admitted against budget.attempts before it starts.
- Task attemptEach run of a task, including every retry.
- Loop roundEach round of a loop task.
- Speculative candidateEach candidate that
speculate()starts.
One agent turn counts as one attempt, even if it makes many model requests. A skipped task or a result restored from the cache uses no attempt.
Report usage
Section titled “Report usage”Agent task helpers report the tokens of their agent automatically: defineAgentTask(), defineIsolatedTask(), defineInteractiveAgentTask() and defineQueuedTask(). A custom task that calls a model reports what it spent.
reportUsage(usage) adds to the totals. reportUsageOnce(receipt, usage) ignores a receipt the task already recorded, even after a checkpoint resume, so a result read twice is counted once. Both work only during the running attempt.
Understand budget limits
Section titled “Understand budget limits”Outpost checks the reported usage before allowing more work to start. The budget applies to these recorded totals; it cannot predict the tokens an in-progress model request will consume.
| Limit | When it is checked | What happens |
|---|---|---|
attempts | Before each attempt | No new attempt starts; running ones finish. WorkflowBudgetExceeded with "attempts". |
usage.input, usage.output, … | Before each attempt and at each usage report | Running attempts are cancelled; nothing else starts. WorkflowBudgetExceeded. |
Token limits, no attempts, incomplete usage | Before each attempt and at each usage report | Running attempts are cancelled; nothing else starts. WorkflowUsageUnavailable. |
A limit stops the run once the total reaches it. The run then ends with status: "failed", the stopped tasks are cancelled, and result.errors holds the error with its dimension, limit and observed values.
Handle incomplete usage
Section titled “Handle incomplete usage”result.usage.tokens.complete === false means some tokens could not be measured: the counters are a lower bound. The marker stays through retries, aggregation and checkpoints.
With token limits and no attempts, incomplete usage stops the run with WorkflowUsageUnavailable. With attempts, the run continues under the attempt limit and Outpost emits a warning.
Copilot CLI and Kimi Code read their final counters from the session after the CLI exits. For them especially, bound each run by attempts and time.
timeoutMs bounds each task attempt and deadlineMs each agent turn: see Limits and cancellation. Choose an agent shows when each agent reports usage.
Keep totals across resumes
Section titled “Keep totals across resumes”A checkpoint saves result.usage. A resumed run starts from the saved totals, so the budget covers the whole run, not the current process.
To continue a run its budget stopped, start it again with a larger budget and authorize the replay of its cancelled tasks: see Durable runs. A durable speculative race must resume with the budget it started with.
Set other limits
Section titled “Set other limits”- Limits and cancellationBound one agent turn in time or silence.
- Concurrency, retries and timeoutsBound each task attempt and a whole run in time.
- Built-in harnessBound model requests, tool calls and tokens within one turn.
API: WorkflowBudget · WorkflowUsage · Usage · TaskContext · WorkflowBudgetExceeded · WorkflowUsageUnavailable.
Include decision usage
Section titled “Include decision usage”Decision tasks contribute their normalized usage to workflow budgets. Model routing also counts against harness, ancestor and workflow budgets, exactly once per router request. Valid usage still counts when truncation rejects the result. A missing receipt is incomplete usage, and strict token budgets reject continuation with unknown consumption.