Build a development workflow
Take a ticket through questions, a reviewed plan, failing tests and verified code.
This example takes a ticket for a CSV export through a complete development workflow. An agent asks the owner for missing requirements and proposes a plan. After approval, agents write failing tests, implement the change and submit it to a reviewer. The accepted work is committed on outpost/shop-142, with separate commits for tests and implementation.
The workflow follows one rule from redline, a ticket-to-merge-request system built on Outpost: code decides, agents judge. Agents answer and edit files; your code validates their answers, runs the tests, refuses edits outside each role’s files and commits.
What this example covers
Section titled “What this example covers”- Interactive tasksThe agent questions the owner, one point at a time, then returns the plan.
- ApprovalsThe owner approves the plan before any file changes.
- Verification loopsWrite, check, feed the rejection back, within a number of rounds.
- Typed responsesThe reviewer returns a validated verdict.
- Sandbox sessionsOne warm sandbox keeps dependencies between agents and test runs.
- Durable runsA checkpoint keeps answers, the plan and finished loops between processes.
Write the script
Section titled “Write the script”Save the files shown in the tabs next to the outpost.config.ts from Installation. They are grouped by role: clarify the ticket, prepare the sandbox, write tests, implement the change and expose the functions your application calls. Each file has one responsibility; the imports connect them.
Clarify the ticket and approve the plan
Section titled “Clarify the ticket and approve the plan”Start with the ticket, its expected plan and the approval gate.
Prepare and reuse the sandbox
Section titled “Prepare and reuse the sandbox”These helpers open one sandbox when delivery starts and reuse it for commands and agent turns.
Write and check the tests
Section titled “Write and check the tests”The test loop rejects changes outside test files and requires the tests to fail before committing.
Implement and review the change
Section titled “Implement and review the change”The implementation loop checks the write boundary, runs the tests and asks a reviewer before committing.
Compose the workflow and its checkpoint
Section titled “Compose the workflow and its checkpoint”Compose the tasks and checkpoint, then close the sandbox after each call.
Submit answers and decisions
Section titled “Submit answers and decisions”Your application imports the entry points from run.ts to submit answers and decisions.
Your application renders questions in a form and plan on a review page. Each call rebuilds the same workflow and checkpoint, so it can run in any process on the machine that holds the repository.
Understand the steps
Section titled “Understand the steps”Each check runs in a fixed order, and the first rejection becomes the next round’s feedback:
| Loop | Check, in order | Rejects when |
|---|---|---|
tests | Write zone | No test changed, or a non-test file changed |
tests | Red run: npm test | The tests already pass: they prove nothing |
code | Write zone | A test file changed |
code | Green run: npm test | The tests fail; their output is the feedback |
code | Reviewer agent, defineJsonResponse() | approved is false |
Agents never commit. Each loop commits only once every check accepts, so a rejected round leaves its changes for the next attempt to fix. answer() and decide() return after the next agent turns, which can take minutes: from a web request, hand them to a job queue worker.
Adapt the example
Section titled “Adapt the example”| Variation | Change |
|---|---|
| Read the ticket | Add a first defineTask() that fetches the ticket from your tracker and returns { key, text }; read it with context.value() in the briefs instead of the constant. |
| An adversarial reviewer | Pass another agent in the reviewer’s request, such as Claude Code with createClaudeHarness(): a different model misses different things (Choose an agent). |
| Stricter checks | Add a typecheck and a lint to the green run, or a second reviewer that checks the tests fail for the right reason, as redline does. |
| Open a merge request | Once progress() returns branch, push it from the host and open a draft pull request with gh pr create --draft --head, or a merge request with glab mr create --draft. |
| Several repositories | One loop pair per repository, chained with after in dependency order (Change several repositories). |
Limits
Section titled “Limits”- Rejection ends the run: A rejected plan skips delivery and the run ends
failed. Start a newrunIdwith the owner’s note in the brief. - Malformed plan:
planthrows, and the run fails without asking the agent again. Use a loop task to let the agent fix its own JSON. - Rounds run out: A loop whose last check rejects fails with
LoopTaskExhausted; finished tasks stay in the checkpoint and the branch keeps the committed tests. - Same machine: The repository, its worktrees and the framing conversation must be reachable by the process that answers.
- Interrupted round: After a crash right after a commit, the resumed check finds no change and spends a round. Inspect the branch before you resume.
- Retained work: Framing keeps its worktree on an
outpost/interactive-…branch; clean it up when done.
API: defineInteractiveAgentTask · defineApprovalTask · defineLoopTask · defineAgentTask · defineJsonResponse · createSandbox · LoopTaskExhausted.