Observe runs with OpenTelemetry
Collect scoped execution events and export traces and metrics.
Observe a whole run
Section titled “Observe a whole run”Create an observation hub when you want workflow, agent and resource events in one place. Add sinks that receive the events, then pass the hub as observation to your run. Each event includes its execution scope.
The sink prints workflow transitions, sandbox and Git operations and the agent’s events, in seq order. dispatch() accepts the same observation option for a single task.
The event envelope includes the context of the run.
API reference: Observation.
Carry the scope into your own tasks
Section titled “Carry the scope into your own tasks”defineAgentTask and defineIsolatedTask attach their dispatch to the task’s scope. In a task you write with defineTask, pass context.observation to each dispatch.
Without it, the dispatch still reports to its own observe, but its events never reach the hub. observation.child(scope, sinks) derives a hub that adds scope fields; sinks passed to a child receive only the events emitted below it.
Speculation takes the same observation option, and the recovery and retention helpers accept a hub as their last argument.
What the hub receives
Section titled “What the hub receives”The hub receives agent activity alongside workflow transitions and resource operations.
API reference: ObservationEvent.
An operation event pairs started with finished or failed through its id. Only the terminal event carries durationMs.
Built-in harness events
Section titled “Built-in harness events”The built-in harness reports its loop with these kinds. They also reach observe.
API reference: AgentEvent.
tool-result keeps only a 2,000-character preview of a result; subscribe to tool-output for the full stream.
Delivery and failures
Section titled “Delivery and failures”A sink that returns nothing runs during emission, so keep it fast. A sink that returns a promise gets its own ordered queue. An optional flush() on the sink runs whenever the hub drains.
dispatch() and start() drain their deliveries before they return. flush() drains the hub at any time; close() drains it and stops accepting events.
| Situation | What happens |
|---|---|
A sink queue already holds capacity events (default 1,024). | New events for that sink are dropped and dropped increases. |
A delivery exceeds deliveryTimeoutMs (default 5,000). | The sink is disabled. Its pending promise keeps running. |
| A sink throws or rejects. | The error joins errors and the run’s observerErrors. The run is unaffected. |
The sink prints the numbered workflow events, then the script prints done 0 0. Check dropped and errors before treating a trace as complete.
When agent output is too large
Section titled “When agent output is too large”A single protocol line above 16 MiB stops a CLI agent. The hub and observe receive a raw event holding its first 2,000 characters, with bytes and truncated: true, then stopped with reason oversized-event. The dispatch fails with code process (Errors).
Export to OpenTelemetry
Section titled “Export to OpenTelemetry”Install @opentelemetry/api and an OpenTelemetry SDK, then register the SDK and its exporters before you create the observer. Its sink turns hub events into linked spans and metrics.
The trace nests outpost.workflow, outpost.task, outpost.task.attempt and outpost.dispatch spans, with one span per operation such as outpost.sandbox.acquire. Without a registered SDK, the API handles export nothing.
API reference: createOpenTelemetryObserver.
telemetry.close() ends the spans still open. Your application flushes and shuts down the SDK. Pass onError to receive instrumentation errors; they never change a run’s outcome.
Without a hub
Section titled “Without a hub”Pass the observer as telemetry to start() for workflow, task and attempt spans, or to dispatch() for a dispatch span. Operation spans need the hub.
Handle agent events asynchronously
Section titled “Handle agent events asynchronously”createCustomReporter() builds an observe callback from handlers keyed by event kind. Handlers may be asynchronous; they run on a bounded queue like hub sinks.
The dispatch waits for pending handlers before it returns and reports the first handler error in result.observerErrors. report.flush() rethrows that error. The second argument takes onError, called for each failure, plus the hub’s capacity and deliveryTimeoutMs.
Limits
Section titled “Limits”- The hub is a live, in-memory stream: it stores nothing, and a slow sink loses events. To read events after the run, use the dispatch’s journal, itself a sink with the same
capacityanddeliveryTimeoutMsbounds. - A disabled sink stays disabled for the hub’s lifetime, and
errorskeeps the first 100 errors. - A closed hub ignores new events. A hub reused across runs keeps its
errorsanddroppedcount, so each run’sobserverErrorsincludes earlier errors. - Events emitted on a remote worker stay on that worker’s hub.
API: createObservationHub · ObservationHub · Observation · ObservationEvent · OperationEvent · createOpenTelemetryObserver · OpenTelemetryObserver · createCustomReporter.
Observe decisions and selections
Section titled “Observe decisions and selections”Decision evaluations emit lifecycle summaries with source decision. Routed harnesses emit model-route agent events identifying the effective model, selection reason and optional native confidence. Pass observation to decide() for a direct evaluation; decision tasks and harnesses propagate workflow, task, pass and subagent scopes. Full states and answers require a verbose hub. Valid decision usage is accounted synchronously, independently of sink delivery or failures.