Skip to main content
agent.stream() runs an agent exactly like agent.send() and reports the run as a stream of typed events: model text as it arrives, every tool call, step boundaries, approval requests, and a final run.done. Every event is a plain JSON object, so you can forward it to a browser over Server-Sent Events or a WebSocket without converting it.
The full API has the same method: AgentExecutor.stream(options) takes the same options as AgentExecutor.execute() (approval store, checkpoints, hooks, tracing, toolConcurrency, …). These events, the AgentEvent union, are the SDK’s one event system: the stream, the listener options, session.on(), the UI hooks, the server routes, ACP and channels all carry them.

The AgentRun handle

stream() returns an AgentRun:
  • The run starts immediately. You do not have to iterate: await run.result on its own drives the run to completion.
  • result resolves with the same ExecutionResult that send() returns, and rejects with the same error send() would reject with. An aborted run resolves with finishReason: 'aborted'; a run paused for approval resolves with finishReason: 'awaiting-approval' and approvalId. Leaving result unawaited never causes an unhandled rejection.
  • Breaking out of the for await loop early aborts the run (through the same path as signal, see Cancellation). result then resolves with finishReason: 'aborted'.
  • No backpressure; nothing is dropped. The run never waits for the consumer. Events are buffered until you read them, so a slow consumer sees every event, in order. Iterating after the run finished still yields all of its events.
  • enqueue(input) adds user input to the run while it runs: it joins the transcript before the next model call. See Queued input.
  • steer(input) redirects the run: a model call that has not produced anything yet is aborted and made again with the input. See Steering.
  • One consumer. An AgentRun can be iterated once; a second for await throws. To fan out, collect the events yourself.
  • Invalid options (no provider, agent or input, a bad toolConcurrency) throw synchronously from stream().

Listening without iterating

To observe every run without iterating one, give the agent a listener: createAgent({ onEvent }), or onAgentEvent in the AgentExecutor.execute() / stream() / resumeAfterApproval() options. It is called synchronously with each AgentEvent as it happens, on send() as on stream(), for session turns and for runs resumed after an approval. It gets the same events, in the same order, as iterating the stream would, with one difference: send() / execute() generate each model step whole, so a step’s text comes as a single text.delta (as it does on stream() with a provider that cannot stream). Sub-agents’ events arrive tagged with subagent.

Migrating from onEvent / ExecutionEvent

ExecuteOptions.onEvent with ExecutionEvent is the SDK’s older event callback. It is deprecated: it still works (it logs a one-time console.warn), with its events now derived from the run’s AgentEvents, and it will be removed in a future major version. Replace it with onAgentEvent (or createAgent({ onEvent }), which already takes AgentEvents): AgentEvent also reports what ExecutionEvent never did: step boundaries, approval requests, permission decisions, budgets, guardrails, queued and steered input, compaction, reasoning, provider retries and agent drift (see the schema).

Streaming a session turn

session.stream(input, { signal }) streams one turn of a multi-turn session and returns the same AgentRun. The run sees the conversation so far, and when it ends its turn is saved to the session’s store, just as session.send() saves it. run.done is delivered after the save, so the transcript is complete when the loop ends. An aborted or failed run, or one you stop reading early, is not saved.

Streaming after an approval

A run that paused for an approval can continue as a stream too. streamResumeAfterApproval() takes the arguments of resumeAfterApproval() and returns an AgentRun whose result is what resumeAfterApproval() returns. Its events are run.start, the decided call’s tool.start and tool.done (tool.error for a rejection), then the continuation’s events exactly as in a fresh run, up to run.done; a further pause ends it with approval.requested. Cancellation, enqueue() and steer() work as on any run, and the approval can come from another process, since everything is read from the approval store. For createAgent() agents use agent.approvals.streamResolve() or streamAnswer(). A model chosen per run (model as a function) is the one the paused run used.

Event schema (version 1)

Every event has these fields: The event types and their extra fields: Optional fields are left out when they have no value. They are never undefined, so JSON.parse(JSON.stringify(event)) returns an equal object.

Ordering guarantees

  • run.start is first and run.done is last, each exactly once.
  • Each step.start is followed by exactly one step.done with the same step, before the next step.start. Everything a step does happens between the two.
  • Inside a step: compaction.start / compaction.done (if the request was compacted, before the model call), provider.retry / provider.fallback events (if the model call fails), then text.delta events, then text.done, then the tool events.
  • Tool calls of one step run in parallel (see toolConcurrency): tool.start events come in the model’s call order and tool.done / tool.error events in completion order. Match them by toolCallId.
  • input.queued comes when run.enqueue() is called (after run.start, even for input queued before the run got going), so it can fall inside a step. Its input.applied comes between the step.done of the step that was running and the next step.start, which carries the step it names.
  • input.steered comes when run.steer() is called. With mode: 'immediate' the running step ends next with step.done ('steered', no text events from its aborted call), then input.applied and the next step.start.
  • When a call needs approval, the calls before it run and report, then approval.requested, step.done ('awaiting-approval') and run.done ('awaiting-approval') follow. Calls after it never start.
  • When a guardrail blocks, guardrail.tripped comes just before the end: for an input guardrail right after run.start (no step starts); for an output or tool guardrail inside the step, followed by step.done and run.done ('guardrail'). A blocked output has no text.done; calls after a blocked tool call never start.
A text-only run:
A run with one tool call:

TypeScript

AgentEvent is a discriminated union: narrowing on event.type gives the payload. Each event type is exported too (TextDeltaEvent, ToolDoneEvent, RunDoneEvent, …), and AgentEventOf<'tool.done'> picks one by name.

Versioning

v changes only when an existing event changes incompatibly (a field removed, renamed or retyped). New event types and new optional fields can be added without changing v, so ignore event types you do not know.

Sub-agents

When the agent delegates with the task tool or a createDelegateTool() tool (see Sub-agents), the sub-agent’s run is streamed inside the same stream: its steps, text.deltas, text.dones, tool events and errors appear between the lead’s tool.start and tool.done (or tool.error) for that call, each with a subagent field:
toolCallId is the lead’s tool call that started the sub-agent; a sub-agent of a sub-agent has depth: 2 and the enclosing one as parent. Events without subagent are the top-level run’s: the ordering guarantees above hold for them, and separately for each sub-agent’s steps (several sub-agents running in parallel interleave). run.start and run.done belong to the top-level run only, so they still come exactly once. A sub-agent that pauses for approval is reported once, by the top-level approval.requested (which carries the sub-agent’s call).

Token streaming and providers

When the provider implements stream() (all built-in providers and mockModel do), each model step is streamed and text.delta events arrive as the model produces text. Tool calls are assembled from the stream. When a provider has no stream(), or its supportsStreaming(model) returns false, the step falls back to generate() and its text arrives as a single text.delta followed by text.done. Streaming changes only how one model step is obtained. Hooks, argument validation, approvals, parallel tool calls, checkpoints, cancellation, tracing spans and the other execute() callbacks behave exactly as in send(). send() and execute() themselves still use generate(). A custom provider’s stream() should yield text-delta chunks and a final finish chunk with finishReason and usage. Tool calls can be yielded as tool-call chunks, or resolved on the toolCalls promise of the StreamResult. An error chunk fails the step.

Cancellation

Pass signal to stop a run from outside, or break out of the loop. Both end the run the same way as an aborted send(): the signal is checked between stream chunks, before every model call and every tool call, and reaches the provider and the tools. The stream ends with step.done ('aborted', when a step was running) and run.done ('aborted'), and result resolves with finishReason: 'aborted'.

Queued input

run.enqueue(input) adds user input (a string, content parts or a Message[]) to a run that is still going, for example a follow-up the user types while the agent works. The run does not stop: the input joins the transcript at the next safe point - after the current step’s tool results, never between a tool-call turn and its results - and the next model call sees it, as if the user had typed it. Input queued while the model writes its final reply gets one more step, so the model answers it in the same run.
enqueue() returns { id, applied }. id is on the input.queued and input.applied events. applied is false when the run had already finished: the input was not taken, so send it yourself, as a new turn (session.send()). Otherwise it is a promise that resolves to true once the input is in the transcript, or to false when the run stopped before its next model call (aborted, paused for approval, out of maxSteps or budget, or failed); the input is then left to you too.
  • Several inputs queued before the same model call are applied together, in the order they were queued.
  • With checkpointing (sessionId or a durable session), an input that is still waiting is saved at the end of the run’s checkpoint right away, the same slot as input queued behind unanswered tool calls (see Durable execution). A crash, or a run that fails, does not lose it: agent.resume() applies it (applied of a run that failed in-process still resolves false). A run that finishes, pauses or is aborted does not keep it.
  • Without a stream, pass an InputQueue as ExecuteOptions.inputQueue and call its push(): the same queue that is behind run.enqueue(). One queue serves one run.
  • In a session, agent.session({ turnPolicy: 'queue' }) makes a send() or stream() made while a turn runs (or waits to start) join that turn this way. See Sessions.
Queued input waits for the step that is running. To redirect the run at once, steer it (next section).

Steering

run.steer(input) redirects a running run to new user input, for example when the user changes their mind while the agent is still thinking:
  • If a model call is in flight and has not emitted text or tool calls yet, it is aborted (with a signal of its own, not the run’s: the run goes on), its partial output is discarded, input is appended as a user message and the model is called again. applied is 'immediate', and the aborted call’s step ends with step.done 'steered'. A provider that ignores the abort signal is not waited for; its late reply is dropped.
  • If the model has already emitted text (or tool calls), input waits for the next safe point like enqueue(): applied is 'queued'.
  • Tool calls already running finish and their results are kept; calls of that turn that have not started yet are not run and get the usual “cancelled before it ran” result. Then the input is applied.
  • applied is false once the run has finished: send the input as a new turn.
steer() returns { id, applied, joined }: id is on the input.steered and input.applied events, and joined is a promise of whether the input made it into the transcript (as enqueue()’s applied). The aborted call counts as a step of maxSteps. With checkpointing, the input is saved at the end of the checkpoint as soon as it is taken and before the model is called again, so a crash during the redirect does not lose it; the discarded partial turn is never checkpointed. Without a stream, call steer() on the run’s ExecuteOptions.inputQueue. In a session, agent.session({ turnPolicy: 'steer' }) makes a send() or stream() made while a turn runs steer that turn (see Sessions).

Example: terminal

Example: Server-Sent Events

Events are JSON-serializable, so an SSE endpoint is one res.write per event. Abort the run when the client disconnects.
In the browser:
For a React chat UI over this kind of endpoint, see React: useLoushoAgent() POSTs the input and reads the same data: lines. lousho dev serves this format for a session per browser tab: POST /chat with { sessionId, input } streams agent.session({ id }).stream(input) as data: lines ending with event: done, and approvals are decided through POST /chat/:sessionId/approvals/:id (see CLI).

Example: the full API

AgentExecutor.stream() accepts every execute() option. Here a tool that needs approval ends the stream with approval.requested: