Building and running agents
Dynamic config
createAgent()’s model, instructions (or prompt) and tools each take
the static value or a function of the run, (ctx) => value | Promise<value>
(type PerRun<T>). ctx is { sessionId?, input, metadata? } (type
RunConfigContext): the sessionId of send() / stream() or the
agent.session() id, the run’s user input, and the metadata call option of
send(), stream(), session.send() and session.stream(). The functions
run once when a run starts, before the first model call, and again on every
session turn. Everything else applies to what they return: fallbackModels
and retry, projectInstructions, memory, skills, sub-agents, MCP tools,
permissions, approvals and guardrails. A dynamic agent used as a sub-agent
resolves with the task prompt as input.
LOUSHO_CONFIG_RESOLVER_FAILED
(error.field names the option, error.cause is the thrown error): send()
rejects, stream() ends with an error event and a session keeps its
transcript as it was. A run paused for an approval or a question keeps the
ctx and the model it resolved in the paused snapshot: resuming it, even from
another process, uses that model (it never switches mid-turn) and calls the
tools function again with the same ctx. A run resumed from a crash
checkpoint (agent.resume(id)) resolves again with input: [] and no
metadata. With only static values nothing changes: the agent is built once,
when it is created.
UI bindings
@lousho/build-ai-agent/react exports useLoushoAgent(source, options?), a
React hook that runs an agent in process ({ agent, sessionId? }) or over HTTP
({ url }) and returns messages, status, pendingApproval,
send(), stop(), approve() and reject(). Its framework-neutral parts,
reduceAgentEvents() and parseEventStream(), are exported too. See
React. @lousho/build-ai-agent/vue exports the same
useLoushoAgent() as a Vue 3 composable, with the state as refs. See
Vue. @lousho/build-ai-agent/svelte exports loushoAgent(), the
same state and actions as a Svelte store ($agent). See Svelte.
Sub-agents
Passsubagents: { researcher, writer } (agents from createAgent() with a
description, or a { list, resolve } catalog) to createAgent() or
AgentExecutor.execute(): the lead gets one task tool and a prompt listing,
each sub-agent runs on the task prompt alone, and it inherits the lead run’s
signal, hooks (ctx.subagent), tracing, approval store and event listeners
(event.subagent). maxSubagentDepth (default 1) bounds nesting. See
Sub-agents.
Approvals
AcreateAgent() agent pauses on a needsApproval tool instead of failing:
send() (and session.send()) resolves with finishReason: 'awaiting-approval'
and an approvalId. agent.approvals.list() returns the pending calls and
agent.approvals.resolve({ id, approved, note? }) runs or rejects the call and
resolves with the continued run’s result (continuing the session it paused
in). Pauses are kept in a per-agent InMemoryApprovalStore unless you pass
approvalStore (e.g. SqliteStore.approvals) or a store; approve: (call) => boolean | string
decides each call in code without pausing (stream() still ends at the
pause). With askQuestion: true the agent can ask the user a question
(kind: 'question'), answered with agent.approvals.answer({ id, answer }).
See Approvals.
Structured output
Passoutput: zodSchema to createAgent() (or AgentExecutor.execute() /
stream()): the final reply must be a JSON object matching the schema, and
send() / run.result resolve with it validated as result.object, typed
z.output<typeof schema> (result.text keeps the raw JSON). Tools still run
first. Each model call carries a responseFormat: { type: 'json', schema }
hint, which the ai-SDK providers map to JSON mode. An invalid reply gets one
repair step listing the issues (it counts against maxSteps); if that is
invalid too, the run ends with finishReason: 'output-invalid' and
outputError: { message, issues }. See Structured output.
Skills
Passskills: [defineSkill({ name, description, content }), ...(await loadSkills(dir))] to
createAgent() or AgentExecutor.execute(): only names and descriptions go in
the system prompt and the model loads bodies through an auto-registered
load_skill tool. See Skills.
Memory
Passmemory: [defineMemory({ name, scope, provider })] to createAgent() to
keep items across conversations: each run recalls a slot’s newest items into
the system prompt on its first model call, and the model gets
remember_<name> / recall_<name> tools. scope is 'global', 'session'
or a function of { sessionId, metadata }; inMemoryMemory() and
fileMemory({ dir }) are the built-in providers. See Memory.
Durable execution
sessionId + checkpointStore make a run crash-safe and a session
multi-turn: the run is checkpointed after every model response, every tool
result and every pause, and calling execute() again with the same
sessionId resumes an unfinished run (without re-calling the model for a
turn it already has), continues a finished conversation with the new
input, or throws SessionAwaitingApprovalError while an approval is
pending. Tools run at-least-once across a crash; execute receives the
call’s toolCallId to use as an idempotency key. See
Durable execution for the exact guarantees.
Cancellation
Pass anAbortSignal to stop a run: agent.send(input, { signal }),
AgentExecutor.execute({ ..., signal }) or
resumeAfterApproval(..., { signal }).
AbortSignal.timeout(ms):
- The signal is checked before every model call and every tool call. It is
passed to the provider (
GenerateOptions.signal, sent to theaiSDK asabortSignal) and to each tool asexecute(args, { abortSignal }), so in-flight work can stop early. The built-inhttpToolpasses it tofetch, and agents created withcreateDelegateTool()are aborted along with their parent. - An aborted run resolves (it does not reject) with
finishReason: 'aborted'and the messages and steps so far. A rejection caused by the abort, such as anAbortError, is not treated as a failure: it is not retried by the provider retry and fallback wrappers and is not compacted into a provider error. - The run’s events end with
run.donewithfinishReason: 'aborted'. - With
sessionId+checkpointStore, the state is checkpointed. Callingexecute()again with the samesessionIdresumes where the run stopped; newinputis appended as the next user message (see Durable execution). Tool calls the run never reached get an{ error }result saying they were cancelled, so the conversation stays valid for the provider. - An already-aborted signal returns at once without calling the provider.
Finish reasons
result.finishReason says why a run ended: the model’s own reason for its last
turn ('stop', 'length', 'tool_calls', 'content_filter', 'error'),
'awaiting-approval' (paused on a tool call that needs a human), 'aborted'
(cancelled with signal), 'max-steps', 'output-invalid' (the reply
did not match the output schema even after the repair step, see
Structured output), 'budget-exceeded' (a
limits budget such as maxTokens or maxCostUsd tripped; result.budget
says which, see Budgets), or 'guardrail' (an
input, output or tool guardrail blocked; result.guardrail says which, see
Input and output guardrails). 'max-steps' means the maxSteps
budget (default 10) ran out while the model still wanted to continue, so the
reply may be empty or partial; a run that finishes naturally within the budget
keeps its 'stop'. Steps carried over by initialSteps or an approval resume
count against the budget, and result.steps is the number of steps taken. The
same reason is on the finish event and on run.done in
streaming.
Parallel tool calls
When the model asks for several tools in one turn, they run concurrently.toolConcurrency (on createAgent() and AgentExecutor.execute()) caps how
many run at once: a positive integer, or 'unbounded' (the default). Use 1
for strictly sequential execution, e.g. when your tools share state that is
not safe to touch concurrently.
- Transcript order is call order. Tool results are appended in the order the model requested the calls, not the order they finish, so the next provider request is deterministic.
- Events. Calls start in call order. A call’s
tool-callevent,onToolCall, argument validation,preToolCallhooks andneedsApprovalcheck run just before it starts, one call at a time. Itstool-resultevent fires when it finishes, so results arrive in completion order. WithtoolConcurrency: 1events alternate call/result exactly as before. - Approvals. The first call that needs approval stops the batch: the
calls before it run (concurrently) and their results are recorded, then the
run pauses on that call (
finishReason: 'awaiting-approval'). Calls after it never start in this run.resumeAfterApproval()records the paused call’s result (or rejection) and then runs those later calls the same way, so every call of the turn gets exactly one result - see Durable execution. - Failures are isolated. A tool that throws gets its own error result;
its siblings carry on. A propagating error (
PropagatingToolError, such as the delegation depth guard, or a throwing hook) stops new calls from starting, waits for the running ones to settle, then rejects the run. No tool is left running detached. - Cancellation. An abort while a batch runs resolves with
finishReason: 'aborted'. Calls that finished keep their results; the rest get a “cancelled” result. Running tools see the abort through theirabortSignal. - Checkpoints. With
sessionId+checkpointStore, the model’s turn is checkpointed before any call starts, then again each time the in-order run of finished calls grows (with1, after every call). A resumed run never runs a recorded call again, and never asks the model again for a turn it already checkpointed.
Declarative specs
Providers
Message.content is a string or a list of ContentParts (text, image,
file); the built-in providers send image parts on user messages. See
Multimodal input.
See Providers for how a model string is resolved and which model runs.
Testing
Exported from@lousho/build-ai-agent/testing (see Testing agents).
Tools
See Tools for a guide to defining and registering tools.
displayName, needsApproval (boolean or predicate),
requiresSandbox and sandboxExecute. Defining a tool validates its name,
description and zod input immediately; registering two tools with the same
name throws an error naming the conflict.
A defined tool carries its schema as inputSchema (the same zod schema as
input) and its execute function directly. These are the canonical fields of
a ToolDescriptor; the .tool object (an ai v4 { description, parameters, execute }) is legacy, still built for compatibility, and used only for
descriptors that do not set inputSchema / execute.
inputSchema is converted to a
zod schema, and the model’s arguments are validated against it before the
server is called. The converter handles type (including arrays such as
["string", "null"]), enum with mixed types, const, anyOf / oneOf
(a union; anyOf: [X, { type: "null" }] becomes X.nullable()), allOf
(merge / intersection), local $ref into $defs / definitions,
properties / required, additionalProperties (boolean or schema), items,
default, and the constraints minimum, maximum, exclusiveMinimum,
exclusiveMaximum, minLength, maxLength, pattern, minItems and
maxItems. Where JSON Schema is ambiguous the conversion is permissive:
unknown keywords, empty schemas, tuple items and unresolvable refs become
z.any(), a recursive $ref is expanded once and z.any() is used for the
inner occurrence, an invalid pattern regex is skipped, and objects keep
extra properties unless additionalProperties is false. The converter never
throws on schema content.
Skipped tools. If one tool still cannot be converted, only that tool is
skipped; the rest of the server’s tools load. Pass a logger to receive a
warning naming the server, the tool and the reason, and onSkip to collect
what was left out:
isError: true is a normal tool failure: the
model receives { "error": "McpToolError", "toolName": "...", "message": "..." }
where message is the server’s text content. Successful results are
JSON-serializable: if the server returns structuredContent it is the result
object; otherwise the result is { text, content }, where text joins all
text parts and content keeps every part in order (text, image, audio,
resource, resource_link; an unrecognised part type is kept as
{ type: 'unknown', raw }). When structuredContent arrives together with
non-text parts the result is { structuredContent, text, content } with
content holding only the non-text parts, so images, audio and resources are
never dropped.
Todo tools
createTodoTools(options?) gives long-running agents a plan to track:
todo_write replaces the whole list ({ id?, content, status } items, status
pending | in_progress | completed, at most one in_progress) and returns
the list plus counts; todo_read returns it. Ids are assigned automatically and
stay stable when a later write repeats an item’s content. Invalid lists (for
example two in_progress) reach the model as a structured tool error so it can
retry. The list lives in memory per call; pass store ({ get, set }) to
persist it, and onChange to update a UI.
Tool errors
A tool call can fail in two ways. In both, the run continues and the model receives a structured JSON error as the tool result (never the stringnull),
so it can recover. tool-result events, tracing, onToolResult and
postToolCall hooks see the call as an error.
Argument validation. Before a tool runs, the model’s arguments are parsed with the tool’s zod
inputSchema (for a legacy descriptor, tool.parameters). Validation happens first, so pre-tool hooks, the
needsApproval predicate and execute all receive the parsed value
(defaults, coercions and transforms applied). Tools without a zod schema are
passed through unchanged.
If the arguments do not match, execute is not called, the run continues, and
the model receives a structured error as the tool result so it can retry
(tool-result events, tracing and postToolCall hooks see it as an error
result; preToolCall hooks are skipped because there is no valid call):
ToolArgumentsValidationError (with a typed issues array) is exported from
the package root.
Thrown errors. If execute throws, the model receives the error name,
the tool name and the message only (never a stack trace). Messages are capped
at 2,000 characters and end with ... (truncated) when cut:
{ error, toolName, message, kind } shape; see Errors for the kind values.
Errors extending PropagatingToolError (for example the delegation depth
guard) are the exception: they are rethrown and abort the run instead of being
shown to the model.
Models, tokens and cost
Dependency-free helpers for budgeting and context decisions.estimateTokens(input, { model?, estimator? })is a heuristic (about 4 characters per token for English, more for CJK and other scripts, plus per-message overhead and tool-call JSON). Expect roughly 15-20% error on English: fine for compaction and budgets, not for billing. Plug in a real tokenizer withsetTokenEstimator(fn)oroptions.estimator.getModelInfo(id)matches the exact id, thenprovider/id, then dated snapshots (gpt-4o-mini-2024-07-18resolves togpt-4o-mini). Unknown models returnundefined.registerModel(info)adds or overrides an entry; the latest registration wins.estimateCost(usage, model)returns USD, orundefined(not0) when the model or its prices are unknown.
src/models/modelData.ts). Providers change prices and models, so override entries with registerModel when you need billing-grade numbers.
Usage and cost of a run
EveryExecutionResult (and agent.send() result) carries usage, the running total of the whole run:
- Reported vs estimated. Providers report
promptTokens/completionTokens; the built-in providers also pass oncachedInputTokens/reasoningTokenswhen the ‘ai’ SDK’s provider metadata carries them (usage.cachedInputTokensandusage.reasoningTokensonly appear then). A backend that reports nothing givesundefinedusage, never zeros. For that step the executor falls back toestimateTokensand setsusage.estimated(a leading~informatUsage). A custom provider should leaveGenerateResult.usageunset rather than fill in zeros. - Cost.
costUsdis the sum ofestimateCostper model. It isundefined, never a misleading partial sum, as soon as any model used has unknown pricing;byModelshows which ones are priced. - Delegation. A delegated child’s usage is added to the parent’s totals and
byModel, and is also shown on its own asusage.delegated({ inputTokens, outputTokens, totalTokens, costUsd, modelCalls, estimated, runs }). - Resume. A run resumed from a checkpoint, or after an approval, continues from the saved totals instead of restarting at zero. Checkpoints written by older versions start from their saved token counts with an unknown cost.
- Events and traces. The
finishevent (and every lifecycle event that carriedusage) now carries the running totals;text-completealso hasstepUsage, andonLLMResponsereceives the call’s usage as a third argument. Thechatspan’sgen_ai.usage.*attributes use the same numbers, withlousho.usage.estimatedset totruewhen they are estimates. - Streaming.
agent.stream()events carry the same accounting:step.doneandrun.donehaveusagewithinputTokens,outputTokens,estimatedandcostUsd(run-level alsomodelCalls), alongside the olderpromptTokens/completionTokens. promptTokensandcompletionTokensonusageremain as deprecated aliases ofinputTokensandoutputTokens.
Context compaction
createCompactionHook({ thresholdPercent?, contextWindow?, protectedTokens?, strategy?, onCompaction? })
returns an AgentHook that, before each model call above 90% (by default) of
the model’s context window, replaces tool results older than the newest
40,000 tokens with a [pruned: <tool> result, N chars] marker. It edits the
run’s transcript in place, so pruning persists in checkpoints and
result.messages. twoPhaseStrategy({ model }) (recommended) prunes first
and, if the run is still too big, replaces old turns with a summary written
by model; summarizeStrategy() only summarizes. pinMessage(message) marks
a message that is never pruned or summarized. compactMessages(messages, options)
does the same once, by hand (async), and CompactionStrategy is the
pluggable interface (compact() may be async). createAgent({ compaction: true })
installs the hook on an agent (createAgent({ hooks }) takes any other
AgentHooks), and stream() emits compaction.start / compaction.done.
See Context compaction.
Flows, evals, observability and security
-
FlowBuilder/FlowExecutor- multi-step workflow graphs; see Flows. -
defineEval(), scorers such asexactMatch,toolCallOrderandbudget, checks such asincludesandatLeast, andllmJudge()- agent evals run under vitest orlousho eval; see Evals. -
withSpan()andTraceExporter- tracing forAgentExecutor.execute().TraceExporteris a bring-your-own-exporter interface (no exporter ships by default); for real OpenTelemetry spans, importcreateOtelTraceExporter()from the@lousho/build-ai-agent/otelsubpath (requires the optional peer dependency@opentelemetry/api) instead of hand-rolling the OTel bridge - seeexamples/tracing/run-otel.ts. Spans follow the OpenTelemetry GenAI semantic conventions (flows are traced too); see observability. -
NoopSandbox/SubprocessSandbox- sandboxing for tools that opt in viarequiresSandbox;runGuardrails()and guardrails such ascreateCommandGuardrail(),createDiffSizeGuardrail()andsecretScanGuardrail. See Guardrails and sandboxing. -
HookRegistry,AgentHook,HookContext,ToolCallHookContext,GenerateHookContext- pre/post agent hooks (run before/after a tool call or an LLMgenerate, can mutate args/messages/results or throw to abort the step; tool-call hooks can also deny a call, replace its result or modify its input, see Hook outcomes). Available from the package root and from@lousho/build-ai-agent/hooks. Agent Forge’s canvas hook editor (docs/agent-forge.md) compiles the hooks a user attaches to a node into aHookRegistrythis way, run sandboxed viaSandboxAdapterrather than in the host process. -
EncryptionUtils,sha256,StorageService- supporting utilities; see Utilities.
Hook outcomes
A tool-call hook can change what happens, not just watch it. ApreToolCall
hook may return:
- nothing: the call goes on unchanged (mutating
ctx.argsin place still works). { deny: reason }: the call does not run. The model gets the samekind: 'denied'tool error aneedsApprovaldeny produces (see Approvals), streams seetool.error, andonPermissionDecisionrecordsdecision: 'deny'withhookandreason.{ result: value }: the call does not run andvalueis its result. Thetool.doneevent and the transcript’stoolmessage (metadata) carryreplacedByHook: '<hook name>'.{ input: args }: the call runs withargs. They are validated against the tool’s input schema again; a mismatch becomes akind: 'validation'tool error whose message names the hook, and the tool does not run.
postToolCall hook may return { result: value } to replace the result the
model sees (to redact or truncate it); nothing keeps it. Hooks run in
registration order: the first deny or result of a pre-hook skips the
pre-hooks after it, inputs chain (each hook sees the previous one’s as
ctx.args), and each post-hook sees the result an earlier one replaced. A hook
that throws still rejects the run, as before. Sub-agents inherit the lead
agent’s hooks and their outcomes.
For one tool call, the order is: argument validation, preToolCall hooks,
permission rules (permissions), tool guardrails, the tool’s needsApproval,
the approval pause, execution, postToolCall hooks. Hooks run before every
approval decision, so rules, needsApproval and the human all see (and approve)
the input a hook produced; the pending approval’s args show it. When an
approved call is resumed, the pre-hooks run again: they may still deny it or
supply its result, but an { input } that differs from the approved input is
refused with a tool error instead of running.
Flow expressions
oneOf branch conditions (and the Agent Forge router node’s branch conditions)
and evaluator node expressions are evaluated by a small built-in expression
evaluator. It never compiles or runs host code: there is no eval,
new Function or vm in src/flows. {{name}} placeholders are bound as
values, never pasted into the expression text: a bare {{score}} >= 90 uses the
variable’s value, and inside a string literal, '{{classify}}' === 'refund'
interpolates the value’s text into that literal after the expression has been
tokenized. A variable whose value contains quotes, backslashes or operators
(for example x' === 'x' || 'a) is therefore just data and cannot change the
condition’s logic. A missing or null variable is an empty string inside a
literal; used bare it contributes nothing, so {{missing}} >= 90 is a syntax
error and the condition counts as not matched. The expression is evaluated
against the flow’s variables.
Precedence, loosest to tightest:
||, &&, equality, relational, + -,
* / %, unary, member access.
Not supported, and rejected with an ExpressionError that names the
expression, the character position and this list of supported forms: any other
function or method call, assignment (=, +=, ++), ternaries, template
strings, object/array literals, access to constructor, __proto__ or
prototype, and globals (process, require, globalThis, …). Only a
flow’s own variables, and only their own properties, are reachable.
Failure behaviour is unchanged: a oneOf condition that cannot be evaluated
counts as not matched (false), and an evaluator expression that cannot be
evaluated fails the flow with Failed to evaluate expression: ..., including
the ExpressionError detail.
Triggers
Trigger adapters (@lousho/build-ai-agent/triggers) wake an agent up from an
inbound webhook, a schedule or a Slack message. Wire any of them with
listen(agent, onEvent), where onEvent runs the agent.
For a surface people talk to, use a channel instead:
defineChannel() and mountChannels() (package root) map each conversation
on the surface to a session and send the reply, and any approval or question
the agent pauses on, back to it. httpChannel() and webhookChannel() are
built in; WebhookTriggerAdapter uses webhookChannel() for its auth.
Webhook authentication
WebhookTriggerAdapter starts an HTTP server. Always set auth for a
webhook that is reachable from outside your machine: without it, anyone who
can reach the port can run your agent (and spend your tokens). If you listen
on a non-loopback host with no auth, the adapter logs a one-time warning
through options.logger.
header (default x-signature-256), algorithm (sha256 or
sha1, default sha256), prefix (default sha256=; '' for a bare
digest), timestampHeader and toleranceSeconds. Signatures and bearer
tokens are compared in constant time. A request that fails authentication gets
a generic 401 {"error":"Unauthorized"} - the response never says which check
failed - and the reason (never a secret or signature) is logged at warn
level. Serve webhooks over HTTPS (terminate TLS in front of the adapter) so
tokens and payloads are not sent in clear text.
Slack request signatures
SlackTriggerAdapter.handleRequest({ headers, rawBody }) handles a raw Slack
Events API request and returns the { status, body } to send back. Set
signingSecret for any endpoint reachable from outside your machine: without
it, anyone who can reach the endpoint can run your agent, and listen() logs a
one-time warning through options.logger. With it, every request is verified
as Slack documents
before the body is parsed: HMAC-SHA256 over v0:{X-Slack-Request-Timestamp}:{raw body},
compared in constant time with X-Slack-Signature (v0=<hex>), and requests
more than five minutes old are rejected. Failures get a generic
401 {"error":"Unauthorized"}; the reason (never a secret or signature) is
logged at warn level. The signed url_verification handshake is answered
after verification.
handleRequest answers Slack after the agent finishes; Slack expects a reply
within three seconds, so for slow agents acknowledge first and run the agent in
the background.
Cron schedules
CronTriggerAdapter takes either a fixed intervalMs or a real cron
expression:
*, lists (1,15), ranges (1-5), steps (*/15,
10-40/10), month names (JAN) and weekday names (MON), with 0 and 7
both meaning Sunday, plus @hourly, @daily, @weekly and @monthly. As in
classic cron, when both day-of-month and day-of-week are restricted a day
matches if either does. An invalid expression throws a CronExpressionError
naming the field and showing a valid example. parseCronExpression(expr, timezone).nextRun(after) is exported if you need the next fire time.
Around daylight-saving changes, a time that does not exist (spring forward) is
skipped for that day, and a time that happens twice (fall back) fires once;
an every-hour schedule keeps firing hourly. The timer is re-armed after each
run from the scheduled time (no drift, no double fire), and stop() clears it.
Deployment
DeploymentAdapter,registerAdapter(),getAdapter(),listAdapters()- the adapter registry behindlousho build(see Deployment).