Skip to main content
How the pieces fit:

Building and running agents

Dynamic config

createAgent()’s model, instructions (or prompt) and tools each take the static value or a function of the run, (ctx) => value | Promise<value> (type PerRun<T>). ctx is { sessionId?, input, metadata? } (type RunConfigContext): the sessionId of send() / stream() or the agent.session() id, the run’s user input, and the metadata call option of send(), stream(), session.send() and session.stream(). The functions run once when a run starts, before the first model call, and again on every session turn. Everything else applies to what they return: fallbackModels and retry, projectInstructions, memory, skills, sub-agents, MCP tools, permissions, approvals and guardrails. A dynamic agent used as a sub-agent resolves with the task prompt as input.
A function that throws fails the run with LOUSHO_CONFIG_RESOLVER_FAILED (error.field names the option, error.cause is the thrown error): send() rejects, stream() ends with an error event and a session keeps its transcript as it was. A run paused for an approval or a question keeps the ctx and the model it resolved in the paused snapshot: resuming it, even from another process, uses that model (it never switches mid-turn) and calls the tools function again with the same ctx. A run resumed from a crash checkpoint (agent.resume(id)) resolves again with input: [] and no metadata. With only static values nothing changes: the agent is built once, when it is created.

UI bindings

@lousho/build-ai-agent/react exports useLoushoAgent(source, options?), a React hook that runs an agent in process ({ agent, sessionId? }) or over HTTP ({ url }) and returns messages, status, pendingApproval, send(), stop(), approve() and reject(). Its framework-neutral parts, reduceAgentEvents() and parseEventStream(), are exported too. See React. @lousho/build-ai-agent/vue exports the same useLoushoAgent() as a Vue 3 composable, with the state as refs. See Vue. @lousho/build-ai-agent/svelte exports loushoAgent(), the same state and actions as a Svelte store ($agent). See Svelte.

Sub-agents

Pass subagents: { researcher, writer } (agents from createAgent() with a description, or a { list, resolve } catalog) to createAgent() or AgentExecutor.execute(): the lead gets one task tool and a prompt listing, each sub-agent runs on the task prompt alone, and it inherits the lead run’s signal, hooks (ctx.subagent), tracing, approval store and event listeners (event.subagent). maxSubagentDepth (default 1) bounds nesting. See Sub-agents.

Approvals

A createAgent() agent pauses on a needsApproval tool instead of failing: send() (and session.send()) resolves with finishReason: 'awaiting-approval' and an approvalId. agent.approvals.list() returns the pending calls and agent.approvals.resolve({ id, approved, note? }) runs or rejects the call and resolves with the continued run’s result (continuing the session it paused in). Pauses are kept in a per-agent InMemoryApprovalStore unless you pass approvalStore (e.g. SqliteStore.approvals) or a store; approve: (call) => boolean | string decides each call in code without pausing (stream() still ends at the pause). With askQuestion: true the agent can ask the user a question (kind: 'question'), answered with agent.approvals.answer({ id, answer }). See Approvals.

Structured output

Pass output: zodSchema to createAgent() (or AgentExecutor.execute() / stream()): the final reply must be a JSON object matching the schema, and send() / run.result resolve with it validated as result.object, typed z.output<typeof schema> (result.text keeps the raw JSON). Tools still run first. Each model call carries a responseFormat: { type: 'json', schema } hint, which the ai-SDK providers map to JSON mode. An invalid reply gets one repair step listing the issues (it counts against maxSteps); if that is invalid too, the run ends with finishReason: 'output-invalid' and outputError: { message, issues }. See Structured output.

Skills

Pass skills: [defineSkill({ name, description, content }), ...(await loadSkills(dir))] to createAgent() or AgentExecutor.execute(): only names and descriptions go in the system prompt and the model loads bodies through an auto-registered load_skill tool. See Skills.

Memory

Pass memory: [defineMemory({ name, scope, provider })] to createAgent() to keep items across conversations: each run recalls a slot’s newest items into the system prompt on its first model call, and the model gets remember_<name> / recall_<name> tools. scope is 'global', 'session' or a function of { sessionId, metadata }; inMemoryMemory() and fileMemory({ dir }) are the built-in providers. See Memory.

Durable execution

sessionId + checkpointStore make a run crash-safe and a session multi-turn: the run is checkpointed after every model response, every tool result and every pause, and calling execute() again with the same sessionId resumes an unfinished run (without re-calling the model for a turn it already has), continues a finished conversation with the new input, or throws SessionAwaitingApprovalError while an approval is pending. Tools run at-least-once across a crash; execute receives the call’s toolCallId to use as an idempotency key. See Durable execution for the exact guarantees.

Cancellation

Pass an AbortSignal to stop a run: agent.send(input, { signal }), AgentExecutor.execute({ ..., signal }) or resumeAfterApproval(..., { signal }).
For a time limit, use AbortSignal.timeout(ms):
How it behaves:
  • The signal is checked before every model call and every tool call. It is passed to the provider (GenerateOptions.signal, sent to the ai SDK as abortSignal) and to each tool as execute(args, { abortSignal }), so in-flight work can stop early. The built-in httpTool passes it to fetch, and agents created with createDelegateTool() are aborted along with their parent.
  • An aborted run resolves (it does not reject) with finishReason: 'aborted' and the messages and steps so far. A rejection caused by the abort, such as an AbortError, is not treated as a failure: it is not retried by the provider retry and fallback wrappers and is not compacted into a provider error.
  • The run’s events end with run.done with finishReason: 'aborted'.
  • With sessionId + checkpointStore, the state is checkpointed. Calling execute() again with the same sessionId resumes where the run stopped; new input is appended as the next user message (see Durable execution). Tool calls the run never reached get an { error } result saying they were cancelled, so the conversation stays valid for the provider.
  • An already-aborted signal returns at once without calling the provider.

Finish reasons

result.finishReason says why a run ended: the model’s own reason for its last turn ('stop', 'length', 'tool_calls', 'content_filter', 'error'), 'awaiting-approval' (paused on a tool call that needs a human), 'aborted' (cancelled with signal), 'max-steps', 'output-invalid' (the reply did not match the output schema even after the repair step, see Structured output), 'budget-exceeded' (a limits budget such as maxTokens or maxCostUsd tripped; result.budget says which, see Budgets), or 'guardrail' (an input, output or tool guardrail blocked; result.guardrail says which, see Input and output guardrails). 'max-steps' means the maxSteps budget (default 10) ran out while the model still wanted to continue, so the reply may be empty or partial; a run that finishes naturally within the budget keeps its 'stop'. Steps carried over by initialSteps or an approval resume count against the budget, and result.steps is the number of steps taken. The same reason is on the finish event and on run.done in streaming.

Parallel tool calls

When the model asks for several tools in one turn, they run concurrently. toolConcurrency (on createAgent() and AgentExecutor.execute()) caps how many run at once: a positive integer, or 'unbounded' (the default). Use 1 for strictly sequential execution, e.g. when your tools share state that is not safe to touch concurrently.
Guarantees, whatever the limit:
  • Transcript order is call order. Tool results are appended in the order the model requested the calls, not the order they finish, so the next provider request is deterministic.
  • Events. Calls start in call order. A call’s tool-call event, onToolCall, argument validation, preToolCall hooks and needsApproval check run just before it starts, one call at a time. Its tool-result event fires when it finishes, so results arrive in completion order. With toolConcurrency: 1 events alternate call/result exactly as before.
  • Approvals. The first call that needs approval stops the batch: the calls before it run (concurrently) and their results are recorded, then the run pauses on that call (finishReason: 'awaiting-approval'). Calls after it never start in this run. resumeAfterApproval() records the paused call’s result (or rejection) and then runs those later calls the same way, so every call of the turn gets exactly one result - see Durable execution.
  • Failures are isolated. A tool that throws gets its own error result; its siblings carry on. A propagating error (PropagatingToolError, such as the delegation depth guard, or a throwing hook) stops new calls from starting, waits for the running ones to settle, then rejects the run. No tool is left running detached.
  • Cancellation. An abort while a batch runs resolves with finishReason: 'aborted'. Calls that finished keep their results; the rest get a “cancelled” result. Running tools see the abort through their abortSignal.
  • Checkpoints. With sessionId + checkpointStore, the model’s turn is checkpointed before any call starts, then again each time the in-order run of finished calls grows (with 1, after every call). A resumed run never runs a recorded call again, and never asks the model again for a turn it already checkpointed.

Declarative specs

Providers

Message.content is a string or a list of ContentParts (text, image, file); the built-in providers send image parts on user messages. See Multimodal input. See Providers for how a model string is resolved and which model runs.

Testing

Exported from @lousho/build-ai-agent/testing (see Testing agents).

Tools

See Tools for a guide to defining and registering tools.
Optional fields: displayName, needsApproval (boolean or predicate), requiresSandbox and sandboxExecute. Defining a tool validates its name, description and zod input immediately; registering two tools with the same name throws an error naming the conflict. A defined tool carries its schema as inputSchema (the same zod schema as input) and its execute function directly. These are the canonical fields of a ToolDescriptor; the .tool object (an ai v4 { description, parameters, execute }) is legacy, still built for compatibility, and used only for descriptors that do not set inputSchema / execute.
Schema support. Each tool’s JSON Schema inputSchema is converted to a zod schema, and the model’s arguments are validated against it before the server is called. The converter handles type (including arrays such as ["string", "null"]), enum with mixed types, const, anyOf / oneOf (a union; anyOf: [X, { type: "null" }] becomes X.nullable()), allOf (merge / intersection), local $ref into $defs / definitions, properties / required, additionalProperties (boolean or schema), items, default, and the constraints minimum, maximum, exclusiveMinimum, exclusiveMaximum, minLength, maxLength, pattern, minItems and maxItems. Where JSON Schema is ambiguous the conversion is permissive: unknown keywords, empty schemas, tuple items and unresolvable refs become z.any(), a recursive $ref is expanded once and z.any() is used for the inner occurrence, an invalid pattern regex is skipped, and objects keep extra properties unless additionalProperties is false. The converter never throws on schema content. Skipped tools. If one tool still cannot be converted, only that tool is skipped; the rest of the server’s tools load. Pass a logger to receive a warning naming the server, the tool and the reason, and onSkip to collect what was left out:
Results. An MCP result with isError: true is a normal tool failure: the model receives { "error": "McpToolError", "toolName": "...", "message": "..." } where message is the server’s text content. Successful results are JSON-serializable: if the server returns structuredContent it is the result object; otherwise the result is { text, content }, where text joins all text parts and content keeps every part in order (text, image, audio, resource, resource_link; an unrecognised part type is kept as { type: 'unknown', raw }). When structuredContent arrives together with non-text parts the result is { structuredContent, text, content } with content holding only the non-text parts, so images, audio and resources are never dropped.

Todo tools

createTodoTools(options?) gives long-running agents a plan to track: todo_write replaces the whole list ({ id?, content, status } items, status pending | in_progress | completed, at most one in_progress) and returns the list plus counts; todo_read returns it. Ids are assigned automatically and stay stable when a later write repeats an item’s content. Invalid lists (for example two in_progress) reach the model as a structured tool error so it can retry. The list lives in memory per call; pass store ({ get, set }) to persist it, and onChange to update a UI.

Tool errors

A tool call can fail in two ways. In both, the run continues and the model receives a structured JSON error as the tool result (never the string null), so it can recover. tool-result events, tracing, onToolResult and postToolCall hooks see the call as an error. Argument validation. Before a tool runs, the model’s arguments are parsed with the tool’s zod inputSchema (for a legacy descriptor, tool.parameters). Validation happens first, so pre-tool hooks, the needsApproval predicate and execute all receive the parsed value (defaults, coercions and transforms applied). Tools without a zod schema are passed through unchanged. If the arguments do not match, execute is not called, the run continues, and the model receives a structured error as the tool result so it can retry (tool-result events, tracing and postToolCall hooks see it as an error result; preToolCall hooks are skipped because there is no valid call):
ToolArgumentsValidationError (with a typed issues array) is exported from the package root. Thrown errors. If execute throws, the model receives the error name, the tool name and the message only (never a stack trace). Messages are capped at 2,000 characters and end with ... (truncated) when cut:
Every other failure (unknown tool, rejected approval, a call that was not run, an MCP error, a refused sandboxed tool) uses the same { error, toolName, message, kind } shape; see Errors for the kind values. Errors extending PropagatingToolError (for example the delegation depth guard) are the exception: they are rethrown and abort the run instead of being shown to the model.

Models, tokens and cost

Dependency-free helpers for budgeting and context decisions.
  • estimateTokens(input, { model?, estimator? }) is a heuristic (about 4 characters per token for English, more for CJK and other scripts, plus per-message overhead and tool-call JSON). Expect roughly 15-20% error on English: fine for compaction and budgets, not for billing. Plug in a real tokenizer with setTokenEstimator(fn) or options.estimator.
  • getModelInfo(id) matches the exact id, then provider/id, then dated snapshots (gpt-4o-mini-2024-07-18 resolves to gpt-4o-mini). Unknown models return undefined.
  • registerModel(info) adds or overrides an entry; the latest registration wins.
  • estimateCost(usage, model) returns USD, or undefined (not 0) when the model or its prices are unknown.
The built-in context windows and prices are a dated snapshot (see the retrieval date and sources at the top of src/models/modelData.ts). Providers change prices and models, so override entries with registerModel when you need billing-grade numbers.

Usage and cost of a run

Every ExecutionResult (and agent.send() result) carries usage, the running total of the whole run:
  • Reported vs estimated. Providers report promptTokens/completionTokens; the built-in providers also pass on cachedInputTokens/reasoningTokens when the ‘ai’ SDK’s provider metadata carries them (usage.cachedInputTokens and usage.reasoningTokens only appear then). A backend that reports nothing gives undefined usage, never zeros. For that step the executor falls back to estimateTokens and sets usage.estimated (a leading ~ in formatUsage). A custom provider should leave GenerateResult.usage unset rather than fill in zeros.
  • Cost. costUsd is the sum of estimateCost per model. It is undefined, never a misleading partial sum, as soon as any model used has unknown pricing; byModel shows which ones are priced.
  • Delegation. A delegated child’s usage is added to the parent’s totals and byModel, and is also shown on its own as usage.delegated ({ inputTokens, outputTokens, totalTokens, costUsd, modelCalls, estimated, runs }).
  • Resume. A run resumed from a checkpoint, or after an approval, continues from the saved totals instead of restarting at zero. Checkpoints written by older versions start from their saved token counts with an unknown cost.
  • Events and traces. The finish event (and every lifecycle event that carried usage) now carries the running totals; text-complete also has stepUsage, and onLLMResponse receives the call’s usage as a third argument. The chat span’s gen_ai.usage.* attributes use the same numbers, with lousho.usage.estimated set to true when they are estimates.
  • Streaming. agent.stream() events carry the same accounting: step.done and run.done have usage with inputTokens, outputTokens, estimated and costUsd (run-level also modelCalls), alongside the older promptTokens/completionTokens.
  • promptTokens and completionTokens on usage remain as deprecated aliases of inputTokens and outputTokens.
Prices come from the model registry above, so to get a cost for a custom or fine-tuned model, register it under the id you pass as the model:

Context compaction

createCompactionHook({ thresholdPercent?, contextWindow?, protectedTokens?, strategy?, onCompaction? }) returns an AgentHook that, before each model call above 90% (by default) of the model’s context window, replaces tool results older than the newest 40,000 tokens with a [pruned: <tool> result, N chars] marker. It edits the run’s transcript in place, so pruning persists in checkpoints and result.messages. twoPhaseStrategy({ model }) (recommended) prunes first and, if the run is still too big, replaces old turns with a summary written by model; summarizeStrategy() only summarizes. pinMessage(message) marks a message that is never pruned or summarized. compactMessages(messages, options) does the same once, by hand (async), and CompactionStrategy is the pluggable interface (compact() may be async). createAgent({ compaction: true }) installs the hook on an agent (createAgent({ hooks }) takes any other AgentHooks), and stream() emits compaction.start / compaction.done. See Context compaction.

Flows, evals, observability and security

  • FlowBuilder / FlowExecutor - multi-step workflow graphs; see Flows.
  • defineEval(), scorers such as exactMatch, toolCallOrder and budget, checks such as includes and atLeast, and llmJudge() - agent evals run under vitest or lousho eval; see Evals.
  • withSpan() and TraceExporter - tracing for AgentExecutor.execute(). TraceExporter is a bring-your-own-exporter interface (no exporter ships by default); for real OpenTelemetry spans, import createOtelTraceExporter() from the @lousho/build-ai-agent/otel subpath (requires the optional peer dependency @opentelemetry/api) instead of hand-rolling the OTel bridge - see examples/tracing/run-otel.ts. Spans follow the OpenTelemetry GenAI semantic conventions (flows are traced too); see observability.
  • NoopSandbox / SubprocessSandbox - sandboxing for tools that opt in via requiresSandbox; runGuardrails() and guardrails such as createCommandGuardrail(), createDiffSizeGuardrail() and secretScanGuardrail. See Guardrails and sandboxing.
  • HookRegistry, AgentHook, HookContext, ToolCallHookContext, GenerateHookContext - pre/post agent hooks (run before/after a tool call or an LLM generate, can mutate args/messages/results or throw to abort the step; tool-call hooks can also deny a call, replace its result or modify its input, see Hook outcomes). Available from the package root and from @lousho/build-ai-agent/hooks. Agent Forge’s canvas hook editor (docs/agent-forge.md) compiles the hooks a user attaches to a node into a HookRegistry this way, run sandboxed via SandboxAdapter rather than in the host process.
  • EncryptionUtils, sha256, StorageService - supporting utilities; see Utilities.

Hook outcomes

A tool-call hook can change what happens, not just watch it. A preToolCall hook may return:
  • nothing: the call goes on unchanged (mutating ctx.args in place still works).
  • { deny: reason }: the call does not run. The model gets the same kind: 'denied' tool error a needsApproval deny produces (see Approvals), streams see tool.error, and onPermissionDecision records decision: 'deny' with hook and reason.
  • { result: value }: the call does not run and value is its result. The tool.done event and the transcript’s tool message (metadata) carry replacedByHook: '<hook name>'.
  • { input: args }: the call runs with args. They are validated against the tool’s input schema again; a mismatch becomes a kind: 'validation' tool error whose message names the hook, and the tool does not run.
A postToolCall hook may return { result: value } to replace the result the model sees (to redact or truncate it); nothing keeps it. Hooks run in registration order: the first deny or result of a pre-hook skips the pre-hooks after it, inputs chain (each hook sees the previous one’s as ctx.args), and each post-hook sees the result an earlier one replaced. A hook that throws still rejects the run, as before. Sub-agents inherit the lead agent’s hooks and their outcomes. For one tool call, the order is: argument validation, preToolCall hooks, permission rules (permissions), tool guardrails, the tool’s needsApproval, the approval pause, execution, postToolCall hooks. Hooks run before every approval decision, so rules, needsApproval and the human all see (and approve) the input a hook produced; the pending approval’s args show it. When an approved call is resumed, the pre-hooks run again: they may still deny it or supply its result, but an { input } that differs from the approved input is refused with a tool error instead of running.

Flow expressions

oneOf branch conditions (and the Agent Forge router node’s branch conditions) and evaluator node expressions are evaluated by a small built-in expression evaluator. It never compiles or runs host code: there is no eval, new Function or vm in src/flows. {{name}} placeholders are bound as values, never pasted into the expression text: a bare {{score}} >= 90 uses the variable’s value, and inside a string literal, '{{classify}}' === 'refund' interpolates the value’s text into that literal after the expression has been tokenized. A variable whose value contains quotes, backslashes or operators (for example x' === 'x' || 'a) is therefore just data and cannot change the condition’s logic. A missing or null variable is an empty string inside a literal; used bare it contributes nothing, so {{missing}} >= 90 is a syntax error and the condition counts as not matched. The expression is evaluated against the flow’s variables. Precedence, loosest to tightest: ||, &&, equality, relational, + -, * / %, unary, member access. Not supported, and rejected with an ExpressionError that names the expression, the character position and this list of supported forms: any other function or method call, assignment (=, +=, ++), ternaries, template strings, object/array literals, access to constructor, __proto__ or prototype, and globals (process, require, globalThis, …). Only a flow’s own variables, and only their own properties, are reachable. Failure behaviour is unchanged: a oneOf condition that cannot be evaluated counts as not matched (false), and an evaluator expression that cannot be evaluated fails the flow with Failed to evaluate expression: ..., including the ExpressionError detail.

Triggers

Trigger adapters (@lousho/build-ai-agent/triggers) wake an agent up from an inbound webhook, a schedule or a Slack message. Wire any of them with listen(agent, onEvent), where onEvent runs the agent. For a surface people talk to, use a channel instead: defineChannel() and mountChannels() (package root) map each conversation on the surface to a session and send the reply, and any approval or question the agent pauses on, back to it. httpChannel() and webhookChannel() are built in; WebhookTriggerAdapter uses webhookChannel() for its auth.

Webhook authentication

WebhookTriggerAdapter starts an HTTP server. Always set auth for a webhook that is reachable from outside your machine: without it, anyone who can reach the port can run your agent (and spend your tokens). If you listen on a non-loopback host with no auth, the adapter logs a one-time warning through options.logger.
HMAC options: header (default x-signature-256), algorithm (sha256 or sha1, default sha256), prefix (default sha256=; '' for a bare digest), timestampHeader and toleranceSeconds. Signatures and bearer tokens are compared in constant time. A request that fails authentication gets a generic 401 {"error":"Unauthorized"} - the response never says which check failed - and the reason (never a secret or signature) is logged at warn level. Serve webhooks over HTTPS (terminate TLS in front of the adapter) so tokens and payloads are not sent in clear text.

Slack request signatures

SlackTriggerAdapter.handleRequest({ headers, rawBody }) handles a raw Slack Events API request and returns the { status, body } to send back. Set signingSecret for any endpoint reachable from outside your machine: without it, anyone who can reach the endpoint can run your agent, and listen() logs a one-time warning through options.logger. With it, every request is verified as Slack documents before the body is parsed: HMAC-SHA256 over v0:{X-Slack-Request-Timestamp}:{raw body}, compared in constant time with X-Slack-Signature (v0=<hex>), and requests more than five minutes old are rejected. Failures get a generic 401 {"error":"Unauthorized"}; the reason (never a secret or signature) is logged at warn level. The signed url_verification handshake is answered after verification.
handleRequest answers Slack after the agent finishes; Slack expects a reply within three seconds, so for slow agents acknowledge first and run the agent in the background.

Cron schedules

CronTriggerAdapter takes either a fixed intervalMs or a real cron expression:
Supported syntax: *, lists (1,15), ranges (1-5), steps (*/15, 10-40/10), month names (JAN) and weekday names (MON), with 0 and 7 both meaning Sunday, plus @hourly, @daily, @weekly and @monthly. As in classic cron, when both day-of-month and day-of-week are restricted a day matches if either does. An invalid expression throws a CronExpressionError naming the field and showing a valid example. parseCronExpression(expr, timezone).nextRun(after) is exported if you need the next fire time. Around daylight-saving changes, a time that does not exist (spring forward) is skipped for that day, and a time that happens twice (fall back) fires once; an every-hour schedule keeps firing hourly. The timer is re-armed after each run from the scheduled time (no drift, no double fire), and stop() clears it.

Deployment

  • DeploymentAdapter, registerAdapter(), getAdapter(), listAdapters() - the adapter registry behind lousho build (see Deployment).