Skip to main content
A long agent run fills its context window mostly with tool results: file contents, search hits, API responses the model read once and no longer needs in full. Compaction keeps such a run under the model’s limit in two phases: first it replaces old tool results with a short marker, and if that is not enough it replaces the oldest turns with a summary written by a (cheaper) model.

Compacting an agent

With createAgent(), turn compaction on with the compaction option. true installs the hook with its defaults (prune old tool results above 90% of the model’s window); an object configures it, and summarizer selects twoPhaseStrategy() with that model (a "provider/model" string or an LLMProvider):
The object takes strategy, thresholdPercent, contextWindow and protectedTokens (as in the table below) and summarizer; combining summarizer with strategy is a configuration error (pass twoPhaseStrategy({ model }) as the strategy instead). createAgent({ hooks }) takes any other AgentHooks, which run before the compaction hook, so compaction sees what they added. Hooks and compaction apply to send(), stream() and sessions.

Stream events

agent.stream(), session.stream() and AgentExecutor.stream() report each compaction as two events, inside the step and before the model call that triggered it (see Streaming): Every compaction.start is followed by exactly one compaction.done. When the strategy could not shrink anything, tokensAfter equals tokensBefore; when it failed (or the summarizer failed and the hook fell back to pruning), error is set and the run continues. summary is true when old turns were replaced by a summary (the text itself is not sent; use onCompaction for it). A non-streaming send() emits no events. Hooks add their own events with ctx.emit?.(...) on the preGenerate context, which exists only in streamed runs.

Compacting a run

createCompactionHook() returns an AgentHook named compaction. Register it and pass the registry to AgentExecutor.execute(). The recommended setup is the two-phase strategy with a cheap summarizer model:
Before every model call the hook estimates the request’s size with estimateTokens (see Models, tokens and cost). When it is above thresholdPercent of the context window, the hook runs its strategy (and waits for it, if it is async) and calls onCompaction with the token counts, the pruned toolCallIds, the summary if there is one and the strategy name. Compaction never fails a run. If the strategy throws, or the summarizer call fails, the run continues (with the pruned conversation, or unchanged) and onCompaction receives the error. The hook rewrites the run’s transcript itself, not a copy: the request’s messages array is the run’s message list, and the hook changes it in place. A pruned result or a summary therefore stays in later steps, in checkpoints (a resumed session does not bring the full history back) and in result.messages, and it is not compacted again on the next step. Compaction is lossy: if the model needs a pruned result again, it has to call the tool again.

Strategies

Pruning tool results

pruneToolResultsStrategy() replaces the content of each tool result older than the protected tail with a marker such as [pruned: search result, 18234 chars]. It never changes:
  • system messages, user messages (the first one included) or assistant turns, so every tool call keeps its result message and the transcript stays valid for every provider;
  • the results of the latest assistant turn, which the model has not read yet, even when they are bigger than protectedTokens;
  • pinned results, results that are already a marker, or results shorter than a marker.

Summarizing old turns

summarizeStrategy() sends the messages before the protected tail to the summarizer with a focused instruction and replaces them with one user message:
The summary keeps the transcript valid for every provider:
  • system messages and pinned messages stay where they are, and the summary goes where the first summarized message was;
  • an assistant turn is summarized or kept together with all its tool results, so no tool call loses its result (and no result loses its call);
  • the last message and the latest assistant turn’s results are always kept.
If the summary call fails (or returns nothing), the strategy falls back to pruning and returns the failure as error.

Two phases

twoPhaseStrategy() takes the same options. It prunes tool results first, which is free, and calls the summarizer only when the pruned conversation is still above thresholdPercent of the context window. The summarizer then reads the pruned conversation. If it fails, the pruned conversation is kept and error is reported.

Pinning messages

pinMessage(message) returns a copy of the message marked as pinned (it sets metadata.pinned: true; Message.metadata is application data that is never sent to the model). The built-in strategies never prune or summarize a pinned message; a pinned tool result keeps its whole assistant turn. Pin what the agent must always see in full, such as the task statement or a key document:
Custom strategies should honor isPinned() too.

Compacting by hand

compactMessages(messages, options) runs a strategy once, whatever the size of the conversation, and resolves to a new array (its input is not modified). Use it to shrink a stored transcript before you continue it:

Compacting a session

session.compact(options?) compacts a session’s transcript now, whatever its size, saves it to the session’s store (memory, file, SQLite or KV) and resolves to { messagesBefore, messagesAfter, tokensBefore, tokensAfter, strategy, error? }. It runs options.strategy, else the strategy of agent.session({ compaction }) (the same value as createAgent({ compaction }), and the agent’s own compaction when the session does not set one), else prunes old tool results. Pinned messages stay. A session with no messages resolves without doing anything. It rejects with LOUSHO_SESSION_BUSY while a turn is running, and with LOUSHO_SESSION_TURN_PENDING / LOUSHO_SESSION_AWAITING_APPROVAL while a durable turn is unfinished. Listeners added with session.on() get compaction.start and compaction.done with trigger: 'manual'. session.clear() empties the transcript instead and emits context.cleared.

Writing a strategy

A strategy is a name and a compact() function, which may be async. It receives the messages, a token counter for the request’s model, the context window, protectedTokens, thresholdTokens (the size to get under) and the run’s abort signal, and returns the new messages with before/after token counts (and optionally a summary or an error). Return the input array unchanged when there is nothing to do: the hook then leaves the run alone and does not call onCompaction.
A custom strategy gets the same compaction.start / compaction.done events as the built-in ones. onCompaction (on createCompactionHook()) still receives the summary text and the Error.