Skip to main content
Pass a sessionId and a checkpointStore to AgentExecutor.execute() and the run survives a crash, a restart, an abort or a human approval that takes days. Nothing here depends on a particular host: any CheckpointStore (LocalStorageCheckpointStore on Node, the Cloudflare KV store the cloudflare-worker build wires up, or your own three-method implementation) and any ApprovalStore work the same way.
With createAgent(), one store option does the wiring: pass a sessionId to send() (or stream()) and the run is checkpointed in store.checkpoints; after a crash, agent.resume(sessionId) finishes it (and returns null when nothing is pending). No AgentExecutor needed:
store.approvals holds approval pauses, and agent.approvals.resolve() keeps checkpointing the run under its sessionId. Sessions are checkpointed too: agent.session({ id }) checkpoints every turn and agent.resume(id) (or session.resume()) finishes an interrupted one - see Durable sessions.

What is checkpointed, and when

A checkpoint is the run’s whole transcript plus its step count, usage, businessState and a status. It is written under the sessionId: A run that rejects (a provider error, a throwing hook, a PropagatingToolError, a process crash) keeps its last 'in-progress' checkpoint.

What happens when you call execute() again with the same sessionId

It depends on the stored status:
  • Unfinished run ('in-progress': it crashed, was aborted, or a store write failed). The run resumes. If its last model turn has tool calls without a result, exactly those calls run first, through the normal tool path (argument validation, hooks, needsApproval, toolConcurrency), without calling the model again. Then the loop continues, with the remaining maxSteps budget; usage continues from the checkpointed totals (the same holds for resumeAfterApproval(), from the snapshot).
    • New input is appended as a user message after those tool results. It is never placed between a tool-call turn and its results, so the transcript stays valid for every provider. This is what lets a user interrupt a run (abort it) and redirect it with a new message.
    • input that re-sends the message the run started from (the natural “retry the same request after a crash”) is treated as a retry and is not appended again. input: [] also just resumes.
  • Finished run ('finished'). The session continues as a conversation: the stored messages are kept and input becomes the next user turn. Pass only the new message(s); if you re-send the stored history followed by new messages, the history is recognised and not duplicated. steps, usage and toolCalls on the result count from zero for the new run; messages is the whole conversation.
  • Paused awaiting approval ('awaiting-approval'). execute() throws SessionAwaitingApprovalError (with sessionId and approvalId) and does not call the model or any tool: new input must not bypass a pending decision. Resolve it with resumeAfterApproval() first, passing the same checkpointStore, then send the new message.
System messages in input are ignored when it is added to a stored session (the session already has its system prompt). To end a session and start over under the same id, call checkpointStore.delete(sessionId).

Approvals in the middle of a tool batch

When one model turn asks for several tools and one of them needs approval, the calls before it run and are recorded, and the run pauses on it. The approval snapshot records the calls after it (remainingToolCalls). resumeAfterApproval() then:
  1. records the paused call’s result - the tool’s result if approved, or a structured { error, note } rejection result if rejected;
  2. runs the remaining calls through the same batch logic: they can run, fail validation, or pause the run again on another approval (a new approvalId; resolve it the same way, as many times as needed);
  3. calls the model once every call of the turn has exactly one result.
The resumed run uses the same approvalStore for any later approval unless you pass a different one in the options.

Guarantees

  • Every tool call gets exactly one result, in call order, after any sequence of crashes, aborts, pauses and resumes. (An aborted call gets a “cancelled” result; a rejected call gets a rejection result.)
  • A recorded result is never re-executed. Results are checkpointed as the in-order prefix of finished calls grows; a resumed run only runs calls that have no recorded result.
  • A checkpointed model turn is never generated again. If the process dies after the model answered, the resumed run executes that answer’s tool calls instead of asking the model again (no extra cost, no different plan).
  • The provider always receives a valid transcript: tool results directly follow their tool-call turn, and new user input comes after them.

At-least-once tools: make side effects idempotent

A tool that was running when the process died has no recorded result, so it runs again on resume. Tool execution is therefore at-least-once, not exactly-once. For tools with side effects (charging a card, sending an email), either gate them with needsApproval or make them idempotent. Each call’s toolCallId (the model’s id for the call, unchanged when it is re-run on resume) is passed to execute, so it works as an idempotency key:
(toolCallId is set for every tool run, including tools that route through sandboxExecute.)

Checkpoint history

A checkpoint record is overwritten on every save, so the run’s earlier steps are gone once it moves on. A store can also keep a bounded history per session: every save() appends the record to a ring (the newest historyLimit saves, default 50; the oldest are dropped), which is what replaying or forking a run from an earlier step is built on. Which stores keep one: For example, with the in-memory store:
  • history(sessionId, { limit? }) returns { step, savedAt, status, checkpoint } entries, newest first (step is checkpoint.stepIndex, savedAt an ISO timestamp, status is 'in-progress' for a checkpoint without one). The executor saves after every model turn and tool batch, so a session has one entry per save; the entries hold full checkpoints, so lower historyLimit for long transcripts. An unknown session gives [].
  • history is optional on CheckpointStore: a custom store without it keeps working, and getCheckpointHistory() returns undefined for it instead of throwing.
  • KVCheckpointStore keeps one index key per session (<prefix><sessionId>#history, the entry ids oldest first) and one key per entry (<prefix><sessionId>#history/<id>), so no value grows with the history; the latest checkpoint stays at <prefix><sessionId>, so a store written before history existed still loads. A save writes the checkpoint, then the entry, then the index, then deletes the entries it dropped: a crash between writes leaves an unlisted entry at worst (it expires with the TTL), never an unreadable store. KV has no transactions, so two concurrent saves of one session can lose one index update (that entry is missing from history()), and an entry can be missing for a while at another edge location (eventual consistency; history() skips it). Route a session’s requests to one location when that matters. Each save costs one get and two puts more; historyLimit: 0 turns it off.
  • Agent Forge’s file store writes the history as one file per session (.lousho/agents/<id>/checkpoint-history/<session>.json, replaced atomically); a file damaged by a crash reads as no history and is rebuilt by the next save.
  • SqliteStore keeps the history in a new checkpoint_history table, added when an existing database file is opened; prune() also removes history entries older than its cutoff.

Fork and replay

AgentExecutor.fork() starts a new session from a step of another session’s history: “what if the tool had returned something else at step 1?” It takes the newest history entry of fromStep, applies patch, saves the result as the 'in-progress' checkpoint of the new session and returns { sessionId, step, checkpoint }. Resume the fork like any unfinished run:
With createAgent({ store }), agent.fork(sessionId, { fromStep, patch? }) does the same over store.checkpoints, and agent.resume(fork.sessionId) continues the fork:
  • patch is applied in this order: messages(messages) rewrites the transcript; businessState replaces it; toolResult: { toolCallId, result } replaces that call’s result in place (same position, call id and tool name, so the transcript stays valid for every provider), or records it when the call has no result yet; appendInput queues a user message, which is sent after any tool calls still pending.
  • The fork keeps the source checkpoint’s step count and usage, so the rest of the maxSteps budget applies to it. Recorded tool results are not re-run; tool calls without a result run first, as on any resume (so a fork of an 'awaiting-approval' entry goes through needsApproval again; the old approvalId is not carried over).
  • newSessionId names the fork; the default is <sessionId>.fork-<n>, the first n with no checkpoint. A newSessionId that already has a checkpoint is refused.
  • A step the history does not have (the session is unknown, the step never ran, or it was dropped past historyLimit) throws an SDKError with code LOUSHO_CHECKPOINT_NOT_FOUND, whose message lists the steps that are kept. A store without history() cannot fork (ConfigurationError).
  • compareTrajectories(a, b) takes two checkpoints or two transcripts and returns each run’s steps (one per assistant turn: its text and tool calls with their results), divergedAt (the first step that differs, tool call ids aside) and drift, the tool-order, argument, step-count and finish-reason differences in the same shape as lousho eval --drift.

Resuming with a changed agent

Every checkpoint and approval snapshot carries an agent fingerprint: the model id, each tool’s name with a hash of its JSON input schema, and a hash of the system instructions (plus a short SHA-256 over all of them, computed with Web Crypto). Functions and other values that cannot be serialized (a tool’s execute, hooks, callbacks) are not part of it, so changing only code never counts as a change. The parts are stored one by one so a mismatch can say what changed. When a run is resumed (agent.resume(id), session.resume(), a later execute() on an unfinished sessionId, agent.approvals.resolve() and the streamed forms), the resuming agent’s fingerprint is compared with the saved one, before any model call or tool runs. onAgentDrift on createAgent() (and on AgentExecutor.execute() / resumeAfterApproval()) decides what a difference does: A pending tool call whose tool no longer exists (the model’s last turn called it and it has no result yet, or the human approved it) always rejects with LOUSHO_RESUME_TOOL_MISSING, whatever the option.
Notes:
  • Checkpoints and snapshots written before this existed have no fingerprint and resume exactly as before, with no warning.
  • A run of a dynamic agent saves the ctx and the model it chose in the checkpoint (runConfig), and a crash resume uses them: the saved model, and tools and instructions resolved again with the saved ctx (before, it resolved with input: [] and no metadata).
  • The fingerprint covers the agent’s own tools, model and instructions, not the tools that subagents and skills add. A sub-agent that pauses for an approval has its own fingerprint in its nested snapshot, but a resume compares only the top-level agent.
  • A run resumed after an approval continues with the instructions it started with (they are in its transcript), so its later checkpoints record those.

Compatibility with stored data

  • Checkpoints written before status existed are treated as 'in-progress' (before this change, finished runs deleted their checkpoint, so only unfinished runs were ever stored).
  • Sessions saved before a store kept a history have an empty history until their next save.
  • Approval snapshots written before remainingToolCalls existed resume as they used to: the calls after the paused one do not run. Each of them gets an error result saying it was not run, so the transcript stays valid.

Limits

  • A sessionId identifies one logical conversation; running two execute() calls on the same id at the same time is not supported.
  • If you call resumeAfterApproval() without the checkpointStore, the session keeps its 'awaiting-approval' mark and later execute() calls keep throwing SessionAwaitingApprovalError; pass the store (or delete the checkpoint).
  • Workers KV is eventually consistent across locations - see Deployment for the caveat.