sessionId and a checkpointStore to AgentExecutor.execute() and
the run survives a crash, a restart, an abort or a human approval that takes
days. Nothing here depends on a particular host: any CheckpointStore
(LocalStorageCheckpointStore on Node, the Cloudflare KV store the
cloudflare-worker build wires up, or your own three-method implementation)
and any ApprovalStore work the same way.
createAgent(), one store option does the wiring: pass a sessionId to
send() (or stream()) and the run is checkpointed in store.checkpoints;
after a crash, agent.resume(sessionId) finishes it (and returns null when
nothing is pending). No AgentExecutor needed:
store.approvals holds approval pauses, and agent.approvals.resolve()
keeps checkpointing the run under its sessionId. Sessions are checkpointed
too: agent.session({ id }) checkpoints every turn and agent.resume(id) (or
session.resume()) finishes an interrupted one - see
Durable sessions.
What is checkpointed, and when
A checkpoint is the run’s whole transcript plus its step count, usage,businessState and a status. It is written under the sessionId:
A run that rejects (a provider error, a throwing hook, a
PropagatingToolError, a process crash) keeps its last 'in-progress'
checkpoint.
What happens when you call execute() again with the same sessionId
It depends on the stored status:
- Unfinished run (
'in-progress': it crashed, was aborted, or a store write failed). The run resumes. If its last model turn has tool calls without a result, exactly those calls run first, through the normal tool path (argument validation, hooks,needsApproval,toolConcurrency), without calling the model again. Then the loop continues, with the remainingmaxStepsbudget;usagecontinues from the checkpointed totals (the same holds forresumeAfterApproval(), from the snapshot).- New
inputis appended as a user message after those tool results. It is never placed between a tool-call turn and its results, so the transcript stays valid for every provider. This is what lets a user interrupt a run (abort it) and redirect it with a new message. inputthat re-sends the message the run started from (the natural “retry the same request after a crash”) is treated as a retry and is not appended again.input: []also just resumes.
- New
- Finished run (
'finished'). The session continues as a conversation: the stored messages are kept andinputbecomes the next user turn. Pass only the new message(s); if you re-send the stored history followed by new messages, the history is recognised and not duplicated.steps,usageandtoolCallson the result count from zero for the new run;messagesis the whole conversation. - Paused awaiting approval (
'awaiting-approval').execute()throwsSessionAwaitingApprovalError(withsessionIdandapprovalId) and does not call the model or any tool: new input must not bypass a pending decision. Resolve it withresumeAfterApproval()first, passing the samecheckpointStore, then send the new message.
input are ignored when it is added to a stored
session (the session already has its system prompt). To end a session and
start over under the same id, call checkpointStore.delete(sessionId).
Approvals in the middle of a tool batch
When one model turn asks for several tools and one of them needs approval, the calls before it run and are recorded, and the run pauses on it. The approval snapshot records the calls after it (remainingToolCalls).
resumeAfterApproval() then:
- records the paused call’s result - the tool’s result if approved, or a
structured
{ error, note }rejection result if rejected; - runs the remaining calls through the same batch logic: they can run,
fail validation, or pause the run again on another approval (a new
approvalId; resolve it the same way, as many times as needed); - calls the model once every call of the turn has exactly one result.
approvalStore for any later approval unless
you pass a different one in the options.
Guarantees
- Every tool call gets exactly one result, in call order, after any sequence of crashes, aborts, pauses and resumes. (An aborted call gets a “cancelled” result; a rejected call gets a rejection result.)
- A recorded result is never re-executed. Results are checkpointed as the in-order prefix of finished calls grows; a resumed run only runs calls that have no recorded result.
- A checkpointed model turn is never generated again. If the process dies after the model answered, the resumed run executes that answer’s tool calls instead of asking the model again (no extra cost, no different plan).
- The provider always receives a valid transcript: tool results directly follow their tool-call turn, and new user input comes after them.
At-least-once tools: make side effects idempotent
A tool that was running when the process died has no recorded result, so it runs again on resume. Tool execution is therefore at-least-once, not exactly-once. For tools with side effects (charging a card, sending an email), either gate them withneedsApproval or make them idempotent.
Each call’s toolCallId (the model’s id for the call, unchanged when it is
re-run on resume) is passed to execute, so it works as an idempotency
key:
toolCallId is set for every tool run, including tools that route through
sandboxExecute.)
Checkpoint history
A checkpoint record is overwritten on every save, so the run’s earlier steps are gone once it moves on. A store can also keep a bounded history per session: everysave() appends the record to a ring (the newest historyLimit saves,
default 50; the oldest are dropped), which is what replaying or forking a run
from an earlier step is built on.
Which stores keep one:
For example, with the in-memory store:
history(sessionId, { limit? })returns{ step, savedAt, status, checkpoint }entries, newest first (stepischeckpoint.stepIndex,savedAtan ISO timestamp,statusis'in-progress'for a checkpoint without one). The executor saves after every model turn and tool batch, so a session has one entry per save; the entries hold full checkpoints, so lowerhistoryLimitfor long transcripts. An unknown session gives[].historyis optional onCheckpointStore: a custom store without it keeps working, andgetCheckpointHistory()returnsundefinedfor it instead of throwing.KVCheckpointStorekeeps one index key per session (<prefix><sessionId>#history, the entry ids oldest first) and one key per entry (<prefix><sessionId>#history/<id>), so no value grows with the history; the latest checkpoint stays at<prefix><sessionId>, so a store written before history existed still loads. A save writes the checkpoint, then the entry, then the index, then deletes the entries it dropped: a crash between writes leaves an unlisted entry at worst (it expires with the TTL), never an unreadable store. KV has no transactions, so two concurrent saves of one session can lose one index update (that entry is missing fromhistory()), and an entry can be missing for a while at another edge location (eventual consistency;history()skips it). Route a session’s requests to one location when that matters. Each save costs onegetand twoputs more;historyLimit: 0turns it off.- Agent Forge’s file store writes the history as one file per session
(
.lousho/agents/<id>/checkpoint-history/<session>.json, replaced atomically); a file damaged by a crash reads as no history and is rebuilt by the next save. SqliteStorekeeps the history in a newcheckpoint_historytable, added when an existing database file is opened;prune()also removes history entries older than its cutoff.
Fork and replay
AgentExecutor.fork() starts a new session from a step of another session’s
history: “what if the tool had returned something else at step 1?” It takes the
newest history entry of fromStep, applies patch, saves the result as the
'in-progress' checkpoint of the new session and returns
{ sessionId, step, checkpoint }. Resume the fork like any unfinished run:
createAgent({ store }), agent.fork(sessionId, { fromStep, patch? })
does the same over store.checkpoints, and agent.resume(fork.sessionId)
continues the fork:
patchis applied in this order:messages(messages)rewrites the transcript;businessStatereplaces it;toolResult: { toolCallId, result }replaces that call’s result in place (same position, call id and tool name, so the transcript stays valid for every provider), or records it when the call has no result yet;appendInputqueues a user message, which is sent after any tool calls still pending.- The fork keeps the source checkpoint’s step count and usage, so the rest of
the
maxStepsbudget applies to it. Recorded tool results are not re-run; tool calls without a result run first, as on any resume (so a fork of an'awaiting-approval'entry goes throughneedsApprovalagain; the oldapprovalIdis not carried over). newSessionIdnames the fork; the default is<sessionId>.fork-<n>, the firstnwith no checkpoint. AnewSessionIdthat already has a checkpoint is refused.- A step the history does not have (the session is unknown, the step never
ran, or it was dropped past
historyLimit) throws anSDKErrorwith codeLOUSHO_CHECKPOINT_NOT_FOUND, whose message lists the steps that are kept. A store withouthistory()cannot fork (ConfigurationError). compareTrajectories(a, b)takes two checkpoints or two transcripts and returns each run’s steps (one per assistant turn: its text and tool calls with their results),divergedAt(the first step that differs, tool call ids aside) anddrift, the tool-order, argument, step-count and finish-reason differences in the same shape aslousho eval --drift.
Resuming with a changed agent
Every checkpoint and approval snapshot carries an agent fingerprint: the model id, each tool’s name with a hash of its JSON input schema, and a hash of the system instructions (plus a short SHA-256 over all of them, computed with Web Crypto). Functions and other values that cannot be serialized (a tool’sexecute, hooks, callbacks) are not part of it, so changing only code never
counts as a change. The parts are stored one by one so a mismatch can say what
changed.
When a run is resumed (agent.resume(id), session.resume(), a later
execute() on an unfinished sessionId, agent.approvals.resolve() and the
streamed forms), the resuming agent’s fingerprint is compared with the saved
one, before any model call or tool runs. onAgentDrift on createAgent() (and
on AgentExecutor.execute() / resumeAfterApproval()) decides what a
difference does:
A pending tool call whose tool no longer exists (the model’s last turn called
it and it has no result yet, or the human approved it) always rejects with
LOUSHO_RESUME_TOOL_MISSING,
whatever the option.
- Checkpoints and snapshots written before this existed have no fingerprint and resume exactly as before, with no warning.
- A run of a dynamic agent saves the
ctxand the model it chose in the checkpoint (runConfig), and a crash resume uses them: the saved model, and tools and instructions resolved again with the savedctx(before, it resolved withinput: []and nometadata). - The fingerprint covers the agent’s own tools, model and instructions, not
the tools that
subagentsandskillsadd. A sub-agent that pauses for an approval has its own fingerprint in its nested snapshot, but a resume compares only the top-level agent. - A run resumed after an approval continues with the instructions it started with (they are in its transcript), so its later checkpoints record those.
Compatibility with stored data
- Checkpoints written before
statusexisted are treated as'in-progress'(before this change, finished runs deleted their checkpoint, so only unfinished runs were ever stored). - Sessions saved before a store kept a history have an empty history until their next save.
- Approval snapshots written before
remainingToolCallsexisted resume as they used to: the calls after the paused one do not run. Each of them gets an error result saying it was not run, so the transcript stays valid.
Limits
- A
sessionIdidentifies one logical conversation; running twoexecute()calls on the same id at the same time is not supported. - If you call
resumeAfterApproval()without thecheckpointStore, the session keeps its'awaiting-approval'mark and laterexecute()calls keep throwingSessionAwaitingApprovalError; pass the store (or delete the checkpoint). - Workers KV is eventually consistent across locations - see Deployment for the caveat.