@lousho/build-ai-agent/testing ships mockModel: a scripted, deterministic
LLMProvider for unit tests. You write down what the model should say on each
turn, run your agent, and then assert on exactly what the agent sent to the
model. No network, no API keys, no flakiness.
A text reply
A bare string is shorthand for{ text }.
A tool-call flow
Each element of the script is one model turn. Turn 1 asks for a tool; the agent runs it and calls the model again; turn 2 produces the final answer. Tool-call ids are generated deterministically (call_1, call_2, …) unless you pass
id.
An error turn
{ error } makes the model call reject, so you can test how your agent behaves
when the provider fails.
Asserting requests
model.calls holds a deep-frozen snapshot of every request, in order, so later
mutation by the agent can never rewrite history. Use model.lastCall for the
most recent one.
Dynamic turns
Pass a function to compute a turn from the request. It may be async and may return a string or any turn object.Turn reference
Running out of turns
If the agent calls the model more often than you scripted,mockModel throws
an error that names the unexpected call number, shows the last message of the
request and tells you to add a turn. To replay the final turn forever (handy for
loop and step-limit tests), pass { onExhausted: 'repeat-last' }:
model.reset() rewinds the script and clears recorded calls, and
model.assertExhausted() throws if scripted turns remain unused.
Streaming
mockModel also implements stream(): the scripted text is emitted as
text-delta chunks (split on word boundaries), followed by any tool-call
chunks and a final finish chunk, so streaming consumers can be tested too.
Record and replay
mockModel scripts turns by hand. recordReplay is the complementary tool: a
VCR for real model behavior. Run your test once against the real provider and
every generate() / stream() exchange is written to a cassette file. In CI
the same test replays the cassette with no network, no API key and no installed
provider peer package.
() => provider factory that is
only called when recording (so replay never constructs it). It may be
undefined if you only ever replay.
Workflow
- Record locally with a key:
LOUSHO_RECORD=1 OPENAI_API_KEY=... npx vitest run. - Review and commit the cassette (it is stable, 2-space indented JSON, so diffs are readable).
- CI runs
npx vitest runand replays it. Nothing else is needed. - When you change the prompt, tools or flow, replay fails with a
CassetteMismatchError; re-record withLOUSHO_RECORD=1.
mode: 'auto' replays if the cassette file exists and records otherwise.
For defineEval() evals you do not need to wrap the provider yourself:
lousho eval --record writes one cassette per eval case, --replay runs from
them, and --drift reports how each case’s trajectory changed. See
Record, replay and drift.
What replay checks
By default call N is answered by entry N, but only if the request still matches what was recorded: model, each message’s role, content and tool calls, tool names and parameter schemas,temperature and maxTokens. On a mismatch the
error shows the first difference and the re-record hint:
match: 'request': each call finds the
first unused entry with an identical request, in any order. Running out of
entries is also a CassetteMismatchError.
Volatile values do not break matching. Timestamps, dates (2026-10-01,
October 1, 2026), UUIDs and prefixed ids (call_..., toolu_...,
chatcmpl_...) are replaced by placeholders on both sides, and tool-call ids
are ignored (tool calls match by name and arguments). For anything else, pass
normalize:
Errors, streaming and usage
- A provider rejection is recorded and replayed as a rejection with the same
nameandmessage. stream()records the chunks and replays them as an async iterable without delays; passreplayTiming: trueto reproduce the inter-chunk timing. While recording, the real stream is read to the end before it is returned.- Token usage is recorded and replayed, so cost and usage features behave the same in replay.
When the cassette is written
Recording writes the cassette atomically (temp file plus rename) after every call, so a test that fails midway leaves a valid cassette with every call made so far, never a half-written one.await provider.save() forces a write. There
is no process-exit hook. Each record session starts a fresh cassette and
replaces the old file.
Redaction (read this before committing)
Cassettes are meant to be committed, so the wrapper only writes a fixed set of request fields (never headers or provider option bags), never writesrawResponse, and redacts strings that look like API keys (sk-...,
sk-ant-..., Bearer ... tokens and a few other common formats) with
[REDACTED]. That is a safety net, not a guarantee: prompts and model replies
are stored verbatim, so personal data, internal URLs and customer text end up in
the file. Pass redact to scrub them (it runs on every recorded string, in both
requests and responses, and replay applies it when matching):
mockModel or cassettes?
UsemockModel for unit tests of your own logic: you control every turn, can
inject errors and delays, and nothing depends on a model. Use cassettes when
the point is real model behavior (does this prompt and tool set actually lead
to the right tool call?) and hand-scripting turns would only encode your
assumptions. Cassettes need re-recording when prompts change; mockModel tests
only change when the behavior you assert on changes.