defineEval(), live in *.eval.ts files, run under
vitest, and are best run in CI with lousho eval, which
prints a summary and writes JUnit and JSON reports.
Writing a trajectory eval
GivedefineEval() an agent (from createAgent()) and a test function. Send
messages with t.send(), then assert on the run. Using mockModel as the
agent’s provider makes the eval deterministic: no network, no API key, no
flakiness. This is the default way to write an eval; use a real model only
for the judge evals below.
defineEval() also still accepts the original form
({ name, agent, input, provider, score, threshold }): it keeps working
unchanged, and its result shows up in lousho eval as one score assertion.
Assertions
Every assertion is recorded; the case fails at the end and lists every failed gate, so one run shows all that is wrong.
Checks, for
t.check() and t.soft(): includes(text), matches(regex),
equals(value) (yes/no, score 1 or 0), and atLeast(n), atMost(n) (the
number is the score).
Failure messages say what was actually seen:
t.reply (latest reply text), t.result (latest
ExecutionResult), t.toolCalls (every call so far, with parsed arguments).
You can send() more than once; steps, tokens and tool calls add up.
Datasets
Passcases and test runs once per case, each reported separately. A case is
any object; its label (or name, or its input) names it.
agent: () => createAgent(...)) so each case gets its own
agent and its own mockModel script. With a single shared agent, a scripted
model runs out of turns on the second case.
Judge evals
t.judge(rubric) grades the latest reply with an LLM and returns a score from
0 to 1. It needs a judge provider, and lousho never calls a real LLM unless you
configured one: without judge, t.judge() throws an error saying how to fix
it.
*.judge.eval.ts are never picked up by a normal run (and
npm test); only lousho eval --judge (or npm run test:evals:judge in this
repository) runs them.
lousho eval
It spawns the vitest installed in your project (vitest is your dependency; if
it is missing the command prints
npm install --save-dev vitest and exits 2),
then prints one row per case with its scores and duration, totals, and the gate
failures and soft failures listed separately.
Exit code: 1 on any gate failure (or soft failure with --strict, or when
vitest itself fails, for example a file that does not load), 2 when it cannot
run at all (bad arguments, vitest missing), otherwise 0.
How results are collected: each case appends one JSON line to a file named by
the LOUSHO_EVAL_RESULTS environment variable, which lousho eval sets and
reads back. This is more robust than a custom vitest reporter: it works across
vitest versions and worker pools, and needs no module loaded from your project.
JUnit
The report is standard JUnit: onetestsuite per eval, one testcase per case.
A failed gate is a <failure message="..."> carrying the diagnostic message; a
case that threw is an <error>; soft failures are <system-out> notes (or
failures under --strict).
In CI
--tag smoke on every pull request and the full set nightly; run
--judge only where you accept real LLM calls and their cost.
Record, replay and drift
An eval against a real model is slow, costs money and needs a key; the same eval onmockModel only tests the script you wrote. lousho eval sits in
between: it records each case once against the real provider through
recordReplay and replays the recording in CI.
Your eval files do not change.
--recordruns every case with its agent’s real provider and writes one cassette per case next to the eval file:__cassettes__/<eval-name>/<case>.json(names are slugged, for examplerefund-flow/polite.json; an eval withoutcaseswritesdefault.json). Give each case a uniquelabelso the names stay stable. A case whose agent uses more than one provider (a sub-agent on another model) gets<case>.2.jsonand so on. The files use the normal cassette format, with API keys redacted; review and commit them.--replaynever calls the model. A case with no cassette fails withno cassette for "<eval> [<case>]" at ...and the--recordcommand to run; when the agent’s requests changed, the replayed model call fails with aCassetteMismatchErrornaming the first difference. Plainlousho evalwithCIset replays every case that has a cassette and runs the rest live; withoutCIit runs live as before.--driftre-records every case into a temp directory (the committed cassettes are not touched) and compares each with its committed cassette: the ordered tool names, each call’s arguments (JSON-normalized, so key order does not count), the step count and the finish reason. Token usage changes on every real run, so it is compared only with--drift-usage. The summary gets aDrift:table (eval, case, field, committed, current); each drifted case gets a softdriftassertion, so it shows asPASS (soft fail)and as a<system-out>note in the JUnit report. With--strict, drift is a gate failure: the case fails, JUnit has a<failure>, and the run exits 1.
--drift nightly or before a model upgrade, and --replay (or plain
lousho eval in CI) on every pull request. Judges are not recorded: grade with
a real judge only in *.judge.eval.ts files.
How it works: the run loop sends every model call (generate or stream, from
a plain run, a streamed run, a sub-agent or an approval resume) through one
provider-interception seam, setProviderInterceptor() in
src/providers/interception.ts. Nothing is installed by default, so the seam is
free. While a case runs under one of these modes, lousho eval answers it with
the provider wrapped by recordReplay() for that case’s cassette. You can use
the seam yourself to put any wrapper at the model boundary: the interceptor gets
the run’s provider and returns the one to call (it must return the same wrapper
for the same provider, and leave a provider it already wrapped alone).
Run evals against a deployment
The same eval file that gates CI in-process can smoke-test a deployed agent (the node server or Cloudflare Worker). Pointlousho eval at its base
URL and every case runs against the deployment instead of the in-process agent:
POST /chat { sessionId, input }, one
session per case, shared by its t.send() calls) and the SSE stream is read to
run.done. The events become the same result an in-process run produces: tool
calls, final reply, finish reason, steps and usage. So t.calledTool(),
t.completed(), scorers, t.judge(), the summary and --junit work
unchanged. The agent in the file is not used; give the deployment’s own model
and tools the behaviour you want to check.
- A check that needs data the stream does not carry fails and says so (for
example
maxTokenswhenrun.donehas no usage);maxCostUsdis reported as skipped when there is no cost. --urlcannot be combined with--record,--replayor--drift(LOUSHO_CONFIG_CONFLICTING_OPTIONS): cassettes record a provider in-process. Score/threshold evals (score:form) also need an in-process provider and fail with the same code; use trajectory evals.- An unreachable deployment, a non-2xx answer or a truncated stream fails that
case with
LOUSHO_REMOTE_REQUEST_FAILED, and401withLOUSHO_REMOTE_UNAUTHORIZED; the run goes on. The token is never printed or written to a report.
target (this overrides agent; --url overrides both):
remoteTarget({ url, auth?, fetch? }) takes an injectable fetch, which is how
the SDK’s own tests run the real routes in-process without a socket.