OPENAI_API_KEY for the model used here) and a project
directory to run it in. The finished file is at the end of
step 5, and you run it as
npx tsx coding.ts "Fix the failing test".
1. Give the agent a workspace
A workspace is a directory the agent may work in.NodeWorkspace is both a
file system and a shell for it, and createFsTools() and createShellTool()
turn it into tools.
root matters: every path the model passes is confined to it, and ..,
absolute paths and drive letters are rejected (see the
security model). Run the program from
the project directory and use process.cwd() as the root, so the file tools
cannot reach anything outside the project.
The model sees seven tools: read_file, write_file, edit_file,
list_dir, glob and grep from createFsTools(), and shell from
createShellTool(). Their arguments and options are in
Workspace tools.
2. Decide what needs approval
A tool withneedsApproval pauses the run before it executes, so a human
decides. Reading is harmless, so gate the file tools that change things with a
per-tool map. For the shell, pass a predicate over the command string and gate
only the commands you consider risky.
needsApproval: true); the
predicate above lets everything else, such as npm test or git status, run
without asking. The other way to skip the question is an allow list of the
commands you trust, with needsApproval: false, as shown at the top of
Workspace tools. The regular expression is a
convenience, not a security boundary: a shell command can be written in many
ways. For anything you do not control, keep the default and approve every
command, or run the shell in a sandbox (see SandboxShell in
Workspace tools).
3. Run it and stream the answer
Create the agent with the tools, then callagent.stream(). It returns a run
you iterate for typed events: text.delta carries the model’s text as it is
written and tool.start announces each tool call (see
Streaming).
maxSteps bounds the model calls in one run, so a confused agent cannot loop
forever. This program stops at the first gated call: the stream ends with an
approval.requested event and the run finishes as 'awaiting-approval'. The
next step handles that.
4. Handle the pause
When the stream ends,await run.result gives the final result. If the run
paused, result.approvalId names the pending call. Ask the user, then
continue with agent.approvals.streamResolve(), which returns a new run for
the rest of the work. That run can pause again, so loop until there is no
approvalId.
tool.start of the call you just
approved arrives at its start, so the loop prints each call once. And a
rejected call does not fail the run: the model gets a rejection result instead
of the tool output and carries on. To tell it why, pass a note:
streamResolve({ id, approved: false, note: 'Do not touch the lockfile' }).
The try / finally closes the readline interface even when run.result
rejects because the run failed.
The approval.requested event also carries toolName and args, so you can
print what the call will do before you ask. See Approvals
for the whole API.
5. Keep the conversation
So far each task starts from nothing. Anagent.session() keeps the
transcript, so a follow-up such as “now add a test for the empty case” sees the
earlier turns. session.stream(input) is agent.stream() for a conversation
and returns the same kind of run; a turn that pauses on a gated call continues
in the same session when you resolve it through the same agent (see
Streaming a session turn). Turn the
loop into a REPL: one session.stream() per line the user types, with the
same pause handling.
exit ends it. The session lives in memory, so it ends with the process; to
keep a conversation across runs, give the agent a store (see
Choosing a store).
6. Test it without a model
The agent needs no network to test. AMemoryWorkspace keeps the files in
memory and takes a scripted exec for the shell, and mockModel plays the
model. Assert on workspace.snapshot() and workspace.commands.
write_file the way the real agent does.
send() resolves with finishReason: 'awaiting-approval' and an
approvalId, the file is still unchanged, and agent.approvals.resolve()
runs the write and finishes the turn.
Next steps
- Workspace tools: the full tool reference, the security model, and
SandboxShellto run commands in Docker. - Approvals: durable approval stores,
approvecallbacks and deciding from another process. - Streaming: every event type, cancelling a run, and the web stream formats.
- Sessions: stores, limits and durable sessions.
- Durable execution: surviving a crash in the middle of a long run.
- Configuration:
projectInstructions: trueappends yourAGENTS.mdorCLAUDE.mdto the instructions.