Models and the price table
Dependency-free helpers for budgeting and context decisions.estimateTokens(input, { model?, estimator? })is a heuristic (about 4 characters per token for English, more for CJK and other scripts, plus per-message overhead and tool-call JSON). Expect roughly 15-20% error on English: fine for compaction and budgets, not for billing. Plug in a real tokenizer withsetTokenEstimator(fn)oroptions.estimator.getModelInfo(id)matches the exact id, thenprovider/id, then dated snapshots (gpt-4o-mini-2024-07-18resolves togpt-4o-mini). Unknown models returnundefined.registerModel(info)adds or overrides an entry; the latest registration wins.estimateCost(usage, model)returns USD, orundefined(not0) when the model or its prices are unknown.
src/models/modelData.ts). Providers change prices and models, so override entries with registerModel when you need billing-grade numbers.
Usage and cost of a run
EveryExecutionResult (and agent.send() result) carries usage, the running total of the whole run:
- Reported vs estimated. Providers report
promptTokens/completionTokens; the built-in providers also pass oncachedInputTokens/reasoningTokenswhen the ‘ai’ SDK’s provider metadata carries them (usage.cachedInputTokensandusage.reasoningTokensonly appear then). A backend that reports nothing givesundefinedusage, never zeros. For that step the executor falls back toestimateTokensand setsusage.estimated(a leading~informatUsage). A custom provider should leaveGenerateResult.usageunset rather than fill in zeros. - Cost.
costUsdis the sum ofestimateCostper model. It isundefined, never a misleading partial sum, as soon as any model used has unknown pricing;byModelshows which ones are priced. - Delegation. A delegated child’s usage is added to the parent’s totals and
byModel, and is also shown on its own asusage.delegated({ inputTokens, outputTokens, totalTokens, costUsd, modelCalls, estimated, runs }). - Resume. A run resumed from a checkpoint, or after an approval, continues from the saved totals instead of restarting at zero. Checkpoints written by older versions start from their saved token counts with an unknown cost.
- Events and traces. The
finishevent (and every lifecycle event that carriedusage) now carries the running totals;text-completealso hasstepUsage, andonLLMResponsereceives the call’s usage as a third argument. Thechatspan’sgen_ai.usage.*attributes use the same numbers, withlousho.usage.estimatedset totruewhen they are estimates. - Streaming.
agent.stream()events carry the same accounting:step.doneandrun.donehaveusagewithinputTokens,outputTokens,estimatedandcostUsd(run-level alsomodelCalls), alongside the olderpromptTokens/completionTokens. promptTokensandcompletionTokensonusageremain as deprecated aliases ofinputTokensandoutputTokens.