Skip to main content

Models and the price table

Dependency-free helpers for budgeting and context decisions.
  • estimateTokens(input, { model?, estimator? }) is a heuristic (about 4 characters per token for English, more for CJK and other scripts, plus per-message overhead and tool-call JSON). Expect roughly 15-20% error on English: fine for compaction and budgets, not for billing. Plug in a real tokenizer with setTokenEstimator(fn) or options.estimator.
  • getModelInfo(id) matches the exact id, then provider/id, then dated snapshots (gpt-4o-mini-2024-07-18 resolves to gpt-4o-mini). Unknown models return undefined.
  • registerModel(info) adds or overrides an entry; the latest registration wins.
  • estimateCost(usage, model) returns USD, or undefined (not 0) when the model or its prices are unknown.
The built-in context windows and prices are a dated snapshot (see the retrieval date and sources at the top of src/models/modelData.ts). Providers change prices and models, so override entries with registerModel when you need billing-grade numbers.

Usage and cost of a run

Every ExecutionResult (and agent.send() result) carries usage, the running total of the whole run:
  • Reported vs estimated. Providers report promptTokens/completionTokens; the built-in providers also pass on cachedInputTokens/reasoningTokens when the ‘ai’ SDK’s provider metadata carries them (usage.cachedInputTokens and usage.reasoningTokens only appear then). A backend that reports nothing gives undefined usage, never zeros. For that step the executor falls back to estimateTokens and sets usage.estimated (a leading ~ in formatUsage). A custom provider should leave GenerateResult.usage unset rather than fill in zeros.
  • Cost. costUsd is the sum of estimateCost per model. It is undefined, never a misleading partial sum, as soon as any model used has unknown pricing; byModel shows which ones are priced.
  • Delegation. A delegated child’s usage is added to the parent’s totals and byModel, and is also shown on its own as usage.delegated ({ inputTokens, outputTokens, totalTokens, costUsd, modelCalls, estimated, runs }).
  • Resume. A run resumed from a checkpoint, or after an approval, continues from the saved totals instead of restarting at zero. Checkpoints written by older versions start from their saved token counts with an unknown cost.
  • Events and traces. The finish event (and every lifecycle event that carried usage) now carries the running totals; text-complete also has stepUsage, and onLLMResponse receives the call’s usage as a third argument. The chat span’s gen_ai.usage.* attributes use the same numbers, with lousho.usage.estimated set to true when they are estimates.
  • Streaming. agent.stream() events carry the same accounting: step.done and run.done have usage with inputTokens, outputTokens, estimated and costUsd (run-level also modelCalls), alongside the older promptTokens/completionTokens.
  • promptTokens and completionTokens on usage remain as deprecated aliases of inputTokens and outputTokens.
Prices come from the model registry above, so to get a cost for a custom or fine-tuned model, register it under the id you pass as the model: