Skip to main content
The doctrine in one line: unknown stays unknown. A cost Nika cannot prove is never rendered as $0.00, a local model is unpriced — your compute, your electricity — never « free », and an estimate always says which side of the truth it sits on: a floor (≥) or a ceiling (≤).

The vocabulary

Every cost surface speaks the same four words: A task goes UNBOUNDED in one of three ways, and its row names which one:
  • UNBOUNDED — no max_tokens declared · the task has no output limit. You can fix this one: declare max_tokens: and, on a priced model, the report becomes a hard ceiling (the check prints that exact hint).
  • UNBOUNDED — no catalog price (local model; spend not metered) · the task is bounded but the model has no price row. Local models stay here by design — pricing your own hardware would be invention. A cloud model missing from the catalog reads (cloud provider; spend unpriced) instead.
  • UNBOUNDED — for_each over an expression (unknown count) · a fan-out whose item count is known only at run time.

Before a token: the COST rung

nika check audits cost statically, next to the plan and the permits — the same ladder, every time. On weekly-update.nika, a cloud workflow whose one task declares no max_tokens, the COST rung and its hint read:
Declare max_tokens: 4096 on that task and the same rung reads:
The ceiling is $15 per million output tokens × 4096, at the catalog price of the snapshot the rung names; the prompt’s input tokens, exec: and MCP calls are outside it. nika explain <file> narrates the same numbers in beginner words — ≤ $0.0614 worst case · ≥ $0.0614 cheapest path, and for a local model local models: your compute · tokens unpriced — not « free » — and nika run --dry-run carries them onto the plan.

The budget gate: refusal, not remorse

Priced before the first token, refused before the run. nika check finds no max_tokens on the update task and marks it UNBOUNDED. One line, max_tokens: 4096, gives it a worst-case output ceiling of $0.0614: $15 per million output tokens × 4096. With --max-cost-usd 0.05, nika run refuses to start (NIKA-1709, exit 2): no model is called and no run is recorded. Every terminal line and number is captured from the real CLI, offline and with no API key. Prices are catalog estimates; the receipt, the meter and the stamp are illustration.

--max-cost-usd is a block-before-spend gate, not a post-hoc alarm:
  • If the static floor already exceeds the budget, the run refuses to start (exit 2, before any provider is touched):
  • The pre-start refusal prices the effective model — --model override included.
  • If unbounded work rides along, the gate says so loudly on stderr — a budget over work with no ceiling is a promise it names, not one it fakes.
  • Mid-run, the ledger stops the workflow the moment metered spend crosses the budget (NIKA-1704) — settled tasks stay settled, the trace records what was spent.
  • The flag itself is guarded: a non-finite value (NaN, inf) is rejected at parse time — a budget that cannot compare is not a budget.

After the run: the totals stay honest

Every run ends with a totals line that never turns an unknown price into $0.00. nika run hello.nika on the introduction’s mock/echo workflow ends:
Its one call is unpriced, so the total is a word, not a number. A run that calls no model says unmetered. When priced and unpriced calls share a run, the total reads ≥ $… (N unpriced): the ≥ is printed because an unpriced call rode the run. Per-task spend rides the trace itself (cost_usd on the terminal events), so nika trace show and the run report read the recorded ledger — never a summary’s opinion.

Where prices come from — and when they rot

Prices are a catalog fact with provenance, not a constant: nika check --json carries the pricing snapshot (source · date · hash · derived counts) so a cost claim is traceable to the table that produced it, and nika doctor warns when the snapshot is stale enough to distrust — an old price table silently undercounts, which is the one direction honesty cannot tolerate. Cache-aware accounting follows the OpenTelemetry GenAI convention (input includes cache reads), so exported traces mean the same thing your dashboards expect.

The workflow-side controls

  • max_tokens: per task · turns UNBOUNDED into a ceiling on a priced model (a local model stays unpriced).
  • agent: budgets (max_turns · token budgets) bound the loop the same way — exhaustion is a named failure, never a silent overrun.
  • nika check before nika run, always: the cost story is part of the same pre-flight as permits and secrets.
Local-first corollary: a workflow that runs entirely on ollama/…, llamacpp/… or vllm/… reports unpriced (N calls), not a dollar total — Nika never converts « I don’t know the price » into « it’s free ». That distinction is the whole doctrine.