nika test <file> runs it under the mock provider — offline,
deterministic, zero keys — then diffs the output values against
<file>.golden.json (deep JSON equality, path-by-path). Schema
acceptance is a separate gate: when infer.schema is present the
mock must synthesize a valid instance before that JSON is compared.
Predict: if you change only the prompt text, does the golden still
match? Not if the recorded values change. mock/echo echoes the
prompt, so the golden fails. A schema mock that still synthesizes the
same instance can match even when the prompt changed — the golden
never means “shape only”.
Why the mock makes this possible
nika test swaps every infer: / agent: call to the mock provider,
whatever model: the file declares — your ollama/qwen3.5:4b (or
openai/…, or any cloud model) workflow tests offline, unchanged.
The mock is not a stub that returns "ok". It is schema-conformant:
when a task declares a schema:, the mock synthesizes an instance that
validates against it — required fields present, types respected. Every
schema workflow runs offline, end to end, through the same runtime,
bindings, and typed-outputs validation as a live run.
triage.nika
The golden lifecycle
First run — no golden exists yet, sonika test teaches instead of
guessing (exit 3):
--update, review it once, commit it:
triage.nika.golden.json
nika test compares. A match is exit 0; drift renders a
per-path diff and exits 1:
--update and commit the new golden — the diff shows up in code
review, exactly like a snapshot test.
Proving a rule against cases
Anika:decide bundle is a rubric
with consequences, so it wants a corpus, not one happy path. The kernel
enforces the bundle’s own laws, but a fixture’s class is only a label:
nothing asserts that a positive fixture actually recommends. nika test is
how you assert it today.
The pattern that works: the rule lives once, the cases live one per file.
const:, the kernel in
the middle, the outcome out. No model, no network, no clock, so the only thing
under test is the rule.
refund-fresh-returned.nika
refund-fresh-returned.nika.golden.json
6000 that becomes 600, and every case that leaned on it goes red
at once, naming the flip:
Two limits, stated plainly
nika testtakes no--var. Its flags are--update·--answer·--color·--hyperlink·--plain·--ascii. One file is one input, so a ten-case corpus is ten small files sharing one bundle. That is more files than a parameterised table would be, and it is also ten reviewable goldens.- Under the mock, the model leg is a constant.
nika testswaps everyinfer:to the mock provider, which synthesizes a schema-conformant value. A case that runsinfer:thennika:decidetherefore proves the law, and proves nothing about the extraction: the facts were handed to it. Keep the model out of the case files, pin the evidence inconst:, and prove the rubric. Extraction quality is a separate exercise against real seats, where it costs money and is worth what it costs.
Exit codes · the CI contract
nika test refuses to run a dirty file the same way nika run does: a
workflow with check findings exits 2 before anything executes.
Wire it into CI
Offline and deterministic means no keys in CI, no flakes, no spend:.github/workflows/nika-test.yml (fragment)
*.golden.json files in the same directory as the workflows
they guard. A PR that edits a workflow either keeps its golden green or
shows the re-pinned golden in the diff — both are reviewable.
Related
- Concepts · Workflows: typed
inputs:in, typedoutputs:out — the callable contract the golden pins. - Reference · CLI: every
nika testflag. - Reference · Builtins: the
nika:decidebundle grammar, what the kernel enforces, and the receipt the cases above pin. - Guides · Troubleshooting: when
nika checkfindings block a test (exit 2).