workflow_cancelled and --resume picks up from it like any other.)
With most tools that is 7 tasks of tokens re-spent. With Nika the run’s
own trace is the checkpoint: --resume re-executes the workflow,
skipping every task whose recorded work is still valid (ADR-099).
No daemon, no run store, no new artifact — the reader of a file you
already have.
That sentence is the whole architecture, so it deserves saying plainly:
durable-execution systems anchor recovery in standing infrastructure —
a database, an embedded log server, or a managed cloud — and pay for it
with author-facing rules (deterministic code, side effects quarantined
into steps, versioning APIs, replay-test suites). Nika anchors recovery
in the hash-chained trace every run already writes. No server, no
database — the trace is the checkpoint: readable with cat, diffable
with git, owned by you. Two hashes decide skip-or-rerun; recorded work
replays from the trace; everything else honestly re-executes. No
determinism is ever demanded of your workflow.
Record, then resume
.nika/traces/... fails before Nika starts
if the folder is missing. Create it, write a named journal, resume that
same path:
--json still writes an automatic timestamped journal
under .nika/traces/; nika trace verify with no path reads the
latest. The copyable sequence above is the named file the --resume
line uses.
Resume from the file you just wrote:
cache hit line in the render, a
task_cache_hit event in the new trace, and the summary counts skipped
vs live. A resumed run never pretends work happened silently.
--resume prints a notice and runs everything live.The skip rule · two hashes, both must match
A task skips iff the trace holds its completed record and- the task definition hashes the same (the verb body,
with:,extract:,retry:/on_error:,when:,for_each:— as now written), and - the resolved inputs hash the same (what its
${{ }}references actually resolved to — upstream outputs,inputs,const,secrets).
--var → the tasks
that consume it re-run, and the mismatch cascades exactly as far as the
data flows — untouched sibling branches still skip:
infer: / agent: task that matches replays its recorded output —
that is the point: crash-resume without re-spending tokens. There are no
determinism rules, no replay constraints, no workflow versioning:
durability is the engine’s problem, never yours. A task that does not
match simply re-runs live, side effects included.
--from · force a re-run the hashes cannot see
Some changes are invisible to hashing: a rotated secret, external state,
an infer: output you want re-rolled. --from <task_id> forces that
task and its transitive downstream to re-run even on a match —
upstream tasks still cache-hit:
--var key.
The durable human gate · pause and answer
Anika:prompt task waits on a human. At a
terminal, the gate asks you directly (confirm [y/N] · choice by
number or value · input verbatim) and the run continues through the
same verified resume path. Anywhere no human can answer (a pipe · CI ·
an agent · --json) with no usable default:, the run does not hang
and does not fail — it pauses durably: the trace records a
workflow_paused event with the prompt payload, the process exits with
code 4, and the frame prints the exact resume line (your
--var/--model included). Pre-answer in one pass with
--answer <task>=<value> at launch.
Nothing ships without a yes. At a terminal, the gated-ship.nika below asks Ship this build to production? [y/N], and y lets ship run. Where nobody can answer, as in CI, the same run pauses durably, exits 4 and prints the one line that resumes it. Run that line: build is a cache hit, the answer rides in --answer approve=true, and ship runs. Every terminal line is captured from the real CLI, offline and with no model; the door, its light, the keypress, the paste and the CI frame are illustration.
--answer (repeatable):
mode:. confirm wants a boolean
(--answer approve=true / =false — a refusal is a value, and the
when: gate downstream decides what happens with it). input takes a
string, choice one of the declared choices. Like --var, the value
parses as JSON when it parses.
Resumed without an answer, a --json run pauses again — idempotent,
exit 4 every time, so a poller can retry harmlessly. Resumed
interactively (a human at the terminal), the prompt simply asks.
Exit codes
Related
- Concepts · Traces: the NDJSON recorder this rides — record, replay, share.
- Concepts · Workflows:
--varinputs (a changed input re-runs exactly the tasks that consume it). - Guides · Testing: the offline golden gate for the same files.
Further reading
The design has a lineage, and naming it beats pretending novelty:- Build Systems à la Carte — Mokhov, Mitchell & Peyton Jones, ICFP
2018 (doi:10.1145/3236774). The
formal frame for hash-keyed skip-or-rerun:
--resumeis a build system’s verifying trace applied to a workflow DAG. - Durable Functions: Semantics for Stateful Serverless — Burckhardt et al., OOPSLA 2021 (doi:10.1145/3485510). The semantics of replay-based durable execution — and of the determinism obligation that comes with it, which content-addressed resume structurally avoids.
- Nextflow — Di Tommaso et al., Nature Biotechnology 2017
(doi:10.1038/nbt.3820). Scientific
computing proved content-addressed
-resumea decade ago; its ecosystem calls it the single most-loved feature.
A trace is not a bearer artifact
Since 0.118 the opening frame binds the trace to the project that wrote it (project_root_fingerprint). --resume judges that binding FIRST: a
trace copied from another project refuses with exit 3 and one typed line
on the --json stream — « a trace is not a bearer artifact: resume it from
the project that wrote it, or run the workflow afresh here ». An older trace
with no fingerprint is no claim.
--resume-unverified waives the chain check for a trace whose journal was
edited or truncated. It does NOT replay a human’s yes: every recorded
nika:prompt decision re-asks under the opt-out (a decision is a
credential, and the chain that bound it to this run is waived), and the door
says the project binding is a recorded field the waived chain no longer
protects.