Skip to main content
A long run dies at task 7 of 9 — laptop lid, a provider hiccup, a kill. (Ctrl-C is not a death since 0.118: the run cancels at its next wave boundary, in-flight work completes and is counted, the trace seals with workflow_cancelled and --resume picks up from it like any other.) With most tools that is 7 tasks of tokens re-spent. With Nika the run’s own trace is the checkpoint: --resume re-executes the workflow, skipping every task whose recorded work is still valid (ADR-099). No daemon, no run store, no new artifact — the reader of a file you already have. That sentence is the whole architecture, so it deserves saying plainly: durable-execution systems anchor recovery in standing infrastructure — a database, an embedded log server, or a managed cloud — and pay for it with author-facing rules (deterministic code, side effects quarantined into steps, versioning APIs, replay-test suites). Nika anchors recovery in the hash-chained trace every run already writes. No server, no database — the trace is the checkpoint: readable with cat, diffable with git, owned by you. Two hashes decide skip-or-rerun; recorded work replays from the trace; everything else honestly re-executes. No determinism is ever demanded of your workflow.

Record, then resume

release-notes.nika
Redirecting stdout into .nika/traces/... fails before Nika starts if the folder is missing. Create it, write a named journal, resume that same path:
A run without --json still writes an automatic timestamped journal under .nika/traces/; nika trace verify with no path reads the latest. The copyable sequence above is the named file the --resume line uses. Resume from the file you just wrote:
Every skip is visible — a cache hit line in the render, a task_cache_hit event in the new trace, and the summary counts skipped vs live. A resumed run never pretends work happened silently.
A trace recorded by an engine version without resume keys is not an error: --resume prints a notice and runs everything live.

The skip rule · two hashes, both must match

A task skips iff the trace holds its completed record and
  1. the task definition hashes the same (the verb body, with:, extract:, retry: / on_error:, when:, for_each: — as now written), and
  2. the resolved inputs hash the same (what its ${{ }} references actually resolved to — upstream outputs, inputs, const, secrets).
Edit a prompt → that task re-runs. Pass a different --var → the tasks that consume it re-run, and the mismatch cascades exactly as far as the data flows — untouched sibling branches still skip:
An infer: / agent: task that matches replays its recorded output — that is the point: crash-resume without re-spending tokens. There are no determinism rules, no replay constraints, no workflow versioning: durability is the engine’s problem, never yours. A task that does not match simply re-runs live, side effects included.
The honest limit — at-least-once, not exactly-once. A task whose definition and resolved inputs still hash the same is skipped. Edit the verb, change --var / inputs, or pass --from and that work runs again — “completed” is not a forever lock. A crash in the middle of a side effect can also double-fire it on re-run. Make external effects idempotent via their own arguments, or fence the node with --from discipline.

--from · force a re-run the hashes cannot see

Some changes are invisible to hashing: a rotated secret, external state, an infer: output you want re-rolled. --from <task_id> forces that task and its transitive downstream to re-run even on a match — upstream tasks still cache-hit:
An unknown task id is refused before the run, like an unknown --var key.
Secrets participate by name, never by value — a trace carries no secret-derived material, so a rotated secret does not invalidate the cache. That is the one sharp edge: after a rotation, force the affected node with --from.

The durable human gate · pause and answer

A nika:prompt task waits on a human. At a terminal, the gate asks you directly (confirm [y/N] · choice by number or value · input verbatim) and the run continues through the same verified resume path. Anywhere no human can answer (a pipe · CI · an agent · --json) with no usable default:, the run does not hang and does not fail — it pauses durably: the trace records a workflow_paused event with the prompt payload, the process exits with code 4, and the frame prints the exact resume line (your --var/--model included). Pre-answer in one pass with --answer <task>=<value> at launch.

Nothing ships without a yes. At a terminal, the gated-ship.nika below asks Ship this build to production? [y/N], and y lets ship run. Where nobody can answer, as in CI, the same run pauses durably, exits 4 and prints the one line that resumes it. Run that line: build is a cache hit, the answer rides in --answer approve=true, and ship runs. Every terminal line is captured from the real CLI, offline and with no model; the door, its light, the keypress, the paste and the CI frame are illustration.

gated-ship.nika
The paused trace carries everything needed to pick the run back up — hours later, on the same machine, by a different process:
the workflow_paused event (one line of the trace · reformatted)
Resume with the answer bound to the task id via --answer (repeatable):
The value follows the prompt’s mode:. confirm wants a boolean (--answer approve=true / =false — a refusal is a value, and the when: gate downstream decides what happens with it). input takes a string, choice one of the declared choices. Like --var, the value parses as JSON when it parses. Resumed without an answer, a --json run pauses again — idempotent, exit 4 every time, so a poller can retry harmlessly. Resumed interactively (a human at the terminal), the prompt simply asks.

Exit codes

Further reading

The design has a lineage, and naming it beats pretending novelty:
  • Build Systems à la Carte — Mokhov, Mitchell & Peyton Jones, ICFP 2018 (doi:10.1145/3236774). The formal frame for hash-keyed skip-or-rerun: --resume is a build system’s verifying trace applied to a workflow DAG.
  • Durable Functions: Semantics for Stateful Serverless — Burckhardt et al., OOPSLA 2021 (doi:10.1145/3485510). The semantics of replay-based durable execution — and of the determinism obligation that comes with it, which content-addressed resume structurally avoids.
  • Nextflow — Di Tommaso et al., Nature Biotechnology 2017 (doi:10.1038/nbt.3820). Scientific computing proved content-addressed -resume a decade ago; its ecosystem calls it the single most-loved feature.

A trace is not a bearer artifact

Since 0.118 the opening frame binds the trace to the project that wrote it (project_root_fingerprint). --resume judges that binding FIRST: a trace copied from another project refuses with exit 3 and one typed line on the --json stream — « a trace is not a bearer artifact: resume it from the project that wrote it, or run the workflow afresh here ». An older trace with no fingerprint is no claim. --resume-unverified waives the chain check for a trace whose journal was edited or truncated. It does NOT replay a human’s yes: every recorded nika:prompt decision re-asks under the opt-out (a decision is a credential, and the chain that bound it to this run is waived), and the door says the project binding is a recorded field the waived chain no longer protects.