model: at a local server and your data never
leaves the machine. This page is the ladder, from the one-command path to
production servers.
1. The one-command path: Ollama
Install Ollama, then:nika doctor confirms the wiring (it detects the running Ollama server and
prints the exact fix if something is off).
Audit first, then run on your own machine. nika check clears a meeting-notes workflow before any model is called; then a local Ollama model (ollama/llama3.2:3b) turns a sample transcript into typed action items, written to action-items.json with a hash-chained trace. No provider key, no cloud call for the model. Captured from the real CLI; the file is a shorter demo version of the meeting actions example.
2. Any Hugging Face GGUF, still one command
Ollama pulls directly from the Hugging Face Hub. Any public GGUF repo works, no account needed:Q4_K_M is the sane default (quality per
GB), Q8_0 when you have RAM to spare, Q2/Q3 only when memory is tight.
The Hub’s GGUF filter lists
every compatible repo.
3. LM Studio: the visual browser
Prefer a GUI? LM Studio browses Hugging Face, downloads models with a click, and serves an OpenAI-compatible endpoint. Start its server, then:4. Servers: llama.cpp and vLLM
For shared machines and production:- llama.cpp
llama-serverserves any GGUF:model: llamacpp/<model> - vLLM serves full-precision Hub models at
datacenter throughput:
model: vllm/<hub-repo-id>
Checking what’s wired
Thinking time is budgeted for
A local thinking model legitimately needs minutes for one completion on consumer hardware — well past the 30s a cloud round-trip expects. Nika’s provider deadlines are per class: local providers (ollama, lmstudio,
llamacpp, localai, vllm) get ≥ 300s by default, cloud providers
30s. No 30s-everywhere default silently kills a local run mid-think.
When a task needs more than the default, the task’s own timeout: governs
the provider deadline directly:
Swapping between local and cloud
The file does not change shape. One line moves the workflow between a laptop and an API:permits: block with no
net.http entry means the workflow cannot reach the network even if the
model could. Local model + closed permits = fully air-gapped AI work.