Local models
Local models
Holaryn can manage the path from “will this model fit?” to a routed local provider
without turning Ollama or llama.cpp into special cases in the agent loop. Use
Settings → Providers & Models → Local inference desk or the
holaryn local-model command.
Start with an existing runtime
The shortest path on a fresh machine is an existing Ollama model:
holaryn local-model runtimes
holaryn local-model hardware
holaryn local-model add-ollama qwen2.5:7b --id local-qwen
holaryn local-model adopt ollama http://127.0.0.1:11434/v1 --id ollama-local
holaryn local-model register ollama-local local-qwen
holaryn local-model role general local-qwen
holaryn local-model benchmark ollama-local
runtimes only checks this machine. If a runtime is missing, Holaryn reports an
official installation page instead of executing an installer. adopt performs an
explicit health check, but never takes ownership of the process. An adopted
server's Stop control is disabled because its owning runtime must stop it.
An address must not hold a user name, a password, a query or a fragment. One that
an earlier release adopted with such a part is kept in local-models/state.json,
but it is never shown or used; adopt the runtime again with the corrected address
and the same --id to replace it, then rotate the credential it held.
Errors recorded by an earlier release (a benchmark's, a download's, a server's)
are shown with every address cut to its scheme, host, port and path.
The same flow works in Settings: record an Ollama model, adopt the endpoint, scan
it, register the selected model, and assign a role. Once registered, it appears in
the same model pickers and routing paths as every other provider model.
Import a llama.cpp model
For a GGUF file already on disk, supply the license and any provenance you know:
holaryn local-model import D:\models\qwen.gguf --id local-qwen --license Apache-2.0 --quantization Q4_K_M
holaryn local-model start local-qwen --port 8081 --autostart
holaryn local-model health llamacpp-local-qwen
holaryn local-model register llamacpp-local-qwen local-qwen
The import is hashed and copied into content-addressed Holaryn storage. Re-importing
the same bytes reuses the blob. --sha256 can enforce a digest supplied by the
publisher. Holaryn-owned servers bind 127.0.0.1; choose a different port when the
requested one is occupied.
For a large existing file that you do not want Holaryn to copy, run llama.cpp
yourself and use adopt. Adopted endpoints and external model stores are
read-only.
Verified downloads and resume
A download is available only when you explicitly provide immutable artifact
metadata:
holaryn local-model download https://models.example/qwen.gguf --sha256 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef --id local-qwen --license Apache-2.0
Holaryn rejects non-HTTPS URLs, reserves disk space, stores a partial file, resumes
with an HTTP range request, and publishes the model only after the SHA-256 matches.
The model manifest records URL, checksum, publisher, license, and trust source.
The built-in catalog works offline and deliberately does not invent artifact URLs
or checksums for mutable upstream releases.
Hardware fit and roles
holaryn local-model catalog --fit
holaryn local-model role coding local-qwen
holaryn local-model status --json
Fit labels are advisory. Each result shows the parameter and quantization estimate,
requested context, concurrency, memory headroom, disk reserve, probe limitations,
and a confidence label. Apple unified memory is treated as shared memory. Missing
or ambiguous GPU telemetry produces an unknown/low-confidence result rather
than a guess. Expert override remains available because runtime offload behavior
can differ from the estimate.
Roles are general, coding, vision, embedding, and background. They become
ordinary registry tags; general sets the default chat model and embedding
selects embedding modality.
Supervision and diagnostics
Managed processes write logs beneath <state-dir>/local-models/logs. A log that has grown
past 16 MiB is moved to <server-id>.log.1 (replacing the previous one) the next time the
server starts. Health checks record actionable errors. Crashes can restart up to the recorded
limit; port conflicts are reported and Holaryn never kills the process occupying the port. A
conflict keeps the server's autostart setting, provider connection and restart history, so the
next host start (or an explicit restart) tries again once the port is free. Autostart
restores Holaryn-owned servers with the host and host shutdown stops only live child handles
it owns.
holaryn local-model diagnostics
holaryn local-model health llamacpp-local-qwen
holaryn local-model restart llamacpp-local-qwen
holaryn local-model stop llamacpp-local-qwen
Holaryn records each server it starts with the process's PID, boot, executable and
creation time. A server started by holaryn local-model start can therefore be checked,
stopped or restarted by a later holaryn local-model command, by Settings, or by the host,
and a host that restarts after a crash keeps an autostarted server that is still running
instead of launching a second one. A PID the operating system has since reused for another
program never matches that record, so it is never signalled; the server is simply recorded
as no longer running. Servers recorded by older versions, which have no such identity, are
still marked orphaned after a restart: relaunch them or adopt the endpoint deliberately.
A runtime may run behind a launcher, such as a Windows .cmd file or a wrapper script. Stop
and restart end the whole process tree the launcher started, and health checks record the
identities of those processes, so the runtime is still found if its launcher has exited. A
server is reported stopped only once every one of them has exited; otherwise it is recorded
as degraded and can be stopped again.
Safe removal
holaryn local-model remove local-qwen --yes
Removal refuses active models and external storage. For an owned/imported model it
deletes only unshared bytes under Holaryn's blob root; another manifest referencing
the same SHA-256 keeps the shared blob. Settings presents the same scope in a
confirmation dialog.
Offline and privacy behavior
Status, diagnostics, hardware probing, catalog search, fit estimation, roles, and
stored benchmark results are offline. Endpoint scans, health checks, benchmarks,
and downloads happen only when you invoke that action. Holaryn never uploads a
hardware inventory, model file, prompt, or response from this subsystem.
The fixed benchmark sends Reply with exactly: local smoke ok to the selected
local endpoint. Durable results contain only time-to-response, approximate output
throughput, context and declared capability flags, stability, and an error if
present. Prompt and response text are not persisted.
Maintainers can opt into the real-machine acceptance smoke when a loopback Ollama
with at least one model is already running:
$env:HOLARYN_REAL_LOCAL_MODEL_SMOKE = "1"
uv run pytest tests/test_local_models.py -k real_ollama
The test is skipped by default and never installs, downloads, or selects a remote
runtime. It adopts the local endpoint read-only, registers the first reported
model, and runs the same fixed chat benchmark.
For LAN-hosted Ollama guidance and exposure warnings, see
Configuring Ollama. For provider routing details, see
Providers & models.