You are reading Nightly documentation for 0.12.6.dev0+g563af09.

This documentation may describe behavior that differs from Stable.

Open Stable documentation

Security and privacy

Security and privacy

Holaryn Agent is local-first: the agent loop, memory, scheduler, and web dashboard all run on your machine, with no account and no vendor backend. This page explains where your data and secrets live, what the agent exposes on the network, and the layers that keep an autonomous agent safe to run.

Local-first by default

Everything durable — memory, run journals, schedules, parked approvals, chat transcripts — lives in SQLite databases inside a state directory on your machine. The only outbound traffic in a default setup is to the model providers you configure in Providers & Models. Optional features that add network surface (peering, webhooks, Telegram) are strictly opt-in.

Secrets

  • Provider API keys and other secrets live in a write-only secret store (secrets.json in the state directory), managed through Settings or the JSON API. The UI never reads secrets back to the browser. When encrypted local state is enabled, the entire secret store is authenticated and encrypted; missing or wrong keys fail closed and never reset it.
  • Configuration is environment-driven (HOLARYN_* and ANTHROPIC_* variables); env values override the store. Secrets are never written to config files — fields like the memory database password (HOLARYN_MEMORY_DB_PASSWORD) or a Qdrant API key are env-or-store only, and the settings editor cannot smuggle them into the config file.
  • The Qdrant URL, the Ollama host and the memory database URL accept only parts that cannot carry a secret: a scheme, a host, a port and a path (the database URL also its user name and a short list of PostgreSQL options), so a credential cannot ride along in them. Their credentials have settings of their own, among them HOLARYN_MEMORY_DB_SSL_KEY_PASSWORD for an encrypted PostgreSQL client key. Holaryn's errors for these backends name what failed, never what the client library said, and each client is built from the URL's checked parts, never from the text entered (SA-629). If a credential was ever in one of those URLs, rotate it: see memory.
  • A provider connection's base URL and the HACP endpoint follow the same rule (SA-632): a scheme, a host, a port and a path, never a user name, password, query or fragment, and every client receives the URL rebuilt from its checked parts. A base URL an earlier release saved with a credential in it is kept in providers.json but never shown or used until you correct it; see what a base URL may contain. Rotate any credential that was in one. The same holds for the URL overrides of a configuration profile or an agent manifest, and for a local runtime an earlier release adopted. MCP server URLs and custom web-search endpoints may keep a query their service needs; they are shown without it.
  • For an OS service, store provider keys in the secret store (Settings → Providers & Models), not in the generated service definition. Never set them as machine-wide environment variables: every account on the computer can read those. See Running the agent.

Encryption is on by default for a fresh install: the first provider key you save turns it on
with your operating system's key vault (Windows DPAPI or the system credential store), and the
onboarding Recover step lets you set or generate a recovery password that writes an
encrypted recovery backup into the state directory's backups folder. Holaryn never stores that
password; keep it somewhere safe. An install that saved secrets before this default keeps
working unchanged and is not rewritten behind your back: stop the host and run
holaryn encryption enable to encrypt it, and until then each plaintext save logs that command
once. A plaintext secret store is still protected by the owner-only state-directory and per-file
permissions, but not against malware running as your user, copied backups, or an unlocked disk.
Set HOLARYN_ENCRYPTION_DEFAULT=off only for an appliance or test environment that must stay
plaintext.

State directory permissions

The host secures its state directory before loading anything from it, and fails closed if it cannot:

  • POSIX: the directory is created (and kept) mode 0700 — private to the owner.
  • Windows: the default state path lives under ProgramData, whose inherited ACL commonly grants the local Users group write access. The host replaces that with an exact protected DACL granting full control only to the current user, SYSTEM, Administrators, and the explicitly authorized service operator — so another local user cannot plant files the service would execute.
  • Paths that are symlinks or junctions are rejected outright.

Permissions are one layer, not offline protection. Optional
encrypted local state adds per-domain authenticated encryption for
classified secrets, journals, memory, artifacts, device state, and recovery backups. It is
opt-in pending independent security review. Its published inventory calls out uncovered
control-plane stores and residual metadata; full-disk encryption remains recommended.

The scoped secret broker keeps credential values out of prompts,
configs, transcripts, and general extension context. Exact, expiring leases resolve only inside
trusted provider, connector, signing, or process adapters; redaction remains defence-in-depth.

Network exposure

  • The web UI binds 127.0.0.1:8765 by default — reachable only from your own machine. Serving on any non-loopback address requires an explicit --auth-token (the operator token), and a DNS-rebinding guard rejects requests for other Host names and requests that carry no Host header at all.
  • The peer inbox does not exist until you set a peer token (holaryn peer token set): it answers 404, so peering is opt-in. Peer tokens are scoped — they open only the inbox, never operator APIs — and are checked in constant time.
  • Open network mode removes the token requirement on trusted LANs; senders must still resolve to a registered, enabled peer, but any device on the network could then claim a peer's name. Use it only on a network where you trust every device. See Peers and networking.
  • The host listener speaks plain HTTP. For cross-machine traffic, use Tailscale (or another WireGuard-class private network) or terminate TLS in front. Never expose the listener to the open internet. When the listener binds anything other than a loopback address, holaryn serve prints, and the host log records, a warning that the operator token and all traffic cross the network unencrypted. Peer pairing and peer messages use the same plain-HTTP listener; see Peers and networking.
  • The listener bounds its connections: idle or slow connections that never send a complete request are closed after 20 seconds, the oldest idle connection gives way when the connection limit is reached, and beyond that limit requests get 503 instead of hanging. On Linux and macOS the host also raises its open-file limit at startup.
  • The listener ignores X-Forwarded-For and Forwarded headers: the client address is always the real connection peer, so a local process cannot claim another machine's address. Responses carry no Server header naming the HTTP server software.
  • The local web boundary is hardened: streaming request-size limits, bounded voice and peer payloads, sandboxed artifact responses with inert content types, anti-framing and no-sniff headers, a restrictive baseline content-security policy, and a no-referrer policy. Malformed input is a client error rather than a server fault: a NUL byte in a path or JSON nested too deeply to decode answers 400, the static file handler never opens OS device names or stream aliases and serves a folder's index page only when it resolves to a regular file inside the web UI directory, and the peer inbox checks its credential before it reads a request body. Static files take their content type from a fixed table rather than the operating system's registry, content-hashed bundle files are cached as immutable, the page shell, service worker and manifest revalidate on every load, and a missing file is never stored by a cache. Pages, static files, the health probes, bookmark redirects and the MCP App proxy answer HEAD like GET without a body; JSON APIs and sign-in callbacks do not, so a link scanner's HEAD can never consume a one-time sign-in code.
  • Errors a browser opens as a page are readable pages (SA-573). When a GET or HEAD names text/html in its Accept header and sends no operator credential header, a plain-text error (a refused host name, an address with no page, an address that needs a signed-in page, an action address, a server failure) and the MCP sign-in results in the system browser come back as a small HTML page with a language, a title, one heading and a viewport, in English, Spanish or French from the browser's Accept-Language. The page holds only fixed text, the status number and the language code: nothing from the request (host name, path, query, cookies or any header) and nothing from the error it replaces. A malformed Accept or Accept-Language value (one with a control character, or ranges that do not follow the header's grammar) is ignored. It carries a stricter content-security policy than the app (default-src 'none', inline style only, no script), Cache-Control: no-store, and Vary: Accept, Accept-Language, and keeps the route's own cookie, cache and security headers; the headers that described the replaced text (its length, digests, ranges, transfer coding and trailers) are dropped. Every other request — the web UI's own requests, event streams, the desktop app, scripts, monitors — gets exactly the bytes it got before, as do JSON errors, redirects, downloads, streams, the health probes and the MCP App proxy.

Automatic update checks are off on a fresh install. Enabling them sends a bounded GitHub Releases
request containing no prompt or credential data; Check now performs the same egress explicitly.
The update check ignores ambient proxy environment variables. The built-in web fetcher, which carries
no credential, follows them (and an ambient CA bundle), so a proxy set in the environment sees the
host names it fetches; it then refuses the proxied response, because the connection's peer is not
an address it validated. See the
data-flow model. Other optional egress paths —
providers, web retrieval, MCP HTTP/SSE, telemetry, channels, webhooks, and peers — require explicit
configuration or an operator action.

Two browser/network residuals are intentionally documented. Native EventSource cannot attach an
Authorization header, so the web UI opens each event stream with a short-lived stream ticket, and
downloads use the header or a single-use download ticket; the host still accepts a query token
for safe methods from older clients. The host's logs mask the value of a URL ticket or token, and
Holaryn recommends TLS when traffic leaves loopback.

Web retrieval rejects private destinations before it connects, then checks that the connection's
peer is an address its own DNS validation approved. A connection whose peer address cannot be
read is refused, and so, in practice, is every response through an explicit HTTP proxy (the peer
is then the proxy), so the built-in fetcher cannot be used through an explicit proxy, whether set
in the environment or the system settings. To filter what it may reach, use a filter on the
network path that does not proxy the connection, such as a firewall or a transparent egress
filter that allows or blocks destinations by address: a blocked destination then fails as
unreachable. Holaryn has no switch that removes the web tools; to keep them from fetching
without your consent, set "web_fetch" = "ask" and "web_search" = "ask" under
[selective.tools] in an autonomy policy profile, so each fetch
waits for approval and none runs unattended. See the
data-flow model.

The prompt-injection guardrail

Tool results and web content are untrusted input. A guardrail screens every tool result inside the loop before it is journaled or reaches the model. Three modes, set via Settings → Capability Center → "Prompt-injection scan" or HOLARYN_INJECTION_SCAN:

  • warn (default) — flagged output passes through with a caution banner telling the model the content is untrusted data and instructions inside it must not be followed.
  • block — flagged output is quarantined: the model receives only a withholding notice naming the matched rules, never the payload.
  • off — no scanning.

The detector is a curated, offline, high-precision rule set: instruction overrides, covert
directives, prompt exfiltration, role smuggling, hidden HTML/Unicode, terminal escapes,
markdown-image exfiltration, encoded directives, and source data impersonating a control message.
Every flag records its detector version, source, bounded plain-text evidence, confidence, and
recommended control. Open the assistant message's Tool activity disclosure to inspect those
facts safely. Inbound peer and teammate messages, attachments, memory, repository/URL context,
connectors, MCP, and tool results retain untrusted provenance through replay and summarization.

Heuristics are a screen, not a proof — novel phrasings can pass. The hard backstop is an
information-flow policy: untrusted data cannot silently drive a write, command, secret access,
memory save, connector call, or outbound message. Keywordless transitions require your explicit
approval. A source already suspected of injection is blocked from consequential actions; that
block cannot be overridden from the approval dialog. Start a clean operator-directed turn or
remove the source instead. A new operator turn resets active lineage so an old warning does not
create permanent approval noise.

To verify the shipped policy offline, run holaryn bench security --fail-on-bypass. If you suspect
an incident, stop the run, record the event id/source/rule versions without copying the payload,
quarantine the source, and follow the private reporting instructions in SECURITY.md. Never paste
a credential or active attack document into a public ticket.

The approval gate is the safety boundary

Every tool carries a consequence (reversible or not, plus a category), and the autonomy posture decides what happens before a consequential action runs. At the default posture, reversible actions proceed and irreversible ones ask first. In unattended runs (schedules, teams, peers) nothing auto-approves: gated actions park in your Approvals inbox and the run waits. No message from a peer, teammate, or delegated child can approve anything — approval authority is yours alone. Details in Autonomy and approvals.

Enterprise deployments can add explicit organization/workspace scope, directory-backed roles,
deny-overrides organization policy, session revocation, legal-hold-aware retention planning, and a
tamper-evident audit chain. This layer can tighten but never weaken the local autonomy and
information-flow floors. See Organization governance.

Signed updates

The desktop app updates only from published GitHub Releases, after you consent. Every update artifact carries a minisign signature, and the public key is baked into the app bundle — the updater verifies the signature before installing, so a tampered or substituted artifact is rejected. A release-time guard also verifies each published signature was made by the key matching the bundled public key.

Other hardening

  • Plugins and skills are inert until enabled. New, updated, and legacy plugins require explicit activation; install flows disclose the MCP commands and URLs a plugin would run, and installs go through validated staging trees that reject symlinks and junctions.
  • Marketplace trust is layered. Signed packages and indexes are reverified against configured
    trust roots, but signatures are not safety claims. Compatibility, permissions, data access,
    dependencies, scans, human review, advisories, organization policy, and runtime containment are
    enforced as separate signals. Revoked packages roll back or disable.
  • Subprocess output is bounded — a runaway child's stdout/stderr cannot exhaust memory, and its process tree is terminated on overflow.
  • MCP calls that may have caused side effects are never silently replayed after a failure; the same principle drives the resume flow after a host crash, which parks unknown in-flight side effects for your decision instead of retrying them.