You are reading Nightly documentation for 0.12.6.dev0+g563af09.

This documentation may describe behavior that differs from Stable.

Open Stable documentation

Providers & Models

Providers & Models

Holaryn Agent gets its models exclusively from a provider registry you configure — there is no built-in vendor and no fallback: an unconfigured agent refuses to run and tells you exactly what to set up. This page explains connections, models, defaults, nicknames, capabilities, and switching models mid-conversation.

The provider registry

The registry lives in providers.json in the host state directory (find it under Settings → System → Paths); API keys go in the state directory's write-only secret store, never in the JSON file. Manage everything in Settings → Providers & Models, which edits two levels:

  • Connections — a provider endpoint: which adapter it speaks, an optional base URL, and where its API key comes from (an environment variable name, or a key stored in the secret store). A stored key is bound to the endpoint it was entered for: editing the base URL to another host, port, or scheme (for example https to http) clears the stored key, and the save message asks you to enter it again.
  • Credential pools — optional authorized groups of compatible connections with fair scheduling,
    quotas, rate/health state, and drain/revoke controls. Pools contain aliases and opaque connection
    references, never keys.
  • Models — the selectable models, each bound to one connection and optionally one credential
    pool, with its capabilities, an optional nickname, and tags.

The registry is validated on load: connection, pool, and model ids must be unique; every model must
reference a known connection; and a model's fallback connection must belong to its selected pool.

Adding a provider

Settings → Providers & Models → Add provider opens a picker over the built-in provider catalog: a filterable, grouped list of curated presets. The groups are:

  • Major providers — first-party APIs (OpenAI, Anthropic, Google Gemini, Mistral, DeepSeek, xAI, Groq, and more), each authenticated with an API key.
  • Aggregators — many vendors behind one key (OpenCode Zen, Abacus.AI RouteLLM,
    OpenRouter, Hugging Face).
  • Subscriptions (sign in) — your personal ChatGPT or GitHub Copilot plan via a device-code sign-in (see the caveats below), plus API-key coding plans.
  • Local — keyless endpoints on your own machine (Ollama, LM Studio, vLLM, llama.cpp). A local endpoint is just another connection: the preset fills the usual base URL (e.g. http://localhost:11434/v1 for Ollama — change the host to reach it over the LAN), no API key is needed, and a turn routed to it never leaves your network. Two gotchas: Ollama's OpenAI-compatible endpoint defaults to a small context (about 4096 tokens) regardless of what the model supports — bake a larger num_ctx into a model variant with a Modelfile and set the model's declared context window in Holaryn to match, because the agent's compaction math depends on it; and small local models are often unreliable at native tool calling — if a model emits raw tool JSON as text, turn off its native tool calling capability and the adapter falls back to a prompted-JSON tool path instead.

For hardware fit, verified downloads/imports, Ollama and llama.cpp process
supervision, local benchmarks, role assignment, and safe owned-file removal, use
the Local models lifecycle manager. It registers a healthy
endpoint back into this same provider registry.
- Custom — full manual configuration for anything else.

Entries you have already configured are marked, and picking one again simply creates a second connection with a uniquified id. Picking a provider prefills everything — connection id, adapter, endpoint, and the suggested key variable, all still editable under an Advanced disclosure — and the submit button is the whole flow:

  • Add & scan (API-key providers) creates the connection, stores the key you paste inline (a Get a key link opens the provider's console), and opens the model scan.
  • Add & sign in (subscriptions) creates the connection and starts the device-code sign-in immediately, chaining into the scan the moment you approve.
  • A local endpoint goes straight to a keyless scan.

The catalog is a template, not a live binding: creating a connection snapshots the preset's values into your own registry, and later releases update the catalog without ever rewriting an existing connection.

Scanning for models

A scan asks the connection's endpoint for the models it serves and lists them as checkboxes — nothing is ever pre-selected. Filter the list, then use Select all shown (it respects the filter and skips models you already added), Select recommended, or Clear selection, and add the whole selection in one step. Models the catalog recommends for the provider are marked Recommended and sort to the top. Added models get the adapter's default capabilities; edit them afterwards as needed. You can rescan a connection at any time from its Actions → Scan for models… menu.

A few providers (Perplexity, for example) have no model-listing endpoint. For those the scan panel explains that and offers one-click Add buttons that open the manual add-model form prefilled with the provider's typical model ids.

Custom / advanced

The Custom entry (and editing an existing connection) opens the full free-form dialog. Each connection picks one adapter — one per wire protocol, not one per vendor:

Adapter Speaks Typical use
anthropic Anthropic Messages API Claude models
openai OpenAI Chat Completions GPT models
openai_responses OpenAI Responses API Newer OpenAI models
gemini Google Gemini API Gemini models
openai_compat OpenAI Chat Completions against any base URL Ollama, LM Studio, vLLM, any local or self-hosted endpoint

After saving a custom connection you can scan it from its Actions menu, exactly like a catalog one.

What a base URL may contain

A base URL is ordinary configuration: it is saved in providers.json and shown on this page, so
it must not hold a secret (SA-632). It is read by the same rule as the memory backend URLs:

  • http:// or https:// (plain http only to a local or private address, unless the connection
    allows insecure HTTP);
  • a host: a name of letters, digits, hyphens and dots (an international name in its xn-- form,
    no underscore, no dot at the end), an IPv4 address, or an IPv6 address in square brackets;
  • an optional port from 1 to 65535;
  • an optional path of letters, digits, ., _, ~ and - between single slashes. A slash at
    the end is dropped.

A user name, a password, a query (the part after ?) and a fragment (the part after #) are
refused. Store the provider's key with Set API key, or enter the name of an environment
variable that holds the key in the connection's API key environment variable field. Each adapter adds its own endpoint path to the
base URL, so no provider needs a query in it: for Azure OpenAI use the
https://<resource>.openai.azure.com/openai/v1 endpoint, which needs no api-version. Every
client, and this page, use the URL rebuilt from its parts: lower-case scheme and host, without
the default port or a final slash.

A host name with an underscore, such as the Docker service name my_ollama, is refused. Give the service a name or a network alias without an underscore, for example my-ollama, or enter its IP address.

A proxy that takes a user name and password in the URL cannot be used as it is. Holaryn has no setting for a proxy user name and password. Set the proxy up to accept the connection's API key, or to need no sign-in from the machine Holaryn runs on.

A connection saved by an earlier release with a base URL this rule refuses is kept as it is in
providers.json, because it may be the only copy of a credential, but it is never used and its
URL is never shown. Its row on this page says "Not in use" after its id, then what is wrong and
what to do; its Base URL reads "Not shown", it has no Set API key or Scan for models
action, and Delete says that the unseen URL goes too. A run, a model scan or an image request
on it fails with the same message, and Holaryn logs it at start-up. Choose Edit, enter the corrected base URL and
choose Save: an empty Base URL is refused for such a connection, so the saved value is not
lost by accident (enter the adapter's default endpoint to use it). The save says that the earlier
URL is no longer in providers.json; rotate the credential it held, because earlier copies of the
file, logs and backups may still hold it. holaryn onboard configure with the same connection id
also replaces it, and says so.

A flagged ChatGPT or GitHub Copilot connection does not use or refresh its sign-in either, and no
flagged connection is given a key, whether from the secret store or from its environment
variable. A base URL is saved in the form rebuilt from its parts, so HTTPS://Api.Example.com/v1/
is saved as https://api.example.com/v1; HTTP:// in any case is plain HTTP, like http://.
A saved connection that uses plain HTTP to a public host without allow_insecure_http set to
true in providers.json is kept in the same way, with its models, credential pools and default
selections: it is not shown or used, and its row says what to do. Saving another connection
does not remove it.

Credential pools

Credential pools are for organization-approved accounts or endpoints that may share one workload.
They are not a way to combine unrelated users' subscriptions or evade provider limits. All members
must use the same provider wire adapter.

In Settings → Providers & Models → Credential pools:

  1. Add a pool and review the generated JSON template.
  2. Set a safe alias for each connection, its weight/priority, optional
    tenant/project/profile allowlists and validity window, and pool/member request, token, spend, and
    concurrency limits.
  3. Save, then use Preview selection to see which alias is eligible without consuming capacity.
  4. Edit a model and select the pool. Its ordinary connection remains a compatible fallback and must
    be one of the pool members.

The pool card shows active leases, aggregate requests/tokens/cost, provider-reported remaining
capacity, circuit/auth/drain state, and recent selection reasons. Drain gracefully stops new work
but lets in-flight work finish. Revoke and cancel stops new work, cancels durable leases, and
discards late model results. Re-activate only after resolving the underlying account condition.

Weights are deterministic: over time a weight-2 member receives roughly twice the selections of a
weight-1 member at the same priority. Higher member priority is considered first. Reserved priority
slots keep part of pool/member concurrency available for requests at or above the configured
threshold. Sticky conversations stay on their prior account when it remains eligible; disabling
account switching makes an unavailable sticky account throttle instead of silently changing billing
accounts.

See the credential-pool ADR and recovery guide for the full state,
quota, failure, privacy, and troubleshooting contract.

Abacus.AI RouteLLM

Choose Abacus.AI (RouteLLM) in Add provider, open Get a key, and
save your RouteLLM key. The preset fills https://routellm.abacus.ai/v1 and the
OpenAI-compatible adapter. Enterprise users can change the endpoint to the one
shown in their Abacus workspace. Review your account's API credits and billing.

Scan for models and choose a text model, or add route-llm for automatic
selection. Model capabilities remain editable; choose a specific model when
you need predictable capabilities. The CLI equivalent is
holaryn onboard configure --preset abacus --model route-llm; it prompts for
the key with input hidden. ABACUS_API_KEY is the environment-variable option.

Abacus documents bearer-key API authentication. Its account website is where
you obtain that key; this preset does not offer an undocumented OAuth flow.
Text streaming, client tool calls, reported usage and errors use Holaryn's
existing adapter. Abacus-hosted tools, image generation and audio generation
are not enabled by this model connection. These contracts were checked offline;
live Abacus verification has not been performed.
See Abacus's API reference
and account/API billing FAQ.

OpenCode Zen

In Add provider, filter for OpenCode and choose the group containing
your model: Kimi, GLM, MiniMax and more, GPT, Grok and Muse,
Claude and Qwen, or Gemini. Each fills the corresponding adapter and
Zen endpoint. Open Get a key, obtain your Zen API key, and save it in the
dialog. One OPENCODE_API_KEY environment variable can supply all four
connections when you use several model families.

Add & scan shows available models whose endpoint is documented for that
group. The reviewed protocol table is dated 2026-09-12. Models absent from that
table are omitted from scans; after checking their current endpoint in the
Zen documentation, you can add them manually
to a matching connection. A scan does not verify capabilities or model quality.

CLI examples (each prompts for a hidden key):

holaryn onboard configure --preset opencode-zen --model kimi-k2.5
holaryn onboard configure --preset opencode-zen-responses --model gpt-5.5
holaryn onboard configure --preset opencode-zen-messages --model claude-sonnet-4-6
holaryn onboard configure --preset opencode-zen-gemini --model gemini-3.1-pro

These connections use Zen API-key access and its billing. OpenCode Go has a
different subscription endpoint and is not configured by these presets. Account
sign-in obtains your key; no undocumented OAuth method is offered. This is
direct model access; OpenCode CLI delegation is tracked separately. Live Zen
verification has not been performed; adapters were checked with offline HTTP
fixtures. Model-specific capability and retention settings remain under your
control in the model editor.

Nous Portal through the Hermes subscription proxy

Choose Nous Portal (Hermes subscription proxy) in Add provider. This
uses Nous's documented integration for external OpenAI-compatible apps. It
requires Hermes Agent on the same machine as the Holaryn host; inference still
runs in the cloud and uses your Nous subscription.

  1. Install Hermes using its official setup instructions linked from the
    subscription proxy guide.
  2. Run hermes auth add nous and approve the account sign-in in your browser.
  3. Keep hermes proxy start --provider nous running. Its default address is
    http://127.0.0.1:8645/v1.
  4. In Holaryn's preset dialog, leave API key (optional) blank and choose
    Add & scan. Select a model, add it, and make it your default as desired.

For CLI configuration, use holaryn onboard configure --preset nous-portal
after starting the proxy. Add --model anthropic/claude-sonnet-4.6 if you already
know the exact model ID and want to skip discovery. Review the chosen model's
capabilities; availability depends on your account. Nous recommends agentic
models for tool use and distinguishes those from its Hermes 4 chat models in
the Portal guide.

If scanning fails, check that Hermes is running on the Holaryn host and that
its Nous sign-in is current, then scan again. An expired or revoked Portal
session requires signing in again through Hermes. Keep the proxy on loopback:
it has no client authentication. With a remote or containerized Holaryn host,
the proxy must be reachable from that host; do not expose it publicly.

Hermes owns OAuth, refresh and sign-out. Holaryn neither imports its token
files nor uses Hermes's OAuth client ID for a separate login. Removing this
Holaryn connection does not revoke that external session. Stop active Holaryn
runs before switching the proxy's Nous account: Holaryn cannot independently
observe account changes behind this endpoint. This is raw model inference;
Holaryn continues to execute its own tools. The preset does not enable the
Nous Tool Gateway, hosted agents, images, audio or embeddings. No Portal
credential belongs in Holaryn's API-key field. Offline setup and transport
contracts are tested; live Portal sign-in and inference remain unperformed.

Authentication: API key or subscription sign-in

Most connections authenticate with an API key — either the name of an environment variable holding it, or a key stored on this machine (Actions → Set API key, write-only: the UI only ever shows "set" or "not set").

Two connection types can instead sign in with a personal subscription — a short code + URL device flow you approve in a browser. Both use Actions → Sign in with … or, from the terminal, holaryn connection login <id> (with logout / status). Tokens are stored on this machine like any other secret and never shown; the settings page displays only the signed-in state and your account.

In the web UI, pressing Sign in opens your default browser at the provider's page and shows the code with a Copy code button; the panel announces each step in a live region and confirms when you approve. If the browser does not open — a popup blocker, say — the page is still listed as a link and Open sign-in page tries again. The terminal command prints the URL and code for you to open yourself, which is what makes it usable over SSH.

Sign in with ChatGPT — on an OpenAI Responses connection. Requests bill against your ChatGPT Plus/Pro subscription instead of per-token API credits. You approve a short code at https://auth.openai.com/codex/device.

  • Turn on device-code login first. OpenAI requires the account to opt in, and sign-in cannot work until you do: for a personal account, ChatGPT → Settings → Security → "Allow device code login"; for a work/team account, an admin enables it under Workspace Settings → Permissions. If it is off, Holaryn reports that device-code login is not enabled and points you here.
  • Personal use of your own account only. Do not pool accounts or route several people through one subscription.
  • Unofficial and not contractual. It rides an undocumented Codex backend endpoint with no SLA and may stop working without notice.

Sign in with GitHub Copilot — on an OpenAI-compatible connection. Requests use your GitHub Copilot subscription against Copilot's OpenAI-compatible chat endpoint.

  • Unofficial, and against GitHub's editor-only terms. It calls GitHub's internal Copilot endpoints outside a supported editor. GitHub does not authorize this, and its API-abuse terms allow discretionary suspension — using it can get your GitHub account suspended. Personal use of your own account only; it may also break without notice. Use it only if you accept that risk.

Anthropic and Google are prohibited and appear in the menu but disabled: their subscription tokens are scoped to their own apps, with account suspensions for violations. Use an API key for Claude and Gemini.

Models, the default, and nicknames

Every model row has:

  • Id — the label shown in pickers.
  • Capabilities — see below.
  • Nickname — an optional friendly name. Nicknames are shown on the agent's replies and double as @-aliases: @bob what's the weather? runs that turn on the model nicknamed bob. A nickname must be a single word, unique, and must not collide with a profile name, coding backend, or paired peer.
  • Tags — free-form labels (e.g. local, coding) for your own organization.
  • Enabled — disabled models stay manageable in Settings but are hidden from every model chooser.
  • Credential pool — optional fair-capacity scheduling across compatible authorized connections;
    leaving it blank preserves single-connection behavior.

One model is the default: it is what a brand-new chat or session uses. Set it in Settings → Providers & Models; the Session status panel shows which model a fresh session would get right now.

Capability awareness

Each model declares a capability descriptor the agent consults before every turn:

  • Context window — drives the model-aware compaction that keeps long conversations within budget.
  • Native tool calling — whether the model emits real tool calls; when off, the agent emulates tool use with a prompted-JSON protocol.
  • Parallel tool calls — whether the model may request several tools at once.
  • Vision — whether the model accepts images.
  • Prompt cache — whether the provider supports prompt caching. Retention remains off until you
    configure and consent to a per-model prompt cache policy.
  • Reasoning effort — whether the model accepts an effort/reasoning knob; only models that opt in enable the effort selector in chat. The values a model can actually carry come from its adapter's protocol contract (low, medium, high for the OpenAI-style adapters; Anthropic and Gemini do not forward effort at all), and Supported efforts can narrow that to a subset; a subset that keeps no value the adapter can carry is refused when you save the model, naming the values it cannot forward. A value outside the contract is refused before anything is sent, with an explanation of what is supported, and the chat controls, Settings, and the agent itself all refuse exactly the same values.
  • Structured output and output limit — whether typed output is supported and the largest
    requested result the model can accept.
  • Reasoning modes — the specific effort modes the model supports.
  • Routing metadata — expected quality/latency, operator preference, availability, rate
    capacity, locality, and the current immutable pricing snapshot used by optional adaptive routing.

Conformance evidence

Settings → Providers & Models shows each model's conformance status: declared-only (the
descriptor and the adapter contract), verified (the offline transcript round-trip matrix or an
explicit live run completed and observed the contract), stale (observed evidence has expired),
or unknown (a run was made but observed or passed nothing, with a line saying which). A live
run that observed nothing at all is not stored: the model keeps whatever it had before, so a
model with no earlier record still reads as declared-only and one with earlier evidence keeps it;
the failed benchmark report is that attempt's record. Refresh a
model's offline evidence from Settings or with holaryn bench conformance --model ID; live
evidence needs an explicit request budget, and that budget counts provider requests, so a probed
turn is never retried behind it. Evidence belongs to the connection it was recorded on: two
connections that use the same provider model string never share it. Evidence is invalidated
whenever the model identity changes or the adapter code changes, including the shared provider
code an adapter reuses, and is kept in model-conformance.json in the state directory. It
records verdicts and identities only, never conversation content or opaque provider material.
See model tool-capability negotiation.

Switching models per turn

You never have to leave a conversation to try another model:

  • The model picker below the composer switches the chat's model for all following turns.
  • An @-mention — @<model-id> or @<nickname> as the first token of a message — runs that single turn on the named model, then the chat returns to its selected model.
  • The effort selector adjusts the reasoning effort for models whose capability descriptor opts in. A value the selected model's contract cannot carry is refused with an explanation of the supported values; for models that do not forward effort the selection is kept for the chat but nothing is sent. Effort is a provider hint, not an assurance of quality or safety; it is separate from an explicit plan and from the experimental branching reasoning engine.

Routing is sticky in the default static mode. The opt-in adaptive router can
choose per task/role under explicit capability, privacy, cost, latency, and pin constraints.
The separate opt-in provider recovery policy can retry a safe
pre-response failure or move to another eligible model; it never replays a consequential tool or
continues automatically after visible partial output.

Proxies and private certificate authorities

Provider requests carry your API key and the conversation, so Holaryn does not send them through
a proxy or trust a certificate bundle it merely finds in the environment (HTTPS_PROXY,
HTTP_PROXY, ALL_PROXY, SSL_CERT_FILE, SSL_CERT_DIR, and for the clients built on other
libraries grpc_proxy, REQUESTS_CA_BUNDLE, CURL_CA_BUNDLE and
GRPC_DEFAULT_SSL_ROOTS_FILE_PATH). The same rule covers these requests, which also carry a
credential or your content (SA-571, SA-590):

  • model providers: model turns, image generation and Scan for models;
  • provider sign-in: the ChatGPT and Copilot device sign-in and their token refresh;
  • connected apps: Google API calls and the OAuth sign-in and token refresh;
  • the Telegram channel and its Test connection check;
  • the Ollama memory embedder, which sends memory text;
  • remote MCP servers over Streamable HTTP or SSE that authenticate with a static header from
    their configuration (MCP servers that use OAuth sign-in never use those variables);
  • the holaryn commands that call the running host's API (holaryn question, steer,
    resume, the peer, team and coding-job actions and the like) and holaryn capability,
    which send the host's operator token;
  • the Python SDK (holaryn_agent.sdk) clients, which send your public automation API key;
  • public automation API webhook deliveries, which carry signed events and their results;
  • the web search providers (Brave, Tavily, Exa, Google grounding and a credentialed custom
    search endpoint);
  • the OTLP trace, metric and log exporters over HTTP or gRPC, which send the headers and the
    TLS client certificate configured for them (a CA set with HOLARYN_OTEL_CA_CERT_PATH always
    applies; over HTTP the exporters' own OTEL_EXPORTER_OTLP_CERTIFICATE applies only with the
    opt-in, and gRPC never reads it; OTEL_EXPORTER_OTLP_CLIENT_CERTIFICATE and
    OTEL_EXPORTER_OTLP_CLIENT_KEY never apply, see
    observability);
  • the HACP ControlLink WebSocket, which carries the Platform token;
  • mobile Web Push deliveries, which carry this host's push authorization; and
  • the optional Qdrant memory vector backend, which sends HOLARYN_QDRANT_API_KEY (its
    start-up version check, which builds its own connection, runs only with the opt-in).

Instead of an ambient bundle, these clients verify servers against the public certificate
authorities bundled with Holaryn. When one of those variables is set, the host log names it once
(never its value) and says it was ignored. When a server's certificate was not issued by one of
the bundled authorities, the error says so and names HOLARYN_TRUST_ENV and this section
(SA-605): a model turn's failure (whether or not provider resilience is on), Scan for
models
, image generation, ChatGPT and GitHub Copilot sign-in (Settings and
holaryn connection login), the MCP connection test, tool discovery and a tool call, resource
read or prompt during a run for a static-header server, the web search providers, the Telegram
Test connection check and poll log, connected-app errors, the HACP connection log, webhook
delivery records, Web Push delivery errors, Qdrant errors, the Ollama memory embedder, the host
log for the OTLP HTTP exporters (unless HOLARYN_OTEL_CA_CERT_PATH is set, which is then what
failed), holaryn capability, and a note on the exception the Python SDK raises (SA-623 added
the sign-in, the run-time MCP, the embedder, the OTLP and the resilience-off model turn
cases). An expired certificate, a name mismatch, a refused connection, or any failure once the
opt-in is on is reported as before. The OTLP gRPC exporters' failures are logged by the
OpenTelemetry library itself, without this guidance. If your network requires a proxy or a private
certificate authority, set the environment variable HOLARYN_TRUST_ENV=1 for the Holaryn host,
then restart the host. Where to set it depends on how the host runs:

  • Desktop app, when it started the host itself: the host inherits the app's environment
    (the app adds only the host token and its version). On Windows, add HOLARYN_TRUST_ENV with
    the value 1 as a user environment variable (Settings, search for "environment variables",
    then Edit environment variables for your account), then choose Quit from the tray menu
    and start the app again; closing the window is not enough. On macOS and Linux, how the desktop
    session passes variables to apps is outside Holaryn and this guide does not cover it; starting
    the app's executable from a shell where the variable is exported passes it on. When the app
    attached to a host that was already running, set the variable for that host instead (below).
  • The service on Windows: from an elevated session, add the entry HOLARYN_TRUST_ENV=1 to
    the service's Environment value (REG_MULTI_SZ) under
    HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Services\holaryn, keeping its other entries
    (use the installed service name if you changed it), then run
    holaryn service stop --platform windows and holaryn service start --platform windows from
    the same elevated session. The same value carries HOLARYN_ENCRYPTION_PASSWORD_FILE; see the
    service runbook.
  • The service on Linux: add the line HOLARYN_TRUST_ENV=1 to the service's environment
    file (--env-file at install, /etc/holaryn/holaryn.env by default; it must stay root-owned,
    one plain NAME=value per line), then run sudo holaryn service stop --platform linux and
    sudo holaryn service start --platform linux. See "Service environment" in the
    service runbook.
  • The service on macOS: not covered. holaryn service install writes only the state
    directory into the LaunchDaemon's environment and takes no other variables, so there is no
    supported way to set this one yet.
  • A command-line run (holaryn serve, holaryn run): set it in the shell that starts the
    command, for example $env:HOLARYN_TRUST_ENV = "1" in PowerShell or
    export HOLARYN_TRUST_ENV=1 in a POSIX shell, then start the command again.

All of the traffic above then follows those
standard variables, including NO_PROXY; add localhost,127.0.0.1 to NO_PROXY if a local
Ollama or LM Studio should not go through the proxy. With the opt-in the HACP WebSocket also
trusts the operating system's certificate store again. HOLARYN_PROVIDER_TRUST_ENV=1, the
earlier provider-only name, still works as an alias with the same wider meaning; when both are
set, HOLARYN_TRUST_ENV decides (so HOLARYN_TRUST_ENV=0 turns the alias off). The setting is
read when a client is created, so changing it has no effect on a host that is already running.
The holaryn commands and the Python SDK read it from the environment of the process that runs
them, not from the host's.

The holaryn commands, holaryn capability and the Python SDK never send a request for this
machine through a proxy, whatever these variables and the opt-in say: that is the host on this
machine, and a proxy would see its token. This machine is localhost, any loopback address
however it is written (127.0.0.1, 127.1, ::1, an IPv4-mapped ::ffff:127.0.0.1, with or
without a trailing dot), or a name whose every address is loopback; a name that also resolves
to another address, or that does not resolve within two seconds, is treated like any other
host. Certificate checks are a separate decision: with HOLARYN_TRUST_ENV=1 these clients
trust SSL_CERT_FILE and SSL_CERT_DIR for this machine too, so a local TLS front end signed
by a private CA keeps working; without it they use the bundled authorities.

A few clients ignore those variables whatever this setting says, because they connect to
addresses Holaryn checks itself: lifecycle hook webhooks, peer and pairing requests, MCP sign-in
discovery and token requests, and a search endpoint whose policy lists trusted addresses.

Clients that carry no credential keep their library's defaults and follow those variables
whatever this setting says: the built-in web fetcher, the skill index, marketplace and model
download fetchers, and the local-model checks. A proxy in the environment therefore sees the
host names the web fetcher requests, but the fetcher refuses a response that arrives through a
proxy (it checks that it is connected to the address it validated), so web fetches do not work
through a proxy; list the hosts it should reach directly in NO_PROXY. The
data-flow model lists every client
in each group (SA-605).

No silently selected provider

If no connection and default model are configured, the agent fails with instructions rather than
silently assuming a vendor or model. Provider recovery operates only after a real primary model is
configured; it is not a built-in provider. If a run fails with "No runnable provider", the host's
environment is missing the connection's API key. For a Windows service, store the key with
Actions → Set API key; do not set it as a machine-wide environment variable, which every account on
the computer can read. See Troubleshooting.

  • Settings — the Settings page reference.
  • Prompt caching — modes, isolation, diagnostics, provider retention, and
    recovery.
  • Provider recovery — retries, fallback chains, circuits, replay safety,
    and diagnostics.
  • Web Interface — the model picker and @-mentions in chat.
  • Memory — embedding models (a separate choice from chat models).