actual-setup
GitHub配置 Actual Computer 作为 Hermes 推理提供商,支持云端中继和本地设备模式。
Trigger Scenarios
Install
npx skills add NousResearch/hermes-agent --skill actual-setup -g -y
SKILL.md
Frontmatter
{
"name": "actual-setup",
"author": "shl0ms + Hermes Agent",
"license": "MIT",
"version": "2.0.0",
"metadata": {
"hermes": {
"tags": [
"actual",
"actual-inc",
"provider",
"local-inference",
"relay",
"gguf",
"setup"
],
"category": "devops"
}
},
"platforms": [
"linux",
"macos",
"windows"
],
"description": "Set up Actual Computer (actual.inc) inference in Hermes."
}
Actual Computer Setup Skill
Sets up actual.inc (Actual Computer) as a Hermes inference
provider. Actual turns the user's own hardware into a private inference cluster
and exposes an OpenAI-compatible API two ways: a hosted end-to-end-encrypted
relay at https://api.actual.inc (authenticated with an ac_ key), and a local
on-device daemon at http://127.0.0.1:8080 (no auth on loopback). This skill
does not install the Actual daemon for the user — device authorization requires
a human in a browser.
When to Use
- User wants to add actual.inc as an inference provider (cloud relay or local).
- User has an
ac_key and wants Hermes routed through their Actual cluster. - User wants fully-local, on-device inference via the Actual daemon.
- Troubleshooting: Actual requests failing with cryptic 400s or empty streams.
Prerequisites
- Hermes has first-class
actualprovider support (provider idactual, aliasesactual-computer,actualcomputer,aci). Do NOT configure Actual as acustom_providers/providers.actual.*entry on current Hermes — the built-in provider owns the name and handles base-url normalization, the Responses transport, and local no-auth automatically. - Relay mode: an Actual account and an
ac_inference key from https://actual.inc/user/keys. - Local mode: the user has installed the daemon
(
curl -fsSL "https://actual.inc/install" | bash) and completed device authorization by runningactualonce and opening the printedhttps://actual.inc/device?code=...URL in a browser. Relay that URL to the user and WAIT — never invent an email or authorize on their behalf. Codes expire in 5 minutes; re-runactualfor a fresh one.
How to Run
Relay / API mode
- Put the key in
.env(secrets only — never config.yaml): appendACTUAL_API_KEY=ac_...to~/.hermes/.env. - Verify the key and discover models with
terminal:curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY" - Select provider + model:
hermes config set model.provider actual hermes config set model.default "MODEL_ID_FROM_DISCOVERY" - Verify end-to-end:
hermes chat -Q -q "Reply with exactly: ACTUAL_OK" --provider actual -m MODEL_ID
Local mode
- Human has installed + authorized the daemon (see Prerequisites).
- Download and load a model (scriptable once authorized):
actual models search "qwen2.5 0.5b instruct gguf" --limit 8 --no-prompt # Downloads REQUIRE an explicit quantization (409 ambiguous_model_download otherwise): actual models download "Qwen/Qwen2.5-0.5B-Instruct-GGUF/Q4_K_M" actual models list # note the INSTALLED name (differs from download id) actual models load "qwen2.5-0.5b-instruct-q4_k_m" # load by installed name - Point Hermes at the daemon.
ACTUAL_BASE_URLwith a loopback host flips the built-in provider into local no-auth mode automatically — no key needed: appendACTUAL_BASE_URL=http://127.0.0.1:8080to~/.hermes/.env, then:hermes config set model.provider actual hermes config set model.default "INSTALLED_MODEL_NAME" - Verify (reduced toolset — see context-window pitfall below):
hermes chat -Q -q "Reply with exactly: LOCAL_OK" --provider actual -m INSTALLED_NAME -t file,web
Quick Reference
| Thing | Value |
|---|---|
| Hosted relay | https://api.actual.inc/v1 (normalized from bare host automatically) |
| Local daemon | http://127.0.0.1:8080/v1 (no auth on loopback) |
| Key env var | ACTUAL_API_KEY (ac_...) |
| Base URL env var | ACTUAL_BASE_URL (loopback host ⇒ local no-auth mode) |
| Provider id / aliases | actual / actual-computer, actualcomputer, aci |
| Transport | Responses API (codex_responses) — built-in, do not override |
| Cluster pinning | X-Cluster-ID header via providers.actual.extra_headers in config.yaml |
| Model size guide | 0.5B Q4_K_M ~470MB (toy), 7-8B Q4_K_M ~4.5GB (daily driver), 32B ~20GB |
Pitfalls
- reasoning_effort trap (handled by Hermes since the first-class provider).
Actual's SGLang/vLLM backends accept only
none/low/medium/high/max;xhigh/ultraused to fail with a crypticExpecting value: line 1 column 1 (char 0)(a wrapped HTTP 400). The built-in provider clampsxhigh→highandultra→maxon the wire. If a request still 400s this way on an old Hermes, set a per-model cap:agent.reasoning_overrides.<model>: highin config.yaml. - Context-window overflow on small local models. Hermes' default toolset
is ~26k tokens of schemas plus a ~9k-token system prompt. A model loaded
with a 32k context overflows before the first turn, and llama.cpp-family
servers emit a bare
data: [DONE]— Hermes reportsProvider returned an empty stream with no finish_reason. This is NOT an SSE bug. Fixes: restrict tools (-t file,web), load the model with a largern_ctx, or pick a >=64k-context model for the full toolset. Upstream tracking: #51448 (do not file new issues; add evidence there). Related but distinct: #65631 (HTTP-200 SSE carrying a 400), #56516 (reasoning-only streams). - Download ids vs installed names.
actual models downloadtakesrepo/QUANTand 409s without an explicit quantization;actual models loadtakes the INSTALLED name fromactual models list. - Reasoning models returning empty content. GLM/Qwen reasoning variants
emit thinking in a separate
reasoningfield and can burn a smallmax_tokensentirely on reasoning. Give generous max_tokens before assuming failure. - Do not create a custom provider named
actual. Older setup guides (pre first-class support) wroteproviders.actual.*config blocks. On current Hermes the built-in provider wins the name; stale custom blocks are ignored or conflict. Remove them and use the env vars + model.provider flow above.
Verification
# Relay:
hermes chat -Q -q "Reply with exactly: ACTUAL_OK" --provider actual -m MODEL
# Local (small model — reduced toolset):
hermes chat -Q -q "Reply with exactly: LOCAL_OK" --provider actual -m MODEL -t file,web
# Provider status (local no-auth shows key_source=local-offline):
hermes status
For other OpenAI-compatible clients (e.g. OpenCode), see
references/opencode.md.
Version History
- 8430c1b Current 2026-08-20 06:04


