Hermes Agent: The Practitioner's Reference (2026)
# A practitioner's reference to Hermes Agent, Nous Research's open-source self-improving AI agent: provider auth, config files, the skills system, and running it as a multi-platform messaging gateway.
TL;DR: Hermes Agent is an open-source self-improving AI agent from Nous Research. It runs as a CLI and as a multi-platform messaging gateway, stores a durable identity and persistent memory on disk, aggregates skills that improve with use, and works with any OpenAI-compatible LLM provider — Nous Portal, OpenRouter, Anthropic, GitHub Copilot, z.ai, Kimi, MiniMax, DeepSeek, Qwen Cloud, Hugging Face, Google, xAI/SuperGrok, or your own self-hosted endpoint.1219 The current release is v0.21.5 (tag
v2026.9.24, September 24, 2026), a rollup patch on the v0.21 line; What’s New in v0.21.5 covers what changed, and the release sections after it run newest to oldest.52 The hardest part for most new users is provider authentication: Hermes ships 39 providers in its static picker list at tagv2026.9.24and auto-extends that list from 38 bundled provider plugins, plus custom endpoints, and three distinct auth paths (API key in.env, OAuth viahermes model, or custom endpoint inconfig.yaml).53 The auth model is the thing to learn first — everything else is downstream of which provider is resolved.
Hermes Agent operates as a full agent runtime, not a chat wrapper. It reads your filesystem, executes commands in sandboxed backends, scrapes the web, spawns subagents, runs scheduled cron jobs, talks to Telegram/Discord/Slack/WhatsApp/Signal/Email from a single gateway process, and creates its own skills from experience.1 The CLI is a terminal UI built on top of a conversation loop in run_agent.py; the gateway is a long-running process that routes messages from messaging platforms through the same conversation loop.3
The difference between casual and expert Hermes usage comes down to five systems. Master these and Hermes becomes a force multiplier:
- Provider resolution: how auth flows map to API calls
- Configuration hierarchy:
config.yaml+.env+auth.json+SOUL.md+AGENTS.md - Tool + toolset system: what the agent can do, gated per platform
- Skills system: procedural memory the agent creates and evolves
- Gateway + cron + profiles: running Hermes where you live, not just where you are
Key Takeaways
- Provider auth is three paths, not one. API key in
.env, OAuth viahermes model/hermes auth, or custom endpoint inconfig.yaml. Pick the path that matches your provider, not the one that feels familiar. - Switching providers is a single command.
hermes modelinteractively walks you through every supported provider including OAuth logins, and/model provider:modelswitches mid-session without losing history.2 - Two files are the user-editable config surface.
~/.hermes/config.yamlholds settings and~/.hermes/.envholds secrets.auth.json,SOUL.md,MEMORY.md, andskills/are managed by Hermes directly — you can editSOUL.mdby hand, but the rest is touched by the agent itself.4 - Hermes is the successor to OpenClaw. If you’re migrating,
hermes claw migrateimports 30+ categories of state automatically.5 - Auxiliary tasks ride your main model by default. Vision, approval classification, compression, and session titles run as separate “auxiliary” LLM calls, and at the tag
autoroutes every one of them to your main chat model – nothing to configure, but on expensive reasoning models these side tasks add meaningful cost. Point individualauxiliary.<task>slots at cheap fast models when that matters.434
Every section below is grounded in the upstream documentation at hermes-agent.nousresearch.com/docs and the source tree at github.com/NousResearch/hermes-agent. Every factual claim has a footnote pointing at the specific upstream page it came from.
Choose Your Path
| What you need | Go here |
|---|---|
| Install Hermes | Installation — one-line installer or manual steps |
| Sign into a provider | Authentication & Providers — the section you came here for |
| Switch models mid-session | The hermes auth Command and Custom & Self-Hosted Endpoints for /model syntax |
| Run a local LLM | Custom & Self-Hosted Endpoints — Ollama, vLLM, SGLang, llama.cpp, LM Studio |
| Connect messaging platforms | Messaging Gateway — Telegram, Discord, Slack, WhatsApp, Signal, Google Chat, LINE, SimpleX Chat, ntfy, Buzz (28 in the docs comparison table) |
| Write or install a skill | Skills System — progressive disclosure + skill hub |
| Deep reference for every CLI command | Keep reading — and link directly to CLI Commands |
How Hermes Works: The Mental Model
Hermes is structured around a single conversation loop that any entry point can invoke. The entry points are the CLI (cli.py), the messaging gateway (gateway/run.py), the ACP adapter for editor integration, the batch runner, and an API server.3 All of them ultimately call AIAgent.run_conversation() in run_agent.py, which:
- Builds the system prompt from
SOUL.md,MEMORY.md,USER.md, skills, context files, and tool guidance viaagent/prompt_builder.py(the September 2026 decomposition moved it into the newagent/package)3 - Resolves the runtime provider via
runtime_provider.py— this is the step that picks your auth, base URL, and API mode3 - Calls the provider using one of three API modes:
chat_completions,codex_responses, oranthropic_messages3 - Dispatches any returned tool calls through
model_tools.pyand the central tool registry (tools/registry.py)3 - Loops until the model produces a final response, then persists the session to SQLite with FTS53
Understanding this loop matters because every feature — personalities, memory, skills, compression, fallback — attaches to one of these stages. When you’re reading a config key and wondering what it does, the answer is usually “it’s a knob on stage 1, 2, 3, or 4 of the loop above.”
Platform-agnostic core. One AIAgent class serves CLI, gateway, ACP, batch, and API server. Platform differences live in the entry point, not in the agent itself.3 This is why the same slash commands work in the terminal and in Telegram — they’re dispatched from a shared COMMAND_REGISTRY in hermes_cli/commands.py.6
The directory structure is the system. Hermes stores everything under ~/.hermes/ (or $HERMES_HOME for non-default profiles):4
~/.hermes/
├── config.yaml # Settings (model, terminal, TTS, compression, etc.)
├── .env # API keys and secrets
├── auth.json # OAuth provider credentials (Nous Portal, Codex, Anthropic)
├── SOUL.md # Primary agent identity (slot #1 in system prompt)
├── memories/ # Persistent memory (MEMORY.md, USER.md)
├── skills/ # Bundled + agent-created + hub-installed skills
├── cron/ # Scheduled jobs
├── sessions/ # Gateway session state
└── logs/ # agent.log, gateway.log, errors.log (secrets auto-redacted)
Every file above has a specific role; none of them overlap. If you’re looking for “where does Hermes store X,” it’s one of these.
What’s New in v0.21.5 (The September 24 Rollup)
Hermes Agent v0.21.5 (tag v2026.9.24, September 24, 2026) is the current release, the third thin rollup on the v0.21.x line: “This tag rolls up the ~460 PRs merged since v0.21.4 into a stable tagged release for downstream consumers”. Curated notes are again deferred to v0.22.0. The operator-facing changes, checked in source at the tag:5253
- Hindsight memory leaves the core tree. The bundled Hindsight provider and the
hermes-agent[hindsight]pip extra are gone; Hindsight now installs from the plugin catalog, maintained by Vectorize; the release notes omit this. If your config setsmemory.provider: hindsight,hermes updateinstalls the catalog plugin into every profile home that names it, and the first agent start installs it if still missing, unlesssecurity.allow_lazy_installsisfalse(then runhermes plugins install hindsight). Settings,.envkeys, and memory data are untouched. See External Memory Providers. gateway.multiplex_profiles: falseis retired. The gateway rewrites it totruein place and prints a one-time boxed notice. A named profile that must keep its own gateway setsgateway.standalone: truein its ownconfig.yaml; to take one served profile offline,hermes -p <name> gateway stopnow parks it without stopping the host. See Messaging Gateway.- New models in the Nous and OpenRouter pickers: GPT-6 Sol and GPT-6 Luna (each with a
-provariant) and Claude Opus 5.5. - Plugin compat is unchanged:
plugins.allow_deprecated_imports: truestill works.
Updating: hermes update or the installer one-liner; Docker and Hermes Cloud images build from nousresearch/hermes-agent:v2026.9.24.52
What’s New in v0.21.4 (The Second Rollup Patch)
Hermes Agent v0.21.4 (tag v2026.9.21, September 21, 2026) is the second deliberately thin rollup on the v0.21.x line. Its own framing: “Patch release. This tag rolls up the ~1,800 PRs merged since v0.21.3 into a stable tagged release for downstream consumers (Docker images, Hermes Cloud, hosted deployments).” The window since v0.21.3: “5,071 non-merge commits” across “5,169 changed files”, “1,812 merged PRs”, and “2,116 closed issues”. By commit count that is the second-largest tag-to-tag window in the project’s history, behind only v0.21.1’s 5,139; by merged PRs it is the largest. The curated story is deferred again, verbatim: “Full curated release notes for this window ship with v0.22.0, which will document everything from v0.21.0 onward” and “Nothing in this window is skipped.” The note does list what it leaves undocumented, and each item below was checked in source at the tag; where an item changes a standing section of this guide, the pointer is inline.5051
- One gateway per host, and Desktop attaches instead of respawning. The ruling is one
hermes serveand onehermes gateway runper host per OS user, each multiplexing every profile – enforced by a new host-wide singleton layer: a host lock held for the lifetime of the winning process plus a rendezvous record carrying(pid, createTime), so a second invocation can prove the owner is the same live process and attach instead of binding a second port. Staleness is proved, never assumed, and the Desktop app runs the same idea from its side, attaching to the running host backend instead of spawning a second one. Messaging Gateway carries the mechanics.51 - Connectors become one backend-owned operation with one setup card. A
manage_connectionstool call drives a pure-data connection state machine on the backend, and Desktop, TUI, and CLI all render it as the same setup card: a field per missing credential, the card’s verb held until every required field has text. See the Nous free tier subsection.51 --format stream-json: machine-readable one-shot runs.hermes chat -q ... --format stream-jsonemits one JSON object per stdout line for CI runners and orchestrators: asystem/initevent, thentextdeltas /tool_use/tool_resultevents, then one terminalresultenvelope carrying exit code, final text, and token stats. Diagnostics and the session ID stay on stderr, tool output is capped at 5,000 characters per event, the flag implies--quiet, requires-qor--query-file(exit 2 without one), and cannot be combined with--tui. Thehermes chatoptions table carries the flag.51skills.auto_loadpins skills into every session. Names listed underskills.auto_loadinconfig.yamlload fully in every new session – CLI, TUI, gateway, cron, and API alike – resolved once when the prompt is first built. Skills System gains a Pinned Skills subsection.51- A
declineoption for unauthorized DMs.unauthorized_dm_behaviorgains a third value besidepairandignore:declinesends one polite refusal, then goes silent toward that sender for 24 hours. See User Authorization & Pairing.51 mcp.discovery_concurrencycaps MCP discovery connects. Default 4,0= unlimited; every configured server still connects, they just stop arriving all at once. See MCP Integration.51session_searchlearns time bounds and a recall retry. The discovery shape takesafter/beforebounds (ISO or relative durations like7d), and a zero-result multi-word query is retried matching ANY term instead of FTS5’s implicit all-terms AND. See Session Search.51hermes sessions set-journal-mode delete|wal: the offline journal converter. The self-service path for astate.dbstuck in the wrong SQLite journal mode – what previously required a hand-runPRAGMA journal_mode=DELETE. Stop the gateway, dashboard, and every CLI first: it refuses while any foreign process holds the database, flips without waiting out openers, and verifies the SQLite header bytes afterward. Windows has no holder scan, so there it refuses until you pass--forceafter stopping every Hermes process yourself.hermes doctornow points at it. Thehermes sessionsrow in Top-Level Commands carries it.51- Desktop: a font setting, one-click engine updates, and plugin uninstall. Settings gains a font-family field persisted per profile at
desktop.font_family– it overrides the theme’s sans face across chat and UI, the suggestion list leads with accessibility faces (OpenDyslexic, Atkinson Hyperlegible, Lexend), and empty means the theme decides. The managed local-model runtime shows an “Update engine” button when an update is available, with a failed update staying visible for a direct retry. And the Plugins hub uninstalls a plugin behind a confirm dialog.51 - Video catalogs: LTX 2.5 and Kling O3. The FAL video plugin adds
ltx-2.5(Lightricks’ open-source audio-video model: native audio, up to 20 s / 4K image-to-video, camera-motion presets, cheap tier – fal rejects clips over 10 s at 1440p/2160p) andkling-o3(Kuaishou’s frontier family, premium tier: multi-shot native storytelling, optional audio, 3-15 s).51 - The plugin catalog becomes a shipped directory and a website. The repo’s
plugin-catalog/directory grew from 9 entries atv2026.9.14to 228 at this tag – one reviewed YAML per plugin, pinned to an exact commit SHA – and the docs site now builds a page per plugin and per author, each rendering the plugin’s README from the pinned commit. The ten community plugins the release names are all catalog entries at the tag. Plugin System carries the details.51 - And a large fix run across profile/multiplex isolation, cron, kanban, Desktop, and
state.db, named by the release only as a category; the curated record of those fixes is v0.22.0’s to write.50
The plugin-compat revert still has not landed in this window. COMPAT_MANIFEST.md, compat_manifest.json, and the compat shims are all present at v2026.9.21; the only changes to hermes_cli/plugin_compat.py in the window are a process-wide scan cache (a multiplex gateway discovers plugins once per served profile, and re-parsing every plugin’s source cost ~0.4 s per profile on the boot path) and POSIX-form hit paths on Windows (#112576) – the date gate and the literal-boolean escape hatch are unchanged, so plugins.allow_deprecated_imports: true still keeps affected plugins loading. The compat box in Plugin System carries the current state.42
Updating is unchanged: hermes update (git installs) or the shell installer for fresh ones; Docker and Hermes Cloud images build from this tag (nousresearch/hermes-agent:v2026.9.21).50
What’s New in v0.21.3 (The September 14 Patch)
Hermes Agent v0.21.3 (tag v2026.9.14, September 14, 2026) is three more days of main under a two-item note, cut because Cloud agents auto-update to the newest release tag and the remote-gateway sign-in fixes needed to reach them. (1) Remote dashboard sessions no longer get revoked by refresh bursts: both refresh paths on the gateway (the cookie gate and the desktop’s native bearer route) now coalesce concurrent requests carrying the same rotating refresh token into a single flight, so a desktop wake burst can no longer replay an already-rotated token into the Portal’s reuse detection and revoke the whole session – and refresh runs off the event loop, so a slow identity provider no longer freezes /api/status. (2) Long-lived processes stop minting duplicate state.db writer handles: gateway, dashboard/desktop backend, ACP, and CLI readers attach read-only, and in-process writers share the registry handle. Window stats: “1,036 non-merge commits” across “2,642 changed files” and “338 merged PRs”. Everything else is deferred on purpose: the release names what it is not documenting (reasoning-effort pickers on every model picker, OpenRouter OAuth PKCE, HEIF/HEIC/AVIF decoding, a FAL catalog wave, Slack pasted tables and the Agent Sessions API, the state.db WAL refusal on cross-VM filesystems, and more) and commits to the curated record, verbatim: “Full curated release notes for this window ship with v0.22.0, which will document everything from v0.21.0 onward” and “Nothing in this window is skipped.”49
And the plugin-compat deadline arrived on schedule. The 2026-09-14 removal that v0.21.1 announced took effect as a date gate in the shipped code, not as a code revert: at tag v2026.9.14, hermes_cli/plugin_compat.py carries COMPAT_REMOVAL_DATE = 2026-09-14, and removal_in_effect() returns true from that date (or as soon as the manifest file is gone), so an affected external plugin is now disabled at load with the red notice. What has NOT happened yet: the revert that actually deletes the old import paths. COMPAT_MANIFEST.md, compat_manifest.json, and the compat shims are all still present at the tag – and still at v2026.9.21 and on main as of September 22 – which is why plugins.allow_deprecated_imports: true still works as an escape hatch: the old paths still resolve once the loader is allowed to proceed. Two sharp edges: the key must be a literal YAML boolean (the code accepts only the boolean True; a quoted string like "true" or "false" is treated as unset, per the source comment “Literal boolean only”), and the hatch dies the moment the revert lands, because the paths themselves disappear. The compat box in Plugin System carries the current state.4249
What’s New in v0.21.2 (The state.db Patch Release)
Hermes Agent v0.21.2 (tag v2026.9.11, September 11, 2026) exists first of all to fix what v0.21.0 broke. The release says it plainly: “v0.21.0 shipped a large rewrite of the session store’s connection handling, and for some installs it made state.db fragile: second writers cancelling each other’s locks, healthy databases reported as corrupt, one bad row killing sessions list. This release closes that class and rolls up everything else that landed on main in the four days since v0.21.1.” Unlike v0.21.1’s deliberately thin note, this one documents its own highlights; the window stats are “947 non-merge commits” across “1,869 changed files” and “312 merged PRs” from “140 contributors”, and the curated record of the whole v0.21.x stretch remains v0.22.0’s to write.48
- The state.db reliability campaign: six PRs, 44 issues closed. The release’s operator advice comes first, and it should be yours too: if your
state.dbbroke on 0.21.0 or 0.21.1, runhermes doctor– it now names structural damage and full-text-search index damage separately instead of calling everything “FTS write corruption”, and it points athermes sessions recover --inspect-only(offline, non-destructive, profile-pinned; reports canonical-table readability without creating an output database) when a rebuild is not enough. The root-cause work removes every second writer on the store: profile gateways were writing hosted-room state into the rootstate.dbevery 5 seconds, and that coordination now lives in a dedicatedshared-state.db, so profile gateways never open the master session store writable; the dashboard opens read-only first; cron’s lifecycle guard goes through the tracked connection registry instead of a rawopen()on a live database (which cancels the gateway’s POSIX locks – the classic corrupt-your-SQLite recipe); anddoctor --fixrefuses a checkpoint it cannot prove is safe. Beyond the writers: FTS-index damage now degrades search and rebuilds the index later instead of fail-closing the whole turn; one corrupt row no longer killssessions list, export, or insights (bad rows render?with a warning naming the session); sessions never bind to or read another profile’s database; and a read-only open no longer takes the write lock, so a one-shothermesbehind a busy gateway went from a 4-20 s stall-and-fail to 0.01 s.48 - Multi-profile isolation hardening. Since v0.19.0 this guide has repeated the project’s claim that profile routing gives each profile “fully isolated config, skills, memory, and secrets.” At v0.21.2 that promise had holes, now closed: secondary-profile bots inherited the default profile’s allow-lists, adapters could send credentials to the default profile’s host, stdio MCP servers received the default profile’s vault secrets,
MEDIA:delivery could attach another profile’s.env/auth.json/state.db, webhook and Feishu callbacks could drift off the routed profile, and secondary profiles could pick up a sibling’s Nous bearer from per-process memos (#107609-#107630). If you run several profiles under one multiplexed gateway, this is the release where the isolation claim actually holds.48 - The password-blind credential vault. The agent can now sign in, pay, and fill addresses in the browser from 1Password, Bitwarden, or the local Hermes vault without ever seeing a secret, and two-factor codes come from a saved authenticator key (TOTP: a base32 seed or an
otpauth://totpURI; counter-based HOTP is rejected) or are asked for in your UI. Every backend hands the agent the same shape – login metadata plus an opaque namespaced handle (vault_local,op:,bw:), with the password resolved at fill time only; external managers stay locked until you unlock them for the session, and the master password “is never a tool argument, never argv, never persisted.” This builds on v0.19.0’sSecretSourcework, which took provider API keys out of plaintext.env; the vault now does the same for the agent’s browsing credentials.48 - A curated, SHA-pinned plugin catalog.
hermes plugins browselists “every curated plugin catalog entry”,hermes plugins searchqueries the catalog, andhermes plugins installresolves catalog names alongside Git URLs andowner/repo.hermes plugins packrounds it out with “declarative, shareable plugin sets”: onehermes-pack.yamlpinning a set of plugins to exact commit SHAs,pack installfanning out to ordinary pinned installs with capability consent staying per-plugin,pack exportemitting a pack for the current install, andpack showas a dry run. The Plugin System command block carries the new subcommands.48 - Nous free tier and guided first launch. Fresh installs get free inference and connectors out of the box with one command to sign in,
/loginworks from inside a chat, and connector tools (Gmail, Linear, Notion, and the rest) are searchable throughtool_search. The desktop’s guided first launch sits behindHERMES_GUEST_ONBOARDING=1, and only the literal1enables it: the desktop’s own test asserts that'true','0', and an empty value all leave it off, and the launch decision is stamped into the spawned backend’s environment so an inherited value never leaks through. See Nous Tool Gateway.48 - Desktop backend spawn storms are over. Bot Mode used to spawn or dial one backend per profile on launch and on every roster tick, hovering the Bots roster spawned a backend per row, and a profile switch could spawn a duplicate primary. All of it is fixed in this window.48
Updating is unchanged: hermes update from an existing install, or the shell installer for fresh ones.48
What’s New in v0.21.1 (The Rollup Patch)
Hermes Agent v0.21.1 (tag v2026.9.7, September 7, 2026) is deliberately thin: a “Patch release” that “rolls up current main since v0.21.0 for tagged deployments and downstream consumers.” The body gives the window stats – “5,139 non-merge commits across 4,364 changed files (+601,014 / -768,419)” and “632 merged PRs” – and then defers the story: “Full curated release notes for this window will ship with v0.22.0.” That makes this the largest single tag-to-tag window in the project’s history (no previous adjacent-tag window exceeds the 2,790 non-merge commits of v2026.7.20..v2026.7.30), shipped under the smallest release note. Until v0.22.0 writes the curated record, the six clusters below are what changes how you operate Hermes at the tag, each verified in source.41
- The codebase was decomposed, and a plugin-compat clock runs out on September 14. The September 2026 decomposition (PR #102117) split the repo’s large modules into focused files: a new
agent/package holds the conversation loop’s internals (214 top-level modules plus seven subpackages;prompt_builder.pynow lives atagent/prompt_builder.py, whilerun_agent.pyremainsAIAgent’s home), the CLI’s subcommand parsers moved into ahermes_cli/subcommands/package (61 modules), the staticCANONICAL_PROVIDERSlist moved fromhermes_cli/models.pytohermes_cli/models_catalog_static.py(the list itself is unchanged: 39 static entries, still auto-extended from 39 bundled provider-plugin directories), anddelegate_tasksplit across roughly a dozentools/delegate_tool_*modules. Internal import paths are not a stable API, so a newCOMPAT_MANIFEST.mdat the repo root re-exports 1,148 “moved-lazy” public names from their old module paths, each emitting aHermesPluginCompatWarningonce per name per process when resolved. The compat layer was temporary, and its removal took effect on schedule on 2026-09-14, six days after the tag – not as a code revert but as a date gate already in the shipped code. Since that date an affected third-party plugin is not loaded: the CLI banner,hermes doctor, andhermes updateshow a red notice naming the disabled plugin, the desktop shows a one-time modal, andhermes plugins listshows the reason. If you run external plugins, check them:hermes plugins compat <path>prints everyfile:linewith old path -> new path and exits 1 while anything remains (--jsonfor machine-readable output; run it bare to scan the whole installed set). The escape hatch for a plugin whose author has not caught up isplugins.allow_deprecated_imports: trueinconfig.yaml, and it still works: the revert that actually deletes the old import paths still has not landed as of September 22 (verified through tagv2026.9.21and onmain), so the old paths still resolve once the loader is allowed to proceed. The aftermath detail is in What’s New in v0.21.2 above; the Plugin System section carries the current-state box.42 - Gateway conversations no longer rotate on timers, at all. The session-lifecycle doc now states the contract outright: “Inactivity and wall-clock time never rotate a conversation.
/newand/resetcreate an explicit boundary; context compression continues to manage long histories. Legacy timer configuration is ignored. The existingSessionResetPolicydatatype is inert compatibility data, not a runtime policy.” Boundaries are explicit and user-made; if you carried OpenClaw-era session-reset timers throughhermes claw migrate, they are now inert data. See Messaging Gateway.43 - MCP authorization grows a device-code path.
hermes mcp login <name>gains--flow {browser,device}:browseris the existing PKCE flow,deviceis an RFC 8628 device-code login for machines where a browser callback is impractical, and the flag overrides the server’soauth.flowconfig. The window also hardened the rest of the MCP auth surface: profile ownership is enforced throughout OAuth sessions, malformed OAuth metadata caches are ignored instead of wedging a server, and the desktop relays MCP OAuth through client-local callbacks. Related:-t/--toolsetsnow also filters which configured MCP servers are spawned at all, so a one-shothermes -z -t <toolsets>run skips cold-starting servers it does not need. The MCP Integration command block now carriesloginandreauth [--all], both of which predate this window but had not been documented here.44 - Delegation gets honest about background work. Six reliability changes to
delegate_task, all read from the delegate tool source at the tag. (1) A background batch returns as ONE completion by default; opting intodelegation.independent_completionssplits the call into completion units: tasks sharing agroupvalue join and report together, and each ungrouped task reports alone as it finishes. The default is deliberate – the source notes that a per-task flurry of completions “fragmented orchestrators that had no plan for it.” (2) A child’s background processes are killed at its teardown unless the child hands them to the parent withprocess_manage(action="handoff"); un-handed leftovers are named on the result asorphaned_processes, and child processes that exited without ever being read surface asunread_completionswith an output tail. The design stance, per the source docstring: the parent “must hear that from the runtime” rather than trust a child’s claimed “watcher running”. (3)delegation.fallback_providersbecomes a real config surface:nullinherits the parent chain for unpinned children,[]disables fallback, and a child pinned by provider, endpoint, or model gets no fallback unless this setting declares one. (4) The child fallback chain resolves through the canonical normalizer, so malformed entries are dropped instead of breaking a spawn. (5) A crash mid-unit no longer loses finished children: each finished child of a multi-child unit is durably recorded on the unit’s own row and survives into the recovered result as a partial. (6) Subagents never inherit the 1-hour prompt-cache tier: a delegated child is downgraded to the 5-minute tier, because the 1h tier is priced for a person who steps away, not a burst of parallel children.45 - Providers and models. The catalogs pick up OpenAI’s GPT-6 Astra and Astra Pro with
-fast/-flexspeed-tier variants (“2x price, priority tier” / “0.5x price, flex tier”) on the Nous Portal and OpenRouter catalogs. On the ChatGPT/Codex OAuth route, Astra is account-gated (only live account-scoped discovery may advertise it) and gains an opt-in-900kpicker variant that raises the advertised 272K window to a live-verified ~900K; the suffix never goes on the wire. Alongside it:anthropic/claude-fable-5.1,google/gemini-3.7-flashandgemini-3.8-flash,qwen/qwen3.8-max-0902andqwen3.8-flash, and Meta’s Muse Spark 1.3 family (1M context, contributor variants included) plus a Metamuse-imageimage-generation provider plugin. Tavily lands as a web search/extract backend (TAVILY_API_KEY; keyless works when Tavily is selected viahermes tools). A managed llama.cpp runtime makes local models a first-class path (official binaries, one supervisedllama-server, one-click setup from the desktop), and out-of-tree external-process providers get their own resolution branch. Operationally: the picker’s remote catalogs now refresh every 20 minutes (model_catalog.ttl_minutes, default 20; the legacyttl_hourskey is honored only where a user had set it explicitly).46 - Desktop: annotate the page, control the session. The in-app browser gains comment mode: click Annotate, then click any element (or drag a box) on the live page and type a note; saved comments stay as numbered pins and never send a turn by themselves. When you are done, “Add N comments” hands the batch to the composer with a cropped screenshot per pin, and each element comment carries its CSS selector, its markup, and the computed styles that matter for layout, so the agent finds the element in your source instead of guessing from a picture (password and hidden values, and key-shaped attributes, are redacted before markup leaves the page). Larger batches arrive grouped by page region, so twenty-odd comments become a handful of pieces of work that usually touch separate files – which is what makes fanning them out to parallel workers safe. Around it: structured session controls plus session automation controls, drag-to-create sessions, a session import view for foreign coding-agent transcripts,
display.resume_last_session(default true: cold start reopens the last chat or page), a first-open consent prompt offeringbrowser.use_real_profilewhen a Browser pane opens with it off, a built-in optional-skills catalog under Capabilities -> Skills with one-click install, and a new Russian desktop UI locale (the CLI’s 17 locale catalogs are unchanged).47
Updating is unchanged: hermes update from an existing install, or the shell installer for fresh ones. The curated record of this window is v0.22.0’s to write; the clusters above are what changes at the tag.41
What’s New in v0.21.0 (The Pantheon Release)
Hermes Agent v0.21.0 (tag v2026.8.31, August 31, 2026) is the current feature release, and it is the curated record the whole v0.20.x rollup train had been deferring to: “This release rolls up everything from the v0.20.1-v0.20.6 infrastructure patch tags - those windows are fully documented here.” The framing is the sequel to the Herald: “v0.20.0 made Hermes the herald - he spoke, and he carried word to other agents. In v0.21.0 the gods assemble.” The stats line, verbatim: “Since v0.20.0: ~5,800 commits · ~2,475 merged PRs · ~5,680 files changed · ~869,000 insertions · ~135,000 deletions · ~2,100 issues closed · 760+ contributors”.35
The release organizes itself by feature area; so does this summary. Where a feature changes a standing section of this guide, the pointer is inline.
- Bot Mode: your agents become a society, built in. Bot Mode graduates from the bundled
hermes-botsplugin of the v0.20.3 window to a default-on part of the desktop app. Every agent profile gets a name, a deterministic avatar face with randomize/lock controls, and a place in a shared roster; you create Discord-style group chats where multiple bots and you talk in one room, @-mention any bot from the composer, and give rooms names and pictures. The sub-catalog includes attributed agent-to-agent message cards, sender-side delivery notices, paint-first hydration for instant wakes, a Routines pane, and a rebuild on the app design system. The release’s own line: “Before, ‘multi-agent’ meant plumbing; now it looks like a chat app full of coworkers.”35 hermes peer: bot-to-bot DMs between your agents. Any Hermes agent can message any other by handle, across profiles and gateways, from the CLI or from inside a conversation – ask your research bot to hand findings to your coding bot and read the reply where you are. “Replies land in each agent’s canonical Bot Chat, so conversations between agents are durable and inspectable, not fire-and-forget.” Thehermes peercommand (add/list/remove peers,dm) is in the Top-Level Commands table.35- Cron jobs that remember. Scheduled jobs stop being goldfish: cron agents load and update persistent memory like every other agent,
continuity=truecarries each run’s output into the next so a monitor can dedupe against what it already reported, every job gets a durable notepad scratchpad, monitor-mode jobs skip the LLM entirely when nothing changed, reasoning effort can be pinned per job, and cron output can land in a bot’s canonical Bot Chat, where the bot actually responds. The Scheduled Tasks (Cron) section now documents the at-tag mechanics of each.3537 - Live subagent orchestration.
delegate_taskgains control actions: list running children, steer one mid-flight with a course correction, or stop it early and keep the partial result. Child outputs can be validated against a JSON schema, per-delegation cost is surfaced in results, and the defaults are raised to 250 iterations per subagent and 10 concurrent children (the unified caps that replacedmax_async_childrenin v0.19.0).3538 - The MCP command center. MCP servers and the catalog merge into one desktop page with drag-in “paste anything” import, background health checks that nudge you to re-auth before a tool call fails, a fleet cost/usage overlay (schema token estimates, 30-day usage per server), and
hermes://deep links that install an MCP server with explicit confirmation. See MCP Integration.35 - A CLI power wave. Ctrl+P opens a fuzzy command palette (also reachable as
/palette), the/modelpicker filters as you type,/statusshows reasoning mode, pending approvals, and context usage, and the status bar can display live cache-hit %, latency, and tokens/sec with per-field toggles. Plus a global emergency stop, session pin/unpin, rotating composer placeholders, and terminal pets. One naming note: the release bills a dry-run approval checker ashermes approval-check; at the tag that surface ishermes approvals test– noapproval-checksubcommand exists.3540 - The agent drives the desktop’s browser. The in-app browser stops being a window the agent can only look at: Hermes navigates, clicks, and reads it directly, and pages pop out to your system browser with full link context menus.35
- Six new providers and a model catalog wave. Meta Model API (Muse Spark), CommandCode, Tencent TokenPlan, Nebius Token Factory, Ramp Router, and Actual Computer. Three of those (Meta AI, CommandCode, Actual Computer) landed inside earlier rollup windows and were already in the provider matrix; the matrix now adds Tencent TokenPlan, Nebius Token Factory, and Ramp Router, plus the docs’ new Alibaba Token Plan SKU. The catalogs pick up GLM-5.3-Flash, qwen3.8-max/flash, Gemini 3.7 Flash, MiniMax M3 free, Nemotron 3.5 Lightning, and Muse Spark 1.2. Two structural changes ride along:
model_overridesinconfig.yamllets you patch any model’s context window or pricing without waiting on a release, and providers can now ship as pip-installed packages discovered via entry points; a unified selection-guard registry also warns across every picker surface when a model trains on your data.3536 - Security hardening across the board. Writes to protected agent-instruction files (AGENTS.md, skills, memory stores) now always require approval so a prompt-injected agent cannot quietly rewrite its own standing orders; a deep redaction sweep closes secret-leak gaps across terminal errors,
.envreads, checkpoints, and ACP logs; the approval system learns Windows destructive commands; macOS permission grants survive updates via a stable TCC signing identity (hermes desktop --setup-tcc-identity); and the Blender MCP catalog entry and skill were removed after an upstream compromise. See Security Hardening for the at-tag details.3539 - Gateway maturation. Slack gets native live cards (real streaming replies plus opt-in plan/task cards) and outbound link-preview suppression; Telegram gains an inline picker that makes every command and skill searchable via @botname, bypassing Telegram’s command-menu cap; the relay lane matures (native plugin initialization, live-card ops with draft streaming, session-span segmentation, voice-note STT restored); a gateway control socket lets fleet consumers query the gateway and lets updaters pause it gracefully instead of tree-killing it; and the turn-reaper captures wedged worker stacks when the watchdog fires.35
- A skills wave. Eight salvaged productivity skills (document-to-action-items, meeting-action-items, email-inbox-triage, github-issue-to-pr, weekly-review-planning, competitor-news-monitor, product-price-monitor, social-media-content-calendar), HAR-derived API clients (“watch a website once, then call its hidden API directly with no browser”), publish-site, session-librarian, blocked-page-recovery, merge-reconciler, plan-interrogation, and an advisory SKILL.md linter on create.35
Reverted in this window (not shipping): Model Council mode (/council) and the DCP context engine both landed and were reverted; the WS-only gateway server (#94245) was merged then reverted (#96118), so FastAPI remains on the desktop boot path – but the seq-stamped event replay (#94219), lossless desktop reconnect over WebSocket, DID ship. Electron rolled back to 40.10.2. If community coverage of the v0.20.x windows mentioned any of these, they are not in this release.35
Updating is unchanged: hermes update from an existing install, or the shell installer for fresh ones. The rollup subsection under the Herald release below remains the per-tag record of which window landed what.35
What’s New in v0.20.0 (The Herald Release)
Hermes Agent v0.20.0 (tag v2026.8.3, August 3, 2026) was the feature release before v0.21.0 rolled the whole train up; v0.20.1 (August 13) and v0.20.2 (August 16) are stabilization tags on top of it, and v0.20.3 (tag v2026.8.16.2, published August 17), v0.20.4 (tag v2026.8.18, August 18), v0.20.5 (tag v2026.8.19, published August 21), and v0.20.6 (tag v2026.8.27, August 27) keep the rollup train running with feature-bearing windows of their own — see the subsection below. The window since v0.19.0 covers roughly 3,650 commits, 1,400 merged PRs, and 1,200 issues closed across 650+ contributors.55
Three changes break what earlier versions of this guide told you to do. Read these before anything else:
- Node 26 is now required. The installer pins
NODE_VERSION="26"and refuses older runtimes with “Node.js … is too old (Hermes requires Node >=26).” Installers,heal, andupgradeall enforce it. Note the docs site’s installation page still says Node v22 — the installer script and the release notes are the newer, authoritative sources.55 - pip and Homebrew are retired, not merely deprecated. Verbatim: “brew + pip/PyPI wheel channels retired (shell installer / Docker / Nix are the supported channels).” If you are still on a pip or brew install, that path no longer receives releases.55
- The default tool-calling iteration limit moved from 90 to 500. Long autonomous runs stopped hitting an artificial wall, and every budget-pressure threshold below is computed against the new ceiling.
read_filealso defaults to 2,000 lines instead of 500.55
The rest of the release:
- Conversational voice. Streaming TTS with barge-in and on-device wake words.55
- A2A v1.0. An agent-to-agent protocol plugin, closing the long-running request in issue #514.55
- Signed outbound webhooks.
hermes webhookwas inbound-only before; v0.20.0 adds HMAC-signed outbound lifecycle webhooks for session, turn, and tool events.55 - Grounded citations. A new skill with a fact-checking mode.55
- Power-user CLI wave.
!commandruns a shell command instantly without spending a model turn;/initscans the project and writes or updates anAGENTS.md;/diffshows staged, all, or session changes from any surface;/contextbreaks down what is filling the context window;/focusgives a reduced-output view with hidden-line recovery; Ctrl+S stashes a half-written prompt.hermes import-agentmigrates a Claude Code or Codex CLI setup in one command.55 - Secrets surface additions. A command-helper secret source that composes with every vault, one-command token rotation with actionable startup errors, an opt-in encrypted break-glass cache for Bitwarden, vault-injected keys scoped per profile home, and
${env:VAR}SecretRef parity betweenconfig.yamland MCP config. The three-path auth model below is unchanged.55 - Faster warm start.
hermes -wcold start dropped from roughly 14s to 1.8s.55 - Desktop became a platform. Artifacts with versioned cards and sandboxed live preview, a Plugin SDK with Kanban as the founding desktop plugin, a quick-entry global hotkey, multiple GUI windows, SSH remote-backend mode, and RFC 8252 native sign-in.55
The v0.20.3, v0.20.4, v0.20.5, and v0.20.6 rollups (August 17-27)
The project’s cadence after v0.20.0 was high-frequency tagged rollups, and they were more than stabilization. All four deferred curated notes the same way – each stated that “full curated release notes for this window will ship with v0.21.0” – and v0.21.0 has now delivered them: What’s New in v0.21.0 above is the curated record of this whole stretch (the release confirms the windows “are fully documented here”). The blocks below stay as the contemporaneous per-tag record – they preserve which tag landed what, which the curated notes fold together.54233035
v0.20.3 (~250 commits, ~125 PRs since v0.20.2):
- MCP 2.x SDK migration, with 2026-07-28 stateless-protocol support. Hermes moves to the current MCP SDK generation and speaks the stateless revision of the protocol.54
- Bot Mode ships as a bundled plugin (
hermes-bots) carrying the core teammate protocol.54 - A CommandCode provider plugin joins the provider catalog.54
- Cua Driver 0.20 runtime contracts for computer use, plus subprocess Python runtime ownership hardening (PYTHONHOME/PYTHONPATH isolation).54
- Reliability: cron scheduler self-heal (EMFILE recovery, stale-claim reconciliation, wedged-job re-arm), session-handoff data-loss fixes, desktop remote-gateway connection self-healing, and a wave of ecosystem ports (plugin install security scanning,
/worktree,/rollbackhand-edit preservation, UTF-16 file reads).54
v0.20.4 (~146 commits, ~74 PRs since v0.20.3):
- The desktop’s glass surface: matte glass/translucency work with a frost picker and macOS pre-select.54
- A tabbed SESSIONS|BOTS sidebar with per-bot hide/unhide, plus Bot Mode group-chat fixes (long-running member turns, Markdown rendering, cross-machine routing).54
- NVIDIA SkillEvaluator Tier 1 advisory scanning on skill installs - license and security checks run when you install a skill.54
- Cron media-send hardening (configurable timeout, manual-run attachments, missed-fire surfacing), SessionDB event-loop-thread and contention fixes,
hermes updateparked-branch honesty, and kanban native OS notifications.54
v0.20.5 (~746 commits, ~323 PRs since v0.20.4):
- The keyless web tier: web search works on fresh installs with zero API keys - a 5-vendor free rotation with ring failover.23
- A CLI polish wave: fuzzy
/modelpicker, a Ctrl+P command palette, and a richer/status.23 - Bot Mode grows up: group-room threads, foldable conversation summaries, blob-face avatars, and PDF/file attachments with drag & drop.23
- Fleet and worktree tooling:
hermes updatereceipts,hermes update --plan(the release notes call it “fleet--planverification”; the flag lives onhermes update, not on afleetcommand, and per its parser help it will “Show the update plan and exit without changing anything”), andhermes worktree list/prune.2327 - Cron jobs gain persistent memory and per-job reasoning effort, plus execution-discipline and runtime stall guards driven by the Composio eval findings, multi-question clarify, the opencode-free zero-auth provider, and desktop perf work (paint-first Bot Mode hydration, React Compiler in both renderers).23
v0.20.6 (tag v2026.8.27, August 27 – ~1,313 commits, ~525 PRs since v0.20.5):
The release describes itself outright as a “Patch release. This tag rolls up the ~525 PRs merged since v0.20.5 into a stable tagged release for downstream consumers”, and its own description puts the window at ~1,313 commits across ~1,557 files (+177,113 / −21,682) – ~525 merged PRs.30 The headline items from that description:
- Consent-gated real-profile browsing – local browsing can use your default Chromium profile, with a Windows close-with-approval flow.30
- The desktop Browser gets its own OS window, plus a managed SSH remote-update engine and a fleet profile rail.30
- Remote MCP catalog expansion: 50+ live-verified vendor-hosted servers, including Cloudflare, Grafana Cloud, Better Stack, and Railway.30
- Opt-in OS-keychain encryption for stored secrets – no more per-launch macOS Keychain prompts.30
- New models across the pickers: GLM-5.3-Flash, MiniMax M3 free, and MiniMax H3 Max video.30
- TTL result caching for
web_search/web_extractand multi-querytool_searchwith stemming.30 - Lean-tail compression is now the default – the Context Compression section below documents the at-tag config.3031
- Terminal backends became pluggable – see Terminal Backends.3032
- Updater and fleet honesty: updaters pause gateways over the control socket instead of tree-killing them, and image/package-managed installs refuse unsafe in-place updates.30
- Cron durable-incident acks with clearer code-skew failures, Slack link-unfurl controls, and shared Docker container identities.30
Updating is unchanged: hermes update from an existing install, or the shell installer for fresh ones.542330
What’s New in v0.19.0 (The Quicksilver Release)
Hermes Agent v0.19.0 (tag v2026.7.20, July 20, 2026) is named for the messenger god’s own speed: the spine of the release is raw responsiveness, with an ~80% cut in first-turn time-to-first-token on every platform. Around that spine sit terminal billing, password-manager secret sources, smart approvals by default, observable subagents, and crash-proof response delivery. The window since v0.18.0 is the project’s largest yet: ~2,245 commits, ~1,065 merged PRs, ~3,300 issues closed, and 450+ community contributors.56
- ~80% faster to first token, everywhere. Cold submit→dispatch dropped from ~4.3s to ~0.9s across the CLI, gateway, TUI, desktop, and cron alike — Discord capability detection moved off the critical path, the Ollama probe is skipped for known non-Ollama providers, and blocking work was removed from agent init. Perceived latency got its own pass: reasoning models now stream their thinking live by default (
display.show_reasoningis ON), and the response box paints per token instead of per line.56 - Desktop and TUI rendering waves. The desktop app took a ~20-PR speed overhaul: 14× less CPU in the streaming-markdown splitter via incremental block lexing, virtualized review-pane diffs, snappy session switching on large transcripts, and no more per-token sidebar and tool-row re-renders. The TUI now renders streamed markdown incrementally per block.56
- pip and Homebrew installs are deprecated. Both paths were flagged as “unsupported legacy” installs, with removal planned. That removal has since happened in v0.20.0 — the brew and pip/PyPI wheel channels are retired, leaving the shell installer, Docker, and Nix as the supported channels.5655
- Secrets can come from your password manager. A new pluggable
SecretSourceinterface fetches secrets from Bitwarden and 1Password (op://references) at load time, with multiple vaults enabled simultaneously, deterministic precedence, conflict warnings, and per-variable provenance — API keys no longer have to live in a plaintext.env. Future vault providers drop in as plugins.56 - Smart approvals are now the default. When Hermes wants to run a flagged command, an independent LLM reviewer assesses it instead of prompting you for every one — and each verdict covers only that exact command. User-defined deny rules block matching commands even under YOLO mode,
/deny <reason>relays your refusal reason so the agent course-corrects, and the pluginpre_tool_callapprove action (re-landed with rule keys) escalates a tool call to a human gate.56 - Terminal billing:
/subscriptionand/topup. Manage your Nous Portal plan without leaving the terminal — see your plan and remaining allowance, preview exactly what an upgrade costs or when a downgrade takes effect, and apply it with undo. The desktop app gains a matching billing settings tab.56 - Watch subagents work; never lose a finished answer.
delegate_taskdispatches return live transcript files you cantail -fthe moment subagents launch — every tool call, result, and streamed reply, one human-readable log per child. Background delegation completions are durable across restarts, and final gateway responses are recorded in a delivery-obligation ledger instate.dband redelivered on the next boot if the gateway dies mid-send. Themax_async_childrenconfig option is deprecated in favor of unified delegation concurrency caps.56 - One gateway, many profiles. A single multiplexed gateway sharing one bot token can route specific guilds, channels, or threads to different profiles — each with fully isolated config, skills, memory, and secrets — with a
GATEWAY_MULTIPLEX_PROFILESoverride. The routing index moved intostate.db;sessions.jsonis now an optional legacy mirror.56 - Providers and models wave. Fireworks AI lands first-class (cost estimation, #2 slot in the provider picker), joined by DeepInfra and Upstage Solar. The catalogs add GPT-5.6 (Sol/Terra/Luna + Pro, wired end-to-end), grok-4.5 (GA), kimi-k3 (kimi-k2.x retired), and fully wired Claude Sonnet 5. A per-provider
enabled: falseflag andexcluded_providersconfig scrub unused providers from/modelpickers and resolution.56 - Reasoning effort becomes a dial. New
maxandultraeffort tiers land across every surface, with per-model overrides in config, per-slot effort in MoA presets (advisors think hard, the synthesizer stays fast), per-task effort for auxiliary models, and a session-scoped/reasoningin the CLI.56 - CLI and MCP surface.
hermes sessions exportwrites Markdown, Quarto, HTML, prompt-only, and Hugging Face trace formats with an opt-in--redactscrub;/model --oncegives a one-turn model override; slash-skill invocations stack (/skill-a /skill-b do XYZ);--safe-modeaids troubleshooting;hermes config get/unsetround out config management;hermes servebecomes a true headless backend; and MCP tools adopt themcp__server__toolnaming convention.56
If you are upgrading from v0.18.x, two changes deserve attention before anything else: a pip or Homebrew install now warns as unsupported legacy (migrate to the one-line installer), and max_async_children is deprecated in favor of the unified delegation concurrency caps. Everything else is additive — the headline reasons to upgrade are the ~80% first-turn latency cut, smart approvals, and the delivery ledger that makes finished answers crash-proof.
What’s New in v0.18.0 (The Judgment Release)
Hermes Agent v0.18.0 (tag v2026.7.1, July 1, 2026) is named for judgment: the agent verifying its own work instead of claiming success, and ensemble reasoning you can actually inspect. It also closes the entire P0/P1 backlog — roughly 692 highest-priority items resolved in twelve days.22
- Mixture-of-Agents as a first-class model. MoA is now selectable like any other model across all interfaces, and ensemble reasoning is visible: each reference model’s full output renders as its own labelled block with live answer streaming — you can watch the ensemble think instead of getting an opaque merged answer.22
- Completion contracts for
/goal. The agent verifies its own work by running the project’s checks before reporting a goal complete, rather than claiming success — judgment applied to itself.22 /learn— describe anything into a skill. Turn a workflow into a reusable skill by describing it; generated skills comply with the repo’s CONTRIBUTING.md conventions automatically.22/journeytimeline. A visual history of memory and skills over time, with editing, plus a memory graph on desktop.22- Background subagent fan-out. Delegate multiple tasks that run concurrently without blocking the conversation — the v0.17.0 single background subagent becomes a fleet.22
- Desktop Projects. First-class coding Projects with a project/repo/lane organization model.22
- Scale-to-zero gateway. Gateways can go dormant when idle and coordinate drains for seamless deployments — meaningful for anyone running Hermes as an always-on service.22
- Google Vertex AI support. Gemini access through GCP service accounts with automatic OAuth2 token refresh, joining the provider catalog.22
/prompteditor command. Opens$EDITORfor composing multi-line prompts instead of fighting the input line.22
If you are upgrading from v0.17.x, nothing here breaks the CLI. The headline reasons to upgrade are completion contracts (goals that verify themselves), first-class MoA with inspectable ensembles, and /learn for skill capture.
What’s New in v0.17.0 (The Reach Release)
Hermes Agent v0.17.0 (tag v2026.6.19, June 19, 2026) is named for how far the agent now reaches — new messaging channels, new model providers, and deeper desktop and dashboard control. It is additive over v0.16.x; the CLI surface is unchanged.21
- New messaging channels. iMessage now works without a Mac relay via Photon Spectrum (device-code OAuth,
hermes photon login); the WhatsApp Business Cloud API is an official Meta adapter that replaces the bridge-process requirement; SimpleX gains groups, native attachments, text batching, and auto-accept; and Raft joins as a bundled platform plugin with a privacy-by-contract wake-channel design.21 - New models and providers. The catalog adds
z-ai/glm-5.2(1M context),anthropic/claude-fable-5,laguna-m.1,nemotron-3-ultra, andgrok-composer-2.5-fast(Cursor’s model via xAI OAuth, 200k context). The xAI default moved togrok-build-0.1, and Anthropic adaptive models now follow the modern thinking contract (they never send areasoningfield).21 - Desktop and dashboard. Desktop adds background subagents with live “watch-windows” streaming delegated activity (
delegate_task(background=true)), a Composer model selector, rebindable keyboard shortcuts, native OS notifications, per-thread composer drafts, VS Code Marketplace themes, and Japanese + Traditional Chinese UI. The dashboard adds a full profile builder (model/skills/MCPs without editingconfig.yaml), a global profile switcher, a reworked Skills Hub with a security scan, Automation Blueprints (parameterized templates across form, slash command, conversation, and docs), and a secure login that returns 401 behind the OAuth gate.21 - Skills and tools.
image_generatecan now edit and transform a source image, not only create one from scratch, across every supported image provider; thememorytool gained anoperationsarray for atomic batch add/replace/remove in a single call; a newsimplify-codeskill runs a parallel three-agent review-and-cleanup pass gated by a Chesterton’s-Fence risk tier; and a booleanwrite_approvalreplaces the tri-statewrite_mode.21 - Architecture. Background subagents return a handle immediately and re-enter their result as a new turn; an MCP elicitation handler allows mid-tool-call confirmation, and late-connecting MCP tools are exposed between turns (cache-safe); cron becomes a pluggable CronScheduler with a Chronos managed-cron provider; and a new Managed scope (
/etc/hermes) lets an administrator pin user-immutable config, alongside a Gateway-Gateway relay for multi-gateway topologies.21 - New commands.
/version,/billing(interactive terminal billing),hermes photon login(iMessage auth), andhermes curator run --consolidate— consolidation is now opt-in, so routine background curation costs zero tokens.21 - Security. v0.17.0 closes a shell-escape denylist bypass, fails closed on missing approval modules and own-policy gateway adapters, sanitizes the environment for cron job-script subprocesses, redacts secrets in request debug dumps, screens MCP stdio configs for exfil patterns, and bumps urllib3 and PyJWT to clear CVEs.21
If you are upgrading from v0.16.x, nothing here breaks the CLI; it is new channels, models, and surfaces around the same agent. The relay-free iMessage and official WhatsApp adapters and the administrator Managed scope are the headline reasons to upgrade.
What’s New in v0.16.0 (The Surface Release)
Hermes Agent v0.16.0 (tag v2026.6.5, June 5, 2026) is named for the new surfaces it puts in front of the CLI-first agent. The headline is that Hermes is no longer terminal-only.20
- Native desktop app. Hermes Desktop is a new Electron app for macOS, Linux, and Windows with one-click install and in-app self-update. It gives you a streaming chat window, drag-and-drop files, clipboard image paste, a
Cmd+Kpalette, a session list with archive and search, and a status-bar model picker. It can connect to a remote Hermes gateway over a secure WebSocket, authenticating via OAuth or username/password, with per-profile remote hosts and concurrent multi-profile sessions linked by cross-profile@sessionreferences. The desktop UI also ships a full Simplified Chinese (简体中文) translation through a typed i18n layer (display.language; English remains the default).20 - Browser admin panel. The local web dashboard graduated from a status view into a full administration panel: an MCP catalog with enable/disable toggles, credential management, webhook and hook creation, memory configuration, gateway controls, and a System page with check-before-update plus one-click Debug Share. A new Channels page configures every gateway messaging platform (Telegram, Discord, Slack, and the rest) from the browser. Auth is now pluggable: username/password login, a generic self-hosted OIDC provider,
hermes dashboard registerfor a self-hosted OAuth client, and refresh-token session rotation.20 - New CLI and slash commands.
/undo [N]backs up the last N user turns with prefill and soft-delete, and works in the CLI, the TUI, and across messaging platforms. A configurable default interface (clivstui) lands with a--clioverride; the TUI gains a unified/modelcommand and a Sessions overlay.hermes portalis a human-readable alias for the Nous Portal onboarding flow, with new Quick Setup vs Full Setup first-run paths, and two diagnostics arrive:hermes prompt-sizeandhermes sessions optimize.20 - New models and providers. The picker adds
deepseek-v4-flash,MiniMax-M3(1M context, native MiniMax providers),qwen3.7-plus(Nous + OpenRouter), andgemini-3.5-flash(Gemini OAuth + API key). A first-class xAI Grok OAuth provider joins the desktop launcher, the model picker became fuzzy across every surface, multi-endpoint providers group under one row, and catalog refresh moved from daily to hourly.20 (v0.21.1 has since moved the cadence again: the picker’s remote catalogs refresh every 20 minutes,model_catalog.ttl_minutes.46) - Leaner skills and progressive disclosure. The default skill set dropped redundant and dead skills (Spotify moved to a native plugin, Linear to
hermes mcp install linear, and several stale entries were removed), moved more into optional, and added anenvironments:frontmatter relevance gate (kanban/docker/s6) that keeps context-specific skills out of the index until requested.NVIDIA/skillsis now a default trusted Skills Hub tap alongside OpenAI, Anthropic, and HuggingFace. MCP and plugin tools gained progressive (scoped) tool disclosure, and an MCP bug that reported false OAuth success when no token was obtained is fixed.20 - Security. v0.16.0 pins patched Starlette (≥1.0.1) for CVE-2026-48710 (BadHost), moves SSRF URL checks off the event loop in async paths, strips the Bedrock inference bearer token from subprocess env, adds
bws_cache.jsonto the file-safety read guard, addsdocker restart/stop/killto the dangerous-pattern list, and sanitizes invisible-unicode in vetted skill content. The release closed 2 P0 and 62 P1 issues, 16 of them security-tagged.20
If you are upgrading from v0.15.x, none of this is a breaking change to the CLI itself; it is additive surfaces and providers around the same agent. The desktop app and admin panel are the reason to upgrade if you want to run Hermes for non-terminal users or administer a remote gateway from a browser.
What’s New in v0.14.0 (The Foundation Release)
v0.14.0 is less about one headline feature and more about reducing setup drag while widening where Hermes can run.19 The main operational changes:
- Install and startup are lighter.
pip install hermes-agentworks from PyPI, heavy adapters lazy-install on first use, and the launch path defers enough work to cut cold start by roughly 19 seconds. (v0.19.0 has since deprecated pip installs — see Installation.) - Subscriptions can become local API endpoints.
hermes proxyturns OAuth-backed providers such as Claude Pro, ChatGPT Pro, and SuperGrok into an OpenAI-compatible local endpoint for tools like Codex, Aider, Cline, and Continue. - Gateway reach expands. LINE and SimpleX Chat join the gateway (the docs’ platform comparison table lists 28 platforms at tag
v2026.8.31; the Messaging Gateway section states what that number counts), Microsoft Teams is wired end-to-end, Discord history backfill is on by default, and Telegram/Discordclarifyprompts now use native buttons. - Write-time verification improves. After edits, Hermes can surface per-turn file-mutation summaries and language-server semantic diagnostics before the next turn, which moves it closer to evidence-driven agent work.
- Desktop and media tooling broaden.
computer_useworks through cua-driver for non-Anthropic providers,video_generateis unified behind pluggable backends, andvision_analyzesends raw pixels to models that can actually see.
Installation
The one-line installer is the supported install path. It handles Python, uv, Node.js, ripgrep, ffmpeg, the repo clone, the virtual environment, and the global hermes command.7
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
pip and Homebrew installs are deprecated as of v0.19.0. The PyPI package that v0.14.0 introduced (
pip install hermes-agent) and the Homebrew formula are now flagged as “unsupported legacy” installs — every surface warns (without blocking) when Hermes detects one, and removal of PyPI/Homebrew publishing is planned. If you installed via pip or brew, migrate to the installer above.56
Works on Linux, macOS, WSL2, and Android/Termux (the installer auto-detects Termux and switches to a tested Android bundle).7 Native Windows is now a Tier 1 platform rather than the early beta v0.14.0 shipped — install it with iex (irm https://hermes-agent.nousresearch.com/install.ps1). One caveat the docs are explicit about: macOS support is Apple Silicon only, and Intel Macs are unsupported.55 Historically, v0.14.0 added native Windows support in early beta via a PowerShell installer, but WSL2 remains the safer recommendation for production use until the Windows path matures.19
After it finishes:
source ~/.bashrc # or ~/.zshrc
hermes # Start chatting
The only prerequisite is git. The installer auto-provisions Python 3.11 via uv (no sudo required), Node.js 26 (for browser automation and the WhatsApp bridge), ripgrep, and ffmpeg. Node 26 is a hard floor as of v0.20.0 — the installer refuses anything older and installs a Hermes-managed Node instead.557
Verify the install
hermes --version # Check version (global flag; there is no `version` subcommand)
hermes doctor # Diagnose config/dependency issues
hermes status # Show current configuration + auth state
hermes dump # Copy-pasteable setup summary for debugging
hermes doctor tells you exactly what’s missing and how to fix it.7 hermes dump is the diagnostic command to paste into a GitHub issue or Discord thread when asking for help — it’s a plain-text summary of your entire setup with secrets redacted.8
Manual installation
If you need full control — custom Python version, specific extras, Nix/NixOS integration — the manual flow is documented step-by-step in the upstream installation guide.7 Key optional extras you can combine with uv pip install -e ".[<extras>]":
| Extra | What it adds |
|---|---|
all |
Everything below |
messaging |
Telegram & Discord gateway |
cron |
Cron expression parsing |
cli |
Terminal menu UI for setup wizard |
modal |
Modal cloud execution backend |
voice |
CLI microphone input + audio playback |
tts-premium |
ElevenLabs premium voices |
honcho |
AI-native memory (Honcho integration) |
mcp |
Model Context Protocol support |
homeassistant |
Home Assistant integration |
acp |
ACP editor integration support |
slack |
Slack messaging |
pty |
PTY terminal support (interactive CLI tools) |
dev |
pytest & test utilities |
termux |
Tested Android bundle (includes cron, cli, pty, mcp, honcho, acp) |
Termux install command is different — it uses pip with a constraints file, not uv pip:
python -m pip install -e ".[termux]" -c constraints-termux.txt
This is because .[all] on Android pulls faster-whisper via the voice extra, which depends on ctranslate2 wheels that aren’t published for Android.7
Authentication & Providers
At tag v2026.9.7, hermes_cli/models_catalog_static.py carries 39 static CANONICAL_PROVIDERS entries at line 311 (the September 2026 decomposition moved the list out of hermes_cli/models.py; the entries themselves are unchanged, v0.21.0’s tencent-tokenplan included) and auto-extends that list from plugins/model-providers/ (39 bundled directories; nine of them – actual, alibaba-coding-plan, commandcode, deepinfra, meta-ai, nebius-token-factory, opencode-free, router, upstage – have no static entry and are absorbed by the auto-extension because they use the default api_key auth type). The docs providers page tables 45 cloud and subscription providers plus a custom-endpoint row, and separately walks through local and self-hosted servers (Ollama, vLLM, SGLang, llama.cpp, LM Studio, LiteLLM, ClawRouter, and any OpenAI-compatible endpoint).26 At tag v2026.9.24 the static list still holds 39 entries, but plugins/model-providers/ holds 38 directories: the keyless opencode-free provider was removed on September 18 because OpenCode’s free tier now rejects anonymous traffic from outside its own client.53 Custom endpoints and three distinct auth paths sit on top of that. Here is the whole auth surface, organized by path so you can find the one that matches what you have.
The Three Auth Paths
Every provider in Hermes fits into one of three authentication patterns:
Path 1 — API key in .env. Put your key in ~/.hermes/.env and Hermes reads it on startup. Used by OpenRouter, AI Gateway, z.ai/GLM, Kimi/Moonshot, MiniMax (and MiniMax China), Alibaba Cloud/DashScope, Kilo Code, OpenCode Zen, OpenCode Go, DeepSeek, Hugging Face, Google/Gemini, and most third-party providers.2 Since v0.19.0, the key no longer has to sit in a plaintext file: a pluggable SecretSource interface can fetch secrets from Bitwarden or 1Password (op:// references) at load time, with multiple vaults enabled simultaneously, deterministic precedence, conflict warnings, and per-variable provenance — .env remains the fallback. (This is distinct from v0.15.0’s Bitwarden Secrets Manager bootstrap token, which consolidated provider keys behind one token; SecretSource replaces the plaintext file itself, and future vault providers drop in as plugins.)56
Path 2 — OAuth via hermes model or hermes auth. Launches a device code flow, opens a browser, stores credentials in ~/.hermes/auth.json (and can import existing credentials from tools like Claude Code or Codex CLI). Used by Nous Portal, OpenAI Codex (ChatGPT account), GitHub Copilot, and Anthropic (Claude Pro/Max).2
Path 3 — Custom endpoint in config.yaml. For any OpenAI-compatible API — Ollama, vLLM, SGLang, llama.cpp, LM Studio, LiteLLM proxy, Together AI, Groq, Azure OpenAI, or your own self-hosted server. Configured once via hermes model → Custom endpoint, then persisted to config.yaml.2
The Full Provider Matrix
This matrix covers the providers the docs page tables, with the exact setup flow for each; the picker at the tag lists more than the docs do (see the count above), and anything OpenAI-compatible fits through the custom-endpoint row.226
| Provider | Auth path | Setup |
|---|---|---|
| Nous Portal | OAuth | hermes model (OAuth login, subscription-based) |
| OpenAI Codex | OAuth | hermes model (ChatGPT device code, uses Codex models) |
| GitHub Copilot | OAuth or token | hermes model (OAuth device code), or COPILOT_GITHUB_TOKEN / GH_TOKEN / gh auth token |
| GitHub Copilot ACP | Local subprocess | hermes model (requires copilot CLI in PATH + copilot login) |
| Anthropic | OAuth or API key | hermes model (prefers Claude Code credentials), or ANTHROPIC_API_KEY, or ANTHROPIC_TOKEN setup-token |
| OpenRouter | API key | OPENROUTER_API_KEY in ~/.hermes/.env |
| AI Gateway (Vercel) | API key | AI_GATEWAY_API_KEY in ~/.hermes/.env (provider: ai-gateway) |
| z.ai / GLM (ZhipuAI) | API key | GLM_API_KEY in ~/.hermes/.env (provider: zai) |
| Kimi / Moonshot | API key | KIMI_API_KEY in ~/.hermes/.env (provider: kimi-coding). v0.19.0 adds kimi-k3 to the catalogs (kimi-k2.x retired).56 |
| MiniMax (global) | API key | MINIMAX_API_KEY in ~/.hermes/.env (provider: minimax) |
| MiniMax China | API key | MINIMAX_CN_API_KEY in ~/.hermes/.env (provider: minimax-cn) |
| Alibaba Cloud (Qwen) | API key | DASHSCOPE_API_KEY in ~/.hermes/.env (provider: alibaba, aliases: dashscope, qwen) |
| Kilo Code | API key | KILOCODE_API_KEY in ~/.hermes/.env (provider: kilocode) |
| OpenCode Zen | API key | OPENCODE_ZEN_API_KEY in ~/.hermes/.env (provider: opencode-zen) |
| OpenCode Go | API key | OPENCODE_GO_API_KEY in ~/.hermes/.env (provider: opencode-go) |
| DeepSeek | API key | DEEPSEEK_API_KEY in ~/.hermes/.env (provider: deepseek) |
| Hugging Face | API key | HF_TOKEN in ~/.hermes/.env (provider: huggingface, alias: hf) |
| Google / Gemini | API key | GOOGLE_API_KEY or GEMINI_API_KEY in ~/.hermes/.env (provider: gemini) |
| Fireworks AI | API key | First-class provider with cost estimation and cached price columns in the model picker; promoted to the #2 slot in provider pickers. New in v0.19.0.56 |
| DeepInfra | API key | First-class provider with a hardened integration. New in v0.19.0.56 |
| Upstage Solar | API key | First-class provider. New in v0.19.0.56 |
| xAI (Grok) | Native provider / SuperGrok OAuth | First-class provider with direct API access and model catalog (v0.9.0+). v0.14.0 adds SuperGrok OAuth and bumps grok-4.3 to a 1M context window for entitled accounts.21619 v0.17.0 adds grok-composer-2.5-fast (Cursor’s model via xAI OAuth, 200k context) and changes the xAI default to grok-build-0.1.21 v0.19.0 moves grok-4.5 to GA in the catalog.56 |
| xAI Custom Voices | API key | TTS provider with voice cloning. New in v0.13.0; configure under tts: in config.yaml and supply the xAI key in .env.18 |
| Xiaomi MiMo | Native provider | First-class provider with setup wizard and model catalog. Free MiMo v2 Pro on Nous Portal for auxiliary tasks (v0.9.0+).1615 |
| Google AI Studio | API key | GOOGLE_API_KEY or GEMINI_API_KEY in ~/.hermes/.env. Direct Gemini access with auto-detected context lengths via models.dev registry (v0.8.0+).15 |
| Qwen OAuth (Portal) | OAuth | hermes model → “Qwen OAuth (Portal)” (provider: qwen-oauth; browser PKCE login that reuses a local Qwen CLI login). OAuth provider with portal request support (v0.8.0+). The API-key DashScope path above was renamed from Alibaba Cloud to Qwen Cloud in v0.14.0; existing config keys continue to work.151926 |
| OpenCode Free (removed) | Keyless | Removed on September 18, 2026, and absent from tags v2026.9.21 and v2026.9.24. A config that still names opencode-free, free, or opencode_free gets an error naming the removal; switch to opencode-zen (pay-as-you-go, OPENCODE_ZEN_API_KEY) or opencode-go (subscription, OPENCODE_GO_API_KEY) via hermes model. The provider had landed in the v0.20.5 window as “the opencode-free zero-auth provider”.2353 |
| OpenAI API (direct) | API key | OPENAI_API_KEY in ~/.hermes/.env (provider: openai-api, optional OPENAI_BASE_URL)26 |
| Google Vertex AI | OAuth2 / ADC | hermes model → “Google Vertex AI” (provider: vertex; OAuth2 via service-account JSON or Application Default Credentials, billed to your GCP project)26 |
| Azure AI Foundry | Endpoint + key | hermes model → “Azure AI Foundry” (provider: azure-foundry; uses your Azure OpenAI / Foundry endpoint and key; the picker describes it as “OpenAI-style or Anthropic-style endpoint”)26 |
| AWS Bedrock | AWS credentials | hermes model → “AWS Bedrock” (provider: bedrock; standard AWS credentials chain via boto3, IAM or API key; Claude, Nova, Llama, DeepSeek)26 |
| NVIDIA NIM / Build | API key | NVIDIA_API_KEY in ~/.hermes/.env (provider: nvidia; NIM-hosted Nemotron and other models on build.nvidia.com, or a local NIM endpoint via base-URL override)26 |
| Ollama Cloud | OAuth or API key | hermes model → “Ollama Cloud” (provider: ollama-cloud; paste OLLAMA_API_KEY and pick from discovered cloud-hosted models)26 |
| StepFun Step Plan | API key | STEPFUN_API_KEY in ~/.hermes/.env (provider: stepfun; agent and coding models via the Step Plan API)26 |
| MiniMax (OAuth) | OAuth | hermes model → “MiniMax (OAuth)” (provider: minimax-oauth; browser PKCE login for the Coding Plan, global or CN region)26 |
| Meta AI | API key | MODEL_API_KEY in ~/.hermes/.env (provider: meta-ai; Meta Model API, Muse Spark family)26 |
| NovitaAI | API key | NOVITA_API_KEY in ~/.hermes/.env (provider: novita; 200+ models, Model API, Agent Sandbox, GPU Cloud)26 |
| Arcee AI | API key | ARCEEAI_API_KEY in ~/.hermes/.env (provider: arcee; aliases: arcee-ai, arceeai; Trinity models)26 |
| GMI Cloud | API key | GMI_API_KEY in ~/.hermes/.env (provider: gmi; aliases: gmi-cloud, gmicloud). Use the exact model ID returned by GMI’s /v1/models endpoint26 |
| Actual Computer | API key or local daemon | ACTUAL_API_KEY in ~/.hermes/.env for the hosted relay, or ACTUAL_BASE_URL=http://127.0.0.1:8080 for the local daemon with no key on loopback (provider: actual; aliases: actual-computer, actualcomputer, aci)26 |
| Tencent TokenHub | API key | TOKENHUB_API_KEY in ~/.hermes/.env (provider: tencent-tokenhub; aliases: tencent, tokenhub, tencentmaas; Hy3 Preview)26 |
| CommandCode | API key | COMMANDCODE_API_KEY in ~/.hermes/.env (provider: commandcode, alias commandcode-chat; Claude models via commandcode-anthropic, alias commandcode-claude). Works with the GOAT/Pro/Max/Provider plans, not the $1 Go plan, which has no API access. The plugin landed in the v0.20.3 window.5426 |
| Alibaba Cloud (Coding Plan) | API key | DASHSCOPE_API_KEY (provider: alibaba-coding-plan, alias alibaba_coding); a separate billing SKU on a different endpoint from the alibaba DashScope provider26 |
| Tencent TokenPlan | API key | TOKENPLAN_API_KEY in ~/.hermes/.env (provider: tencent-tokenplan, aliases: tokenplan, tencent-lkeap; Hy4 preview via the Anthropic Messages endpoint at api.lkeap.cloud.tencent.com). New in v0.21.0; the picker folds TokenHub and TokenPlan under one display-only “Tencent Hy” row3536 |
| Nebius Token Factory | API key | NEBIUS_API_KEY in ~/.hermes/.env (provider: nebius-token-factory; aliases: nebius, nebius-tf, tokenfactory). New in v0.21.03536 |
| Ramp Router | API key | RAMP_ROUTER_API_KEY in ~/.hermes/.env (provider: router; aliases: ramp-router, ramp, router.com). Ramp’s OpenAI Responses-native LLM gateway at api.router.com with a live account-scoped catalog – valid model IDs are whatever your key’s /v1/models returns, so the picker fetches them instead of hardcoding. New in v0.21.03536 |
| Alibaba Cloud (Token Plan) | API key | ALIBABA_TOKEN_PLAN_API_KEY in ~/.hermes/.env (provider: alibaba-token-plan; mainland-China endpoint: alibaba-token-plan-cn) – the Model Studio flat-token tier, a third Alibaba SKU beside alibaba and alibaba-coding-plan36 |
| Custom endpoint | config.yaml | hermes model → “Custom endpoint” (saved in config.yaml). The docs’ OpenAI-compatible list includes Together AI, Groq, Cerebras (https://api.cerebras.ai/v1), Mistral, Azure OpenAI, LocalAI, and Jan26 |
As of v0.19.0 you can also scrub providers you don’t use: a per-provider enabled: false flag and an excluded_providers config key remove them from /model pickers and built-in provider resolution.56
Anthropic: Three Auth Methods
Anthropic gets its own section because Hermes supports three distinct paths into Claude, and picking the right one matters. From the upstream docs:2
# Method 1: API key (pay-per-token)
export ANTHROPIC_API_KEY=***
hermes chat --provider anthropic --model claude-sonnet-4-6
# Method 2: OAuth through hermes model (preferred)
# Uses Claude Code's credential store when available
hermes model
# Method 3: Manual setup-token (fallback/legacy)
export ANTHROPIC_TOKEN=***
hermes chat --provider anthropic
# Auto-detect Claude Code credentials
hermes chat --provider anthropic # reads Claude Code files automatically
When you choose Anthropic OAuth through hermes model, Hermes prefers Claude Code’s own credential store over copying the token into ~/.hermes/.env. That keeps refreshable Claude credentials refreshable.2 If you already use Claude Code on the same machine, this is the cleanest path.
To pin Anthropic permanently in config.yaml:
model:
provider: "anthropic"
default: "claude-sonnet-4-6"
--provider claude and --provider claude-code also work as shorthand for --provider anthropic.2
GitHub Copilot: Two Modes
Copilot is supported in two modes: direct Copilot API (recommended) and Copilot ACP (which spawns the local Copilot CLI as a subprocess).2
# Direct Copilot API
hermes chat --provider copilot --model gpt-5.4
# Copilot ACP (requires the Copilot CLI in PATH + an existing copilot login)
hermes chat --provider copilot-acp --model copilot-acp
Authentication is checked in this order, per the upstream docs:2
1. COPILOT_GITHUB_TOKEN environment variable
2. GH_TOKEN environment variable
3. GITHUB_TOKEN environment variable
4. gh auth token CLI fallback
5. OAuth device code login via hermes model
Token type matters. The Copilot API does not support classic Personal Access Tokens (ghp_*). Supported types are OAuth tokens (gho_*), fine-grained PATs (github_pat_* with Copilot Requests permission), and GitHub App tokens (ghu_*). If your gh auth token returns a ghp_* token, use hermes model to authenticate via OAuth instead.2
Chinese AI Providers (First-Class Support)
Hermes has built-in support for z.ai/GLM, Kimi/Moonshot, MiniMax (global + China endpoints), and Alibaba Cloud with dedicated provider IDs.2
# z.ai / ZhipuAI GLM
hermes chat --provider zai --model glm-5 # Requires: GLM_API_KEY
# Kimi / Moonshot AI
hermes chat --provider kimi-coding --model kimi-for-coding # Requires: KIMI_API_KEY
# MiniMax (global)
hermes chat --provider minimax --model MiniMax-M2.7 # Requires: MINIMAX_API_KEY
# MiniMax (China)
hermes chat --provider minimax-cn --model MiniMax-M2.7 # Requires: MINIMAX_CN_API_KEY
# Alibaba Cloud / DashScope (Qwen)
hermes chat --provider alibaba --model qwen3.5-plus # Requires: DASHSCOPE_API_KEY
Base URLs can be overridden with GLM_BASE_URL, KIMI_BASE_URL, MINIMAX_BASE_URL, MINIMAX_CN_BASE_URL, or DASHSCOPE_BASE_URL environment variables.2
Z.AI auto-detects the endpoint. When using the z.ai/GLM provider, Hermes probes multiple endpoints (global, China, coding variants) to find one that accepts your API key. The working endpoint is cached automatically — no GLM_BASE_URL needed for most users.2
xAI (Grok) automatically enables prompt caching. When the base URL contains x.ai, Hermes sends the x-grok-conv-id header with every request to route to the same server within a conversation session, reusing cached system prompts and history.2 Automatic; no config needed.
The hermes auth Command
hermes auth is the credential management command for pools and OAuth credentials.6
hermes auth # Interactive wizard
hermes auth list # Show all credential pools
hermes auth list openrouter # Show one provider's pool
hermes auth add openrouter --api-key sk-or-v1-xxx
hermes auth add anthropic --type oauth
hermes auth remove openrouter 2 # Remove by index
hermes auth reset openrouter # Clear cooldowns
Credential pools are how you rotate multiple API keys or OAuth tokens for the same provider — useful for distributing rate limits across multiple keys without changing code.6 The legacy hermes login / hermes logout commands have been removed; use hermes auth instead.6
Custom & Self-Hosted Endpoints
Hermes works with any OpenAI-compatible API endpoint. If a server implements /v1/chat/completions, you can point Hermes at it.2
Interactive setup (recommended):
hermes model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter: API base URL, API key, Model name
Manual config.yaml:
model:
default: your-model-name
provider: custom
base_url: http://localhost:8000/v1
api_key: your-key-or-leave-empty-for-local
Both approaches persist to config.yaml, which is the single source of truth for main-model, provider, and base URL.2 The legacy env vars OPENAI_BASE_URL and LLM_MODEL are no longer read for main-model configuration — use hermes model or edit config.yaml directly.2 (OPENAI_BASE_URL + OPENAI_API_KEY are still honored as a fallback for the auxiliary provider: "main" routing path, so don’t delete them blindly if you’re using them there.)4
Switching custom endpoints mid-session:
/model custom:qwen-2.5 # Custom endpoint with explicit model
/model custom # Auto-detect the model from the endpoint
/model custom:local:qwen-2.5 # Named custom provider "local"
/model custom:work:llama3 # Named custom provider "work"
/model openrouter:claude-sonnet-4 # Back to a cloud provider
/model custom (bare, no model name) queries your endpoint’s /v1/models API and auto-selects the model if exactly one is loaded — useful for local servers running a single model.2
Local LLM Servers (Setup Templates)
The upstream docs have full setup guides for Ollama, vLLM, SGLang, llama.cpp, and LM Studio. Here are the key commands you’ll actually run. Each is designed to produce a working endpoint that Hermes can point at.2
Ollama — easiest local path, zero config:
ollama pull qwen2.5-coder:32b
OLLAMA_CONTEXT_LENGTH=32768 ollama serve # Raise from 4k default
hermes model # Custom endpoint → http://localhost:11434/v1 → qwen2.5-coder:32b
Critical Ollama gotcha: Ollama defaults to very low context lengths (4,096 tokens under 24GB VRAM). You must raise it via OLLAMA_CONTEXT_LENGTH or a Modelfile — the OpenAI-compatible API does not accept context length from the client, so Hermes cannot set it for you.2 For agent use, set at least 16k–32k.
vLLM — high-performance GPU serving:
pip install vllm
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--port 8000 \
--max-model-len 65536 \
--tensor-parallel-size 2 \
--enable-auto-tool-choice \
--tool-call-parser hermes
Tool calling requires --enable-auto-tool-choice and --tool-call-parser <name>. Supported parsers: hermes (Qwen 2.5, Hermes 2/3), llama3_json, mistral, deepseek_v3, deepseek_v31, xlam, pythonic. Without these flags, tool calls will come back as plain text.2
SGLang — fast serving with RadixAttention for KV cache reuse:
pip install "sglang[all]"
python -m sglang.launch_server \
--model meta-llama/Llama-3.1-70B-Instruct \
--port 30000 \
--context-length 65536 \
--tp 2 \
--tool-call-parser qwen
SGLang gotcha: Default max_tokens is 128. Set --default-max-tokens on the server or configure model.max_tokens in config.yaml if responses get cut off.2
llama.cpp / llama-server — CPU and Apple Silicon Metal:
./build/bin/llama-server \
--jinja -fa \
-c 32768 \
-ngl 99 \
-m models/qwen2.5-coder-32b-instruct-Q4_K_M.gguf \
--port 8080 --host 0.0.0.0
--jinja is required for tool calling. Without it, llama-server ignores the tools parameter entirely and the model tries to call tools by writing JSON in its response text — which Hermes can’t parse as actual tool calls.2
LM Studio — desktop app with GUI:
Start the server from the LM Studio app (Developer tab → Start Server), or via CLI: lms server start (starts on port 1234) and lms load qwen2.5-coder --context-length 32768.2 Then point hermes model at http://localhost:1234/v1.
Critical LM Studio gotcha: LM Studio reads context length from model metadata, but many GGUF models report 2048 or 4096 defaults. Always set context length explicitly in the LM Studio model settings — click the gear icon next to the model picker, set “Context Length” to at least 16384 (preferably 32768), and reload the model.2
Named Custom Providers
If you work with multiple custom endpoints (a local dev server and a remote GPU server, for example), define them as named custom providers in config.yaml:2
custom_providers:
- name: local
base_url: http://localhost:8080/v1
# api_key omitted — Hermes uses "no-key-required" for keyless local servers
- name: work
base_url: https://gpu-server.internal.corp/v1
api_key: corp-api-key
api_mode: chat_completions # optional, auto-detected from URL
- name: anthropic-proxy
base_url: https://proxy.example.com/anthropic
api_key: proxy-key
api_mode: anthropic_messages # for Anthropic-compatible proxies
Then switch between them mid-session with the triple syntax:
/model custom:local:qwen-2.5
/model custom:work:llama3-70b
/model custom:anthropic-proxy:claude-sonnet-4
You can also select named custom providers from the interactive hermes model menu.2
Pluggable Provider Architecture (v0.13.0+)
v0.13.0 ships a ProviderProfile ABC plus a plugins/model-providers/ directory so third-party inference providers can drop in without core modifications.18 If a provider speaks an OpenAI-, Anthropic-, or Codex-compatible API mode, you can implement a ProviderProfile subclass that declares the auth path, base URL, model catalog, and caching headers; Hermes resolves it through the same runtime_provider.py path the built-in providers use. This is the architectural change behind the v0.13.0 provider expansion: instead of editing core code to add a provider, you ship a plugin.
OpenAI-Compatible Local Proxy (v0.14.0+)
hermes proxy exposes an OpenAI-compatible local endpoint backed by the OAuth provider Hermes is already signed into — Claude Pro, ChatGPT Pro, SuperGrok, or another compatible configured provider.19 That means tools expecting an OpenAI-style API, including Codex CLI, Aider, Cline, Continue, or custom scripts, can reuse your subscription-backed Hermes auth without a separate API key. Treat the proxy as local developer infrastructure: bind it intentionally, do not expose it broadly, and keep provider-specific terms in mind.
Context Length Detection
Two settings get confused constantly, per the upstream docs:2
context_length— the total context window (combined input + output token budget, e.g. 1,000,000 for Claude Opus 4.7 or 200,000 for Sonnet 4.6). Hermes uses this to decide when to compress history.model.max_tokens— the output cap (max tokens the model may generate in a single response). Unrelated to history length.
Set context_length when auto-detection gets the window size wrong:
model:
default: "qwen3.5:9b"
base_url: "http://localhost:8080/v1"
context_length: 131072 # tokens
Hermes uses a multi-source resolution chain to detect context windows: config override → custom provider per-model → persistent cache → endpoint /models → Anthropic /v1/models → OpenRouter API → Nous Portal → models.dev (community-maintained registry for 3800+ models) → fallback defaults (128K).2 The system is provider-aware, so the same model can have different context limits depending on who serves it (e.g., claude-opus-4.6 is 1M on Anthropic direct but 128K on GitHub Copilot).2
Smart Model Routing: Provider Rotation & Fallback
Hermes does not pin you to one model on one provider. Smart model routing is the set of mechanisms that decide which provider and model actually serves a given request: credential pools spread load across keys, a configured fallback takes over when the primary fails, and the auxiliary slots below route side tasks to cheaper models independently of your main model.26 Configure these three together – they are the difference between an agent that stalls on a rate limit and one that keeps working.
Credential pools. When you have multiple API keys for the same provider, configure a rotation strategy via hermes auth. This is how you distribute rate limits across multiple keys.6
Fallback model. Configure a backup provider:model that Hermes switches to automatically when your primary model fails (rate limits, server errors, auth failures):2
fallback_model:
provider: openrouter # required
model: anthropic/claude-sonnet-4 # required
# base_url: http://localhost:8000/v1 # optional, for custom endpoints
# api_key_env: MY_CUSTOM_KEY # optional, env var name
The fallback swaps model and provider mid-session without losing your conversation. It fires at most once per session.2 Supported providers for fallback: openrouter, nous, openai-codex, copilot, copilot-acp, anthropic, huggingface, zai, kimi-coding, minimax, minimax-cn, deepseek, ai-gateway, opencode-zen, opencode-go, kilocode, alibaba, custom.2
Auxiliary Models
Hermes uses “auxiliary” models for side tasks: image analysis (vision), dangerous-command approval classification, context compression, session-title generation, TTS audio-tag insertion, skill matching, MCP tool dispatch, and the Kanban specifier/decomposer family.434 By default (auxiliary.*.provider: "auto"), every auxiliary task runs on your main chat model – the same provider/model you picked in hermes model. The docs are explicit that this replaced the old cheap-provider auto-detection: “Earlier builds split aggregator users (OpenRouter, Nous Portal) onto a cheap provider-side default. That was surprising … auto now uses the main model for everyone, and per-task overrides in config.yaml still win.”34 Nothing needs configuring to get started; the tradeoff is cost – on expensive reasoning models, auxiliary tasks add meaningful spend, so point individual tasks at cheap fast models when that matters.
Two former auxiliary tasks no longer use an LLM at all: web extraction (“web_extract and browser snapshots truncate long content deterministically and store the full text for read_file paging – no LLM involved”) and session search (the single-shape tool returns DB content directly). Their old auxiliary.web_extract.* and auxiliary.session_search.* blocks are gone from the defaults – leftover values in an existing config.yaml are “harmless leftovers and ignored” – and the flush_memories slot is likewise absent from the defaults at the tag.34
You can configure which model and provider each auxiliary task uses. Every auxiliary slot uses the same knobs: provider, model, base_url (plus api_key, timeout, extra_body, and a per-task reasoning_effort).434
auxiliary:
vision: # vision_analyze + browser screenshots
provider: "auto" # "auto" (= main model), "openrouter", "nous", "main", etc.
model: "" # e.g. "openai/gpt-4o", "google/gemini-2.5-flash"
base_url: "" # Custom OpenAI-compatible endpoint
api_key: "" # Falls back to OPENAI_API_KEY
timeout: 120
download_timeout: 30
approval: # dangerous-command approval classifier
provider: "auto"
model: ""
timeout: 30
compression: # summarizer -- legacy compression.summary_* keys migrate here (config v17)
provider: "auto"
model: ""
base_url: ""
timeout: 120
title_generation: # auto-generated session titles after the first exchange
enabled: true # set false to disable auto-titles
provider: "auto"
model: ""
language: "" # empty follows the conversation; e.g. "English" pins titles to one language
tts_audio_tags: { provider: "auto", model: "" } # Gemini 3.1 TTS hidden audio-tag insertion
skills_hub: { provider: "auto", model: "" } # skill matching and search
mcp: { provider: "auto", model: "" } # MCP tool dispatch
triage_specifier: { provider: "auto", model: "" } # hermes kanban specify: rough one-liner into a concrete spec, promoted to todo
kanban_decomposer: { provider: "auto", model: "" } # hermes kanban decompose: triage task into a graph of child tasks routed to specialist profiles
profile_describer: { provider: "auto", model: "" } # hermes profile describe --auto: 1-2 sentence profile descriptions
More specialist slots follow the same shape – goal_judge (judges /goal contract satisfaction), curator (the skill-usage review fork), background_review (the post-turn self-improvement fork), review (the /review reviewer subagent), moa_reference and moa_aggregator (Mixture-of-Agents), memory_query_rewrite, and monitor – and you don’t have to hand-edit YAML at all: run hermes model and pick “Configure auxiliary models” for a per-task interactive picker.34
The "main" provider option means “use whatever provider my main agent uses” – valid only inside auxiliary:, compression:, and primary fallback entries (fallback_providers:, or legacy fallback_model:). It is not valid for your top-level model.provider setting. If you use a custom OpenAI-compatible endpoint as your main model, set provider: custom in your model: section.4
Why this matters: because auto already rides your main model, the old “configure OpenRouter or auxiliary tasks silently degrade” gotcha is gone – the tradeoff now is cost. If your main model is an expensive reasoning model, route the chatty side tasks to something cheap and fast:
auxiliary:
vision:
provider: "openrouter"
model: "google/gemini-2.5-flash"
compression:
provider: "openrouter"
model: "google/gemini-2.5-flash"
Configuration System
Hermes has a layered configuration system. Understanding the precedence is essential because higher layers override lower ones, and one of the layers is a global provider registry you can’t see in config.yaml.
Config File Layout
Per the upstream docs, these are the files that make up a Hermes configuration:4
~/.hermes/
├── config.yaml # All settings (model, terminal, TTS, compression, memory, toolsets, ...)
├── .env # Secrets (API keys, bot tokens, passwords)
├── auth.json # OAuth provider credentials (Nous Portal, Codex, Anthropic)
├── SOUL.md # Primary agent identity (slot #1 in system prompt)
├── memories/ # Persistent memory (MEMORY.md, USER.md)
├── skills/ # Bundled + agent-created + hub-installed skills
├── cron/ # Scheduled jobs
├── sessions/ # Gateway session state
└── logs/ # agent.log, gateway.log, errors.log (secrets auto-redacted)
config.yaml vs .env — when both are set, config.yaml wins for non-secret settings.4 The rule is:
- Secrets (API keys, bot tokens, passwords) →
.env - Everything else (model, terminal backend, compression settings, memory limits, toolsets) →
config.yaml
Secrets can be referenced from config.yaml using shell-style interpolation:4
auxiliary:
vision:
api_key: ${GOOGLE_API_KEY}
base_url: ${CUSTOM_VISION_URL}
delegation:
api_key: ${DELEGATION_KEY}
Managing Configuration
hermes config # View current configuration
hermes config show # Same as above
hermes config edit # Open config.yaml in your editor
hermes config set KEY VAL # Set a specific value
hermes config get KEY # Print a single value (v0.19.0+)
hermes config unset KEY # Remove a key so the default applies again (v0.19.0+)
hermes config path # Print the config file path
hermes config env-path # Print the .env file path
hermes config check # Check for missing options (after updates)
hermes config migrate # Interactively add missing options
Examples:4
hermes config set model anthropic/claude-opus-4
hermes config set terminal.backend docker
hermes config set OPENROUTER_API_KEY sk-or-... # Saves to .env
hermes config check and hermes config migrate are the commands to run after every hermes update — they catch newly added config options that your file doesn’t yet have.6
Configuration Precedence
Hermes loads configuration from several sources. When multiple sources set the same value, the higher-priority source wins:4
- CLI arguments —
hermes chat --model anthropic/claude-sonnet-4(per-invocation override) - Environment variables — applied at process startup
config.yaml— the primary settings file.env— secrets only- Built-in defaults — applied when nothing else sets a value
CLI flags always win for that single invocation. config.yaml is the long-term source of truth.
Localization (v0.13.0+)
v0.13.0 added 7 locales for CLI and gateway messages: Chinese (Simplified), Japanese, German, Spanish, French, Ukrainian, and Turkish.18 v0.14.0 localizes all gateway commands and the web dashboard, adds 8 more locales, and brings the total to 16.19 At tag v2026.8.31 the locales/ tree holds 17 message catalogs – English plus 16 translations (af, ar, de, es, fr, ga, hu, it, ja, ko, pt, ru, tr, uk, zh-hant, zh).33 Documentation is currently localized in zh-Hans only. Locale resolves from LC_ALL / LANG environment variables or an explicit locale: key in config.yaml. English remains the default and the source of truth for any string a translation hasn’t yet covered.
Profiles — Multiple Isolated Hermes Instances
Profiles give you multiple isolated Hermes instances, each with its own config, sessions, skills, memory, and gateway PID. This is how you run “work Hermes” and “personal Hermes” side-by-side without either seeing the other’s state.6
hermes profile list
hermes profile create work --clone # Clone from current profile
hermes profile use work # Set sticky default
hermes profile alias work --name h-work # Create wrapper script
hermes profile export work -o work-backup.tar.gz
hermes profile import work-backup.tar.gz --name restored
hermes -p work chat -q "Hello from work profile" # One-off without switching
Each profile gets its own HERMES_HOME (~/.hermes-<name>/ by default), so multiple profiles can run the gateway concurrently without stepping on each other.63
The isolation boundary is enforced harder as of v0.21.2: a cluster of fixes closed real cross-profile leaks in multiplexed-gateway setups, where secondary profiles could inherit the default profile’s allow-lists, send credentials to its host, hand its vault secrets to stdio MCP servers, or attach another profile’s .env/auth.json/state.db to MEDIA: delivery (#107609-#107630). If you multiplex profiles, v0.21.2 is the tag where the isolation claim actually holds – see What’s New in v0.21.2.48
CLI Commands
This section is the practitioner’s reference to top-level CLI commands. For the authoritative code-derived reference, see the upstream CLI Commands Reference.6
Global Options
hermes [global-options] <command> [subcommand/options]
| Option | Description |
|---|---|
--version, -V |
Show version and exit |
--profile <name>, -p <name> |
Select which Hermes profile to use |
--resume <session>, -r <session> |
Resume a session by ID or title |
--continue [name], -c [name] |
Resume the most recent session (or match a title) |
--worktree, -w |
Start in an isolated git worktree |
--in <dir> |
Change into DIR before starting or resuming. Combined with --resume latest or -c, the most recent session for DIR’s workspace is picked, and the session stays in DIR (skips the recorded-cwd restore)27 |
--ignore-user-config |
Ignore ~/.hermes/config.yaml and fall back to built-in defaults (credentials in .env are still loaded)27 |
--ignore-rules |
Skip auto-injection of AGENTS.md, SOUL.md, .cursorrules, memory, and preloaded skills27 |
--tui |
Launch the modern TUI instead of the classic REPL27 |
--cli |
Force the classic prompt_toolkit REPL (overrides display.interface=tui)27 |
--dev |
With --tui: run TypeScript sources via tsx (skip dist build)27 |
--yolo |
Bypass dangerous-command approval prompts |
--safe-mode |
Troubleshooting flag — start Hermes in a minimal safe mode to isolate startup problems (v0.19.0+)56 |
--pass-session-id |
Include the session ID in the agent’s system prompt |
Top-Level Commands
| Command | Purpose |
|---|---|
hermes chat |
Interactive or one-shot chat |
hermes model |
Interactively choose default provider and model |
hermes gateway |
Run or manage the messaging gateway |
hermes setup |
Interactive setup wizard |
hermes auth |
Manage credentials — add, list, remove, reset, set strategy |
hermes status |
Show agent, auth, and platform status |
hermes cron |
Inspect and tick the cron scheduler |
hermes webhook |
Manage dynamic webhook subscriptions |
hermes doctor |
Diagnose config and dependency issues |
hermes dump |
Copy-pasteable setup summary for support/debugging |
hermes logs |
View, tail, and filter agent/gateway/error logs |
hermes config |
Show, edit, migrate, query configuration |
hermes pairing |
Approve or revoke messaging pairing codes |
hermes skills |
Browse, install, publish, audit skills |
hermes honcho |
Manage Honcho cross-session memory. Plugin-conditional: the docs state that “Plugin-specific subcommands (e.g. hermes honcho) register automatically when their provider is active”, so it is absent from hermes --help unless Honcho is the active memory provider6 |
hermes memory |
Configure external memory provider |
hermes acp |
Run Hermes as an ACP server (editor integration) |
hermes mcp |
Manage MCP server config; run Hermes as MCP server |
hermes plugins |
Manage plugins |
hermes tools |
Configure enabled tools per platform |
hermes sessions |
Browse, export, prune, delete sessions. v0.19.0 expands hermes sessions export to Markdown, Quarto, HTML, prompt-only, and Hugging Face trace formats, with an opt-in --redact secret-scrubbing pass and age/workspace/platform filters56; hermes sessions recover rebuilds canonical session data from a damaged state.db into a separate NEW database, offline and non-destructive, with --inspect-only reporting canonical-table readability without creating one (v0.21.2’s pointed-to recovery path)48; hermes sessions set-journal-mode delete\|wal (v0.21.4 window) converts a store between WAL and rollback-journal offline – it refuses while any foreign process holds the database or a sidecar, flips without waiting out openers, verifies the SQLite header bytes 18/19 afterward, and takes --db for a non-default store. Stop the gateway, dashboard, and every CLI first; on Windows, where there is no holder scan, it refuses until you pass --force after stopping every Hermes process yourself. The dispatch is deliberately offline so the command never opens the store it converts, and hermes doctor now points at it51 |
hermes insights |
Show token/cost/activity analytics |
hermes claw |
OpenClaw migration helpers |
hermes profile |
Manage profiles (multiple isolated instances) |
hermes completion |
Print shell completion scripts (bash/zsh) |
hermes whatsapp |
Configure and pair the WhatsApp bridge |
hermes --version (-V) |
Print version information. A global flag, not a subcommand: at tag v2026.8.31 the _BUILTIN_SUBCOMMANDS set has no version entry and the startup fast path matches only --version / -V27 |
hermes update |
Pull latest code and reinstall dependencies |
hermes uninstall |
Remove Hermes from the system (--full also deletes config/data) |
hermes backup |
Full backup of config, sessions, skills, and memory (v0.9.0+)16 |
hermes import |
Restore from a backup archive — migrate between machines or roll back (v0.9.0+)16 |
hermes dashboard |
Launch the local web dashboard for browser-based agent management (v0.9.0+)16 |
hermes serve |
Run the backend API server headless — as of v0.19.0 it no longer builds or mounts the web UI56 |
hermes debug share |
Upload a full debug report to a pastebin for sharing when troubleshooting (v0.9.0+)16 |
hermes approvals |
Approval-prompt tools: suggest mines approval history into command_allowlist proposals; test dry-runs the approval verdict for a command without executing it (“never executes it”, with --backend and --json). The v0.21.0 release notes bill this dry-run as hermes approval-check, but no approval-check subcommand exists at tag v2026.8.31 – the surface is hermes approvals test2740 |
hermes bundles |
Create, list, and manage skill bundles (aliases for multiple skills under one /<name> slash command)27 |
hermes checkpoints |
Inspect / prune / clear ~/.hermes/checkpoints/, the shadow store behind /rollback; run with no args for a status overview27 |
hermes computer-use |
Manage the Computer Use (cua-driver) backend (macOS/Windows/Linux)27 |
hermes console |
Open the safe Hermes command console27 |
hermes curator |
Background skill maintenance (curator): status, run, pause, pin27 |
hermes egress |
Manage the iron-proxy egress credential-injection firewall for remote terminal sandboxes (disabled by default)27 |
hermes fallback |
Manage fallback providers, tried when the primary model fails27 |
hermes hooks |
Inspect and manage shell-script hooks: list, test <event>, revoke, and doctor (exec bit, allowlist, mtime drift, JSON validity, synthetic run timing)27 |
hermes import-agent |
Import a Claude Code (~/.claude) or Codex CLI (~/.codex) setup into Hermes27 |
hermes desktop (alias gui) |
Build and launch the native Electron desktop app27 |
hermes kanban |
Multi-profile collaboration board (tasks, links, comments) with boards, a swarm graph (swarm: parallel workers → verifier → synthesizer), and a dispatcher27 |
hermes login / hermes logout |
Deprecated. Use hermes auth to manage credentials, hermes model to select a provider, or hermes setup for full setup27 |
hermes lsp |
Language Server Protocol management: status, list, install <server>, install-all, and restart (tear down running LSP clients; the next edit re-spawns them)27 |
hermes migrate |
Migrate configuration for retired models or deprecated settings27 |
hermes moa |
Configure Mixture of Agents provider/model slots (named presets selectable from the model picker)27 |
hermes journey (aliases learning, memory-graph) |
Timeline of learned skills + memories over time27 |
hermes monitoring |
Inspect gateway monitoring (health & diagnostics export); status shows settings, export state, and redaction posture27 |
hermes pause / hermes resume |
Emergency stop: pause halts cron/kanban dispatch and new gateway turns; resume lifts it27 |
hermes peer |
Bot-to-bot DMs across machines: add, list, remove peer Hermes gateways and dm an agent on one, printing its reply27 |
hermes pets |
Browse, install, and select petdex animated pets27 |
hermes portal |
Set up Nous Portal (login, model pick, Tool Gateway): login (default), info, open, tools; see Nous Tool Gateway2728 |
hermes project |
Manage projects (named, multi-folder workspaces): create, list, show, add/remove folders, rename, set-primary, use, archive, restore, and bind a kanban board27 |
hermes proxy |
Local OpenAI-compatible proxy to OAuth providers: start, status, providers27 |
hermes prompt-size |
Show a byte breakdown of the system prompt + tool schemas; runs offline27 |
hermes send |
Send a message to a configured platform with no agent loop and no LLM (scripts, cron jobs, CI)27 |
hermes skin |
List, switch, and tweak skins (list, use, set)27 |
hermes slack |
Slack integration helpers: manifest prints or writes a Slack app manifest with every gateway command registered as a native slash27 |
hermes sync |
Skill Sync across devices and with your team: status, pull, push, now, enable, disable, device, propose27 |
hermes whatsapp-cloud |
Set up the WhatsApp Business Cloud API integration (distinct from the hermes whatsapp Baileys personal-account bridge)27 |
hermes worktree |
Audit and reclaim accumulated git worktrees and merged branches. list (aliases ls, audit; the default) classifies every tree by age, size, verdict, and reason; prune removes safe trees and deletes fully-merged local branches. Both take --repo <root>; prune adds --dry-run (show the plan without changing anything), --trees-only, and --branches-only. It never deletes uncommitted tracked changes, unique unpushed commits, or in-use trees; untracked-only scratch is archived to ~/.hermes/archive/worktree-prune/ before removal27 |
hermes secrets |
Manage external secret sources (Bitwarden, 1Password) for pulling API keys at process startup27 |
hermes security |
Supply-chain audit (OSV.dev) for the venv, plugins, and MCP servers (audit)27 |
hermes verify |
Detect a project’s run recipe and smoke-test it27 |
hermes chat — The Main Entry Point
hermes with no arguments drops you into interactive chat. hermes chat is the explicit form with options:6
hermes chat -q "Summarize the latest PRs" --oneshot # Answer and exit (without --oneshot, a TTY seeds an interactive session)
hermes chat --provider openrouter --model anthropic/claude-sonnet-4.6
hermes chat --toolsets web,terminal,skills # Enable specific toolsets
hermes chat --quiet -q "Return only JSON" # Programmatic mode
hermes chat --worktree -q "Review repo and open a PR"
Key options:
| Option | Description |
|---|---|
-q, --query "..." |
Query to run. Changed in v0.21.0: on a real TTY it now seeds an interactive session (submitted literally as the first turn); combined with --oneshot or -Q, or on a non-TTY, it answers and exits – the old one-shot behavior40 |
--query-file PATH |
Read the single query from a file instead of the command line (- reads stdin). Nothing is shell-interpreted, so quotes, $(...), and backticks arrive verbatim; mutually exclusive with -q51 |
--oneshot |
With -q/--query-file: answer the query and exit (legacy single-query behavior). Implied on non-TTY stdio and by -Q/--quiet40 |
-m, --model <model> |
Override the model for this run |
-t, --toolsets <csv> |
Enable a comma-separated set of toolsets |
--provider <provider> |
Force a provider (see full list) |
-s, --skills <name> |
Preload one or more skills for this session |
-v, --verbose |
Verbose output |
-Q, --quiet |
Programmatic mode (no banner, spinner, previews) |
--format <fmt> |
Output format for single-query mode (-q or --query-file): text (default) prints the final response as plain text; stream-json emits newline-delimited JSON events (JSONL) – a system/init event, text deltas, tool_use/tool_result events, then one terminal result envelope with exit code, final text, and token stats. stream-json implies --quiet, requires -q or --query-file (exit 2 without one), and cannot be combined with --tui; diagnostics and the session ID stay on stderr, and tool output is capped at 5,000 characters per event (v0.21.4 window)51 |
--resume <session> |
Resume a session directly from chat |
--worktree |
Create an isolated git worktree |
--checkpoints |
Enable filesystem checkpoints before destructive changes |
--yolo |
Skip approval prompts |
--source <tag> |
Session source tag (default: cli; use tool for integrations) |
--max-turns <N> |
Max tool-calling iterations per conversation turn (default: 500 since v0.20.0, or agent.max_turns in config)40 |
hermes setup — Full Wizard
Runs the full setup wizard or jumps into one section:6
hermes setup # Full wizard
hermes setup model # Provider and model only
hermes setup terminal # Terminal backend only
hermes setup gateway # Messaging platforms only
hermes setup tools # Tool enable/disable per platform
hermes setup agent # Agent behavior only
hermes setup --non-interactive
hermes setup --reset # Reset config to defaults before setup
hermes logs — Structured Log Querying
hermes logs is more powerful than tail -f on the log files because it supports filtering by level, session ID, and time range simultaneously.6
hermes logs # Last 50 lines of agent.log
hermes logs -f # Follow in real time
hermes logs gateway -n 100 # Last 100 lines of gateway.log
hermes logs --level WARNING --since 1h # Warnings from the last hour
hermes logs --session abc123 # Filter by session ID substring
hermes logs errors --since 30m -f # Follow errors.log from 30m ago
hermes logs list # List all log files with sizes
Log files live in ~/.hermes/logs/:6
agent.log— all agent activity (API calls, tool dispatch, session lifecycle, INFO+)errors.log— warnings and errors only (a filtered subset of agent.log)gateway.log— messaging gateway activity (platform connections, dispatch, webhooks)
Rotation is automatic via Python’s RotatingFileHandler — look for agent.log.1, agent.log.2, etc.6
hermes doctor — Diagnostics
hermes doctor [--fix] is the first command to run when something is wrong. It checks config validity, dependency presence, API key availability, service status, and can attempt automatic repairs with --fix.6
For sharing diagnostics with someone else, use hermes dump — it produces a compact plain-text summary with redacted API keys, ready to paste into a GitHub issue or Discord thread.6
Slash Commands
Slash commands run inside an active chat session (CLI or messaging platform). They’re dispatched from a shared COMMAND_REGISTRY in hermes_cli/commands.py, which is why most commands work identically across surfaces.9
Session Control
| Command | Description |
|---|---|
/new (alias /reset) |
Start a new session |
/clear |
Clear screen + start new session |
/history |
Show conversation history |
/save |
Save the current conversation |
/retry |
Retry the last message |
/undo |
Remove the last user/assistant exchange |
/title <name> |
Set a title for the current session |
/compress |
Manually compress conversation context |
/rollback [number] |
List or restore filesystem checkpoints |
/stop |
Kill all running background processes |
/status |
Show session, model, token, and context info – since the v0.20.5 window it also surfaces reasoning mode, pending approvals, and context usage4035 |
/queue <prompt> |
Queue a prompt for the next turn. Gotcha: /q is claimed by both /queue and /quit; last registration wins and /q resolves to /quit in practice — always type /queue explicitly.9 |
/resume [name] |
Resume a previously-named session |
/statusbar (alias /sb) |
Toggle context/model status bar |
/background <prompt> (alias /bg) |
Run a prompt in a separate background session |
/btw <question> |
Ephemeral side question (no tools, not persisted) |
/plan [request] |
Load the bundled plan skill to write a plan instead of executing |
/branch [name] (alias /fork) |
Branch the current session |
/goal <target> |
Lock the agent onto a target so it stays on task across turns. Ralph-loop pattern as a first-class primitive. Configurable turn budget. New in v0.13.0.18 |
/subgoal <criterion> |
Add success criteria to an active /goal without restarting the loop. New in v0.14.0.19 |
/handoff <target> |
Transfer the live session — messages, tool calls, and context — to another model, persona, or profile. New in v0.14.0.19 |
/worktree [new [name]\|list\|prune [--dry-run]] |
Inspect, create, or reclaim isolated git worktrees without leaving the session. new creates a worktree under the repo’s .worktrees/ and moves the session into it; list lists them; prune is the same attended reclaim as hermes worktree prune and, per the docs, “never touches the tree the session is running in”27 |
Configuration & Model
| Command | Description |
|---|---|
/config |
Show current configuration |
/model [model-name] |
Show or change the current model |
/provider |
Show available providers and current provider |
/personality [name] |
Set a personality overlay |
/verbose |
Cycle tool progress display |
/reasoning |
Manage reasoning effort and display. v0.19.0 adds max and ultra effort tiers and makes /reasoning session-scoped, alongside per-model and per-MoA-slot effort overrides in config56 |
/skin |
Show or change display skin/theme |
/voice [on\|off\|tts\|status] |
Toggle CLI voice mode |
/yolo |
Toggle YOLO mode (skip approval prompts). Since v0.19.0, user-defined deny rules still block matching commands even in YOLO mode56 |
/fast |
Toggle Fast Mode — priority processing for OpenAI and Anthropic models (v0.9.0+)16 |
/debug |
Quick diagnostics across all platforms (v0.9.0+)16 |
/subscription |
Manage your Nous Portal plan from the terminal — plan and remaining allowance, upgrade/downgrade cost preview, apply with undo (v0.19.0+)56 |
/topup |
Add credit to your Nous Portal balance without leaving the terminal (v0.19.0+)56 |
The /model command is the workhorse for mid-session provider switching:9
/model # Show current model and options
/model claude-sonnet-4 # Switch model (auto-detect provider)
/model zai:glm-5 # Switch provider:model
/model custom:qwen-2.5 # Use model on custom endpoint
/model custom # Auto-detect model from custom endpoint
/model custom:local:qwen-2.5 # Named custom provider
/model openrouter:anthropic/claude-sonnet-4 # Back to cloud
v0.19.0 adds /model --once — a one-turn model override that reverts to your previous model automatically after the response.56 Since the v0.20.5 window the /model picker is also fuzzy: it filters as you type.2335
Tools, Skills & Info
| Command | Description |
|---|---|
/tools [list\|disable\|enable] [name...] |
Manage tools for the current session |
/toolsets |
List available toolsets |
/browser [connect\|disconnect\|status] |
Manage local Chrome CDP connection |
/skills |
Search, install, inspect, or manage skills |
/cron |
Manage scheduled tasks |
/reload-mcp |
Reload MCP servers from config.yaml |
/plugins |
List installed plugins |
/help |
Show all commands |
/usage |
Show token usage, cost, duration |
/insights |
Show usage analytics (last 30 days) |
/platforms |
Show messaging platform status |
/profile |
Show active profile name and home |
/palette |
Open the fuzzy command palette (also Ctrl+P) – fuzzy-search every command and skill. Landed in the v0.20.5 window4035 |
Dynamic Skill Slash Commands
Every installed skill is automatically exposed as a slash command:9
/gif-search funny cats
/axolotl help me fine-tune Llama 3 on my dataset
/github-pr-workflow create a PR for the auth refactor
/excalidraw # Just the skill name loads it and lets the agent ask what you need
Since v0.19.0, slash-skill invocations stack: /skill-a /skill-b do XYZ loads both skills in order in a single turn, with autocomplete and ghost text for the chained names.56
You can also define quick commands in config.yaml that alias a short name to a longer prompt:9
quick_commands:
review: "Review my latest git diff and suggest improvements"
deploy: "Run the deployment script at scripts/deploy.sh and verify the output"
morning: "Check my calendar, unread emails, and summarize today's priorities"
Then type /review, /deploy, or /morning in the CLI.
Prefix Matching
Commands support prefix matching: typing /h resolves to /help, /mod resolves to /model. When a prefix is ambiguous, the first registration in registry order wins. Full command names and registered aliases always take priority over prefix matches.9
Messaging-Specific Commands
Some commands only work on messaging platforms (Telegram, Discord, Slack, WhatsApp, Signal, Email, Home Assistant):9
/status— show session info (no longer messaging-only; see Session Control)/sethome(alias/set-home) — mark the current chat as platform home/approve [session|always]— approve a pending dangerous command/deny [reason]— reject a pending dangerous command. Since v0.19.0,/deny <reason>relays your refusal reason to the agent so it course-corrects instead of retrying blind56/update— update Hermes Agent to latest/commands [page]— browse all commands and skills (paginated)
And some are CLI-only: /skin, /tools, /toolsets, /browser, /config, /cron, /skills, /platforms, /paste, /statusbar, /plugins.9
Tools & Toolsets
Hermes ships with a broad built-in tool registry covering web search, browser automation, terminal execution, file editing, memory, delegation, RL training, messaging delivery, Home Assistant integration, and more.10 Tools are organized into logical toolsets that can be enabled or disabled per platform.
High-Level Categories
| Category | Examples | Description |
|---|---|---|
| Web | web_search, web_extract |
Search the web and extract page content |
| Terminal & Files | terminal, process, read_file, patch |
Execute commands and manipulate files |
| Browser | browser_navigate, browser_snapshot, browser_vision |
Interactive browser automation with text and vision |
| Media | vision_analyze, video_analyze, video_generate, image_generate, text_to_speech |
Multimodal analysis and generation. video_analyze is Gemini-first with extensible support for compatible multimodal providers (v0.13.0+). v0.14.0 adds unified video_generate with pluggable provider backends and sends raw pixels through vision_analyze when the active model is vision-capable.1819 |
| Agent orchestration | todo, clarify, execute_code, delegate_task |
Planning, clarification, code execution, subagent delegation |
| Computer use | computer_use |
Desktop control via cua-driver backend; v0.14.0 makes this work with non-Anthropic vision-capable providers.19 |
| Memory & recall | memory, session_search |
Persistent memory + session search |
| Automation & delivery | cronjob, send_message |
Scheduled tasks, outbound messaging |
| Integrations | ha_*, MCP tools, rl_* |
Home Assistant, MCP, RL training |
Common toolset names include web, terminal, file, browser, vision, image_gen, moa, skills, tts, todo, memory, session_search, cronjob, code_execution, delegation, clarify, homeassistant, and rl.10
Managing Tools
hermes chat --toolsets "web,terminal" # Use specific toolsets
hermes tools # Interactive per-platform tool config
hermes tools --summary # Print enabled-tools summary
Tools can also be toggled mid-session via /tools disable <name> and /tools enable <name>, which resets the session so the new tool set takes effect.9
Terminal Backends
The terminal tool ships seven built-in execution backends – and as of v0.20.6 the set is plugin-extensible, the way the provider picker is (see below):1024
| Backend | Use Case |
|---|---|
local |
Run on your machine (default) — development, trusted tasks |
docker |
Isolated containers — security, reproducibility |
ssh |
Remote server — sandbox, keep agent away from its own code |
singularity |
HPC containers — cluster computing, rootless |
modal |
Serverless cloud execution |
daytona |
Cloud sandbox workspace — persistent remote dev environment |
vercel_sandbox |
Vercel Sandbox cloud microVM – cloud execution with snapshot-backed filesystem persistence. Install hermes-agent[vercel], set terminal.vercel_runtime (node24, node22, or python3.13), and authenticate with VERCEL_TOKEN, VERCEL_PROJECT_ID, and VERCEL_TEAM_ID; the remote workspace root defaults to /vercel/sandbox24 |
Switch backends with hermes config set terminal.backend <name> or in config.yaml:
terminal:
backend: docker # or: local, ssh, singularity, modal, daytona, vercel_sandbox
cwd: "." # Working directory
timeout: 180 # Command timeout in seconds
Plugin backends (v0.20.6+). Third-party sandbox vendors no longer need to land in the core repo: a plugin registers a TerminalEnvironmentProvider at load time via PluginContext.register_terminal_environment_provider, and the registered name becomes selectable through terminal.backend exactly like a built-in. Built-in names are reserved – the registry rejects a provider whose name collides with an in-tree backend – and a registered backend automatically participates in every core surface (the hermes setup backend picker, dashboard probe status, hermes doctor checks, container path/cwd handling, secret stripping) because the core consults the registry at each classification site instead of a hardcoded list of names.32
SSH backend (recommended for security — the agent can’t modify its own code):10
terminal:
backend: ssh
# In ~/.hermes/.env
TERMINAL_SSH_HOST=my-server.example.com
TERMINAL_SSH_USER=myuser
TERMINAL_SSH_KEY=~/.ssh/id_rsa
Docker backend:
terminal:
backend: docker
docker_image: python:3.11-slim
Container resources (applies to docker, singularity, modal, daytona):10
terminal:
container_cpu: 1
container_memory: 5120 # MB (default 5GB)
container_disk: 51200 # MB (default 50GB)
container_persistent: true # Persist filesystem across sessions
With container_persistent: true, installed packages, files, and config survive across sessions.10
All container backends run with security hardening: read-only root filesystem (Docker), all Linux capabilities dropped except DAC_OVERRIDE, CHOWN, and FOWNER, no privilege escalation, PID limits (256 processes), full namespace isolation, persistent workspace via volumes.10
Background Processes
The terminal tool supports background execution with explicit process management:10
terminal(command="pytest -v tests/", background=true)
# Returns: {"session_id": "proc_abc123", "pid": 12345}
process(action="list") # Show all running processes
process(action="poll", session_id="proc_abc123") # Check status
process(action="wait", session_id="proc_abc123") # Block until done
process(action="log", session_id="proc_abc123") # Full output
process(action="kill", session_id="proc_abc123") # Terminate
process(action="write", session_id="proc_abc123", data="y") # Send input
PTY mode (pty=true) enables interactive CLI tools like Codex and Claude Code.10
Sudo
If a command needs sudo, Hermes prompts for your password (cached for the session). Or set SUDO_PASSWORD in ~/.hermes/.env.10
Multi-Agent Kanban (v0.13.0+)
v0.13.0 turns multi-agent collaboration into a first-class primitive: a durable Kanban board that tracks tasks, status, and worker identity across agents and across restarts.18 The board is what makes a swarm of Hermes workers actually finish work instead of stalling on dead handoffs.
| Mechanism | What it does |
|---|---|
| Heartbeats | Each worker pulses while it owns a task. A missed heartbeat marks the worker as suspect and frees the task for reclaim. |
| Reclaim | A different worker can pick up an abandoned task, with full task state and prior partial output. |
| Zombie detection | Workers that exit without marking a task complete are auto-blocked from claiming new work, preventing the swarm from accumulating dead identity. |
| Hallucination gate | Output that fails the gate sends the task back to the board with a noted reason instead of being marked done. |
Per-task max_retries |
Override the default retry budget on a task that you know is fragile. |
| Multi-project boards | One Hermes home can host several independent boards. |
The Kanban board pairs naturally with /goal (locked-target Ralph loop) for the target side and with the existing delegate_task tool for spawn semantics. The result is a swarm pattern where every agent shares one source of truth for what to do next, who is doing it, and what is stuck.
What Is a Hermes Swarm?
A swarm is several Hermes workers running in parallel against one shared Kanban board. It is not a separate subsystem you enable – it is what the board makes possible. The board supplies the one thing parallel agents cannot supply themselves: a single authoritative answer to what should I pick up next, and is anyone already on it?
v0.15.0 promoted this from a pattern to a supported topology, adding swarm topology for parallel worker coordination, auto-decomposition of a high-level goal into subtasks, per-task model overrides, scheduled tasks, and worktree management so parallel workers do not collide in the same checkout.59
| Concern in a naive multi-agent setup | What the board does instead |
|---|---|
| Two workers grab the same task | Task ownership is claimed, with worker identity recorded |
| A worker dies mid-task and the work vanishes | Heartbeat lapses, task is reclaimed with prior partial output |
| A crashed worker keeps “holding” tasks forever | Zombie detection blocks it from claiming new work |
| One expensive model for every subtask | Per-task model overrides – cheap models for mechanical subtasks |
| Parallel workers editing the same files | Worktree management isolates each worker’s checkout |
The practical shape: give the swarm a goal, let auto-decomposition break it into board tasks,
and let workers claim, execute, and return results. Retry budget is per-task (max_retries),
so one fragile subtask does not consume the whole run’s tolerance. Because the board is
durable, a swarm survives restarts – workers reattach and resume against the same state.
A swarm is only as good as its decomposition. The board coordinates workers; it does not make a badly-split goal into a good one. Tasks that share hidden state will still fight, worktrees or not.
Skills System
Skills are on-demand knowledge documents the agent can load when needed. They follow a progressive disclosure pattern to minimize token usage and are compatible with the agentskills.io open standard.11
All skills live in ~/.hermes/skills/ — the primary directory and source of truth. On fresh install, bundled skills are copied from the repo. Hub-installed and agent-created skills also go here.11
Progressive Disclosure
Level 0: skills_list() → [{name, description, category}, ...] (~3k tokens)
Level 1: skill_view(name) → Full content + metadata (varies)
Level 2: skill_view(name, path) → Specific reference file (varies)
The agent only loads the full skill content when it actually needs it.11
SKILL.md Format
---
name: my-skill
description: Brief description of what this skill does
version: 1.0.0
platforms: [macos, linux] # Optional — restrict to OS platforms
metadata:
hermes:
tags: [python, automation]
category: devops
fallback_for_toolsets: [web] # Conditional activation
requires_toolsets: [terminal] # Conditional activation
config: # Config.yaml settings
- key: my.setting
description: "What this controls"
default: "value"
prompt: "Prompt for setup"
---
# Skill Title
## When to Use
Trigger conditions for this skill.
## Procedure
1. Step one
2. Step two
## Pitfalls
- Known failure modes and fixes
## Verification
How to confirm it worked.
Conditional Activation
Skills can show or hide themselves based on which tools are available. This is most useful for fallback skills — free or local alternatives that should only appear when a premium tool is unavailable:11
| Field | Behavior |
|---|---|
fallback_for_toolsets |
Skill hidden when listed toolsets are available |
fallback_for_tools |
Same, but checks individual tools |
requires_toolsets |
Skill hidden when listed toolsets are unavailable |
requires_tools |
Same, but checks individual tools |
Example: the built-in duckduckgo-search skill uses fallback_for_toolsets: [web]. When you have FIRECRAWL_API_KEY set, the web toolset is available and the agent uses web_search — the DuckDuckGo skill stays hidden. Without the API key, the DuckDuckGo skill automatically appears as a fallback.11
Agent-Managed Skills
The agent can create, update, and delete its own skills via the skill_manage tool. This is the agent’s procedural memory — when it figures out a non-trivial workflow, it saves the approach as a skill for future reuse.11
When the agent creates skills:11
- After completing a complex task (5+ tool calls) successfully
- When it hit errors or dead ends and found the working path
- When the user corrected its approach
- When it discovered a non-trivial workflow
Actions:11
| Action | Use for |
|---|---|
create |
New skill from scratch |
patch |
Targeted fixes (preferred — most token-efficient) |
edit |
Major structural rewrites |
delete |
Remove a skill entirely |
write_file |
Add/update supporting files |
remove_file |
Remove a supporting file |
Skill Hub
Browse, search, install, and manage skills from online registries:611
hermes skills browse # Browse all hub skills
hermes skills browse --source official # Browse official optional skills
hermes skills search kubernetes # Search all sources
hermes skills search react --source skills-sh # Search skills.sh directory
hermes skills inspect openai/skills/k8s # Preview before installing
hermes skills install openai/skills/k8s # Install with security scan
hermes skills install skills-sh/anthropics/skills/pdf --force
hermes skills check # Check for upstream updates
hermes skills update # Reinstall changed hub skills
hermes skills audit # Re-scan installed hub skills
hermes skills uninstall k8s
hermes skills publish skills/my-skill --to github --repo owner/repo
hermes skills tap add myorg/skills-repo # Add custom GitHub source
Integrated hub sources:11
| Source | Example | Notes |
|---|---|---|
official |
official/security/1password |
Optional skills shipped with Hermes (builtin trust) |
skills-sh |
skills-sh/vercel-labs/agent-skills/vercel-react-best-practices |
Vercel’s public skills directory |
well-known |
well-known:https://mintlify.com/docs/.well-known/skills/mintlify |
URL-based discovery from sites publishing /.well-known/skills/index.json |
github |
openai/skills/k8s |
Direct GitHub repo/path installs |
clawhub |
— | Third-party skills marketplace |
lobehub |
— | LobeHub agent catalog conversion |
browse-sh |
— | Browserbase skills source |
Default GitHub taps (browsable without setup): openai/skills, anthropics/skills, huggingface/skills, NVIDIA/skills, garrytan/gstack. The claude-marketplace source was removed in v0.20.0; browse-sh replaced it in the source list.1155
Security Scanning
All hub-installed skills go through a security scanner that checks for data exfiltration, prompt injection, destructive commands, supply-chain signals, and other threats.11
Trust levels:11
| Level | Source | Policy |
|---|---|---|
builtin |
Ships with Hermes | Always trusted |
official |
optional-skills/ in the repo |
Builtin trust, no third-party warning |
trusted |
Trusted registries (openai/skills, anthropics/skills) |
More permissive policy |
community |
Everything else | Non-dangerous findings can be overridden with --force; dangerous verdicts stay blocked |
--force can override non-dangerous policy blocks for community skills. It does not override a dangerous scan verdict.11
External Skill Directories
You can point Hermes at additional skill directories scanned alongside the local one:11
skills:
external_dirs:
- ~/.agents/skills
- /home/shared/team-skills
- ${SKILLS_REPO}/skills
Paths support ~ expansion and ${VAR} environment variable substitution. External directories are read-only — when the agent creates or edits a skill, it always writes to ~/.hermes/skills/. Local precedence wins if a skill name exists in both places.11
Pinned Skills: skills.auto_load (v0.21.4 window)
Skill names listed under skills.auto_load in config.yaml are pinned as fully loaded in every new session – CLI, TUI, gateway, cron, and API alike:51
skills:
auto_load:
- team-conventions
- deploy-checklist
The list resolves once, when the agent’s prompt is first built; a missing or disabled name warns and skips rather than failing the session, and HERMES_IGNORE_RULES (the machinery behind --ignore-rules) suppresses the list like the rest of the auto-injected context. This is the standing-orders complement to -s/--skills, which preloads skills for one session only.51
Persistent Memory
Hermes has bounded, curated memory that persists across sessions. Two files make up the agent’s memory, both stored in ~/.hermes/memories/:12
| File | Purpose | Char Limit |
|---|---|---|
MEMORY.md |
Agent’s personal notes — environment facts, conventions, things learned | 2,200 chars (~800 tokens) |
USER.md |
User profile — preferences, communication style, expectations | 1,375 chars (~500 tokens) |
Both are injected into the system prompt as a frozen snapshot at session start. The agent manages its own memory via the memory tool — add, replace, or remove.12
Frozen snapshot pattern: the system prompt injection is captured once at session start and never changes mid-session. This is intentional — it preserves the LLM’s prefix cache for performance. Changes made during a session are persisted to disk immediately but don’t appear in the system prompt until the next session.12
What to Save
Save these (the agent does this proactively):12
- User preferences: “I prefer TypeScript over JavaScript” →
user - Environment facts: “This server runs Debian 12 with PostgreSQL 16” →
memory - Corrections: “Don’t use
sudofor Docker commands, user is in docker group” →memory - Conventions: “Project uses tabs, 120-char line width, Google-style docstrings” →
memory - Completed work: “Migrated database from MySQL to PostgreSQL on 2026-01-15” →
memory
Skip these:12
- Trivial/obvious info
- Easily re-discovered facts
- Raw data dumps (too big for memory)
- Session-specific ephemera
- Information already in context files
Session Search
Beyond MEMORY.md and USER.md, the agent can search its past conversations using the session_search tool. All CLI and messaging sessions are stored in SQLite (~/.hermes/state.db) with FTS5 full-text search. Since the v0.15.0 redesign the tool involves no LLM at all – it returns stored conversation content directly (the source’s phrase: the “single-shape tool returns DB content directly”), which is what made it 4,500x faster and removed its API cost.125934
| Feature | Persistent Memory | Session Search |
|---|---|---|
| Capacity | ~1,300 tokens total | Unlimited (all sessions) |
| Speed | Instant (in system prompt) | One FTS query – no LLM call since v0.15.0 |
| Use case | Key facts always available | Finding specific past conversations |
| Management | Manually curated by agent | Automatic — all sessions stored |
| Token cost | Fixed per session (~1,300 tokens) | On-demand |
Two v0.21.4-window additions, both verified at tag v2026.9.21. First, time bounds: the discovery shape takes after (an inclusive lower bound on session start time) and before (an exclusive upper bound), each an ISO date/datetime – a date-only value means midnight UTC that day – or a relative duration (7d, 24h, 2w), meant only for queries that actually name a time frame; sort remains a ranking bias, not a bound. Second, a zero-result recall retry: FTS5’s implicit AND between terms means a paraphrased multi-word query can miss a stored sentence that lacks even one of its words, so when the exact query and the substring fallbacks all come up empty, the search retries the index matching ANY term, with rank order putting rows that cover more terms first. Hits keep exact-match semantics (the retry is gated on a zero-result miss), and explicit OR/NOT, single-term, and CJK-routed queries are left alone.51
External Memory Providers
For deeper persistent memory beyond MEMORY.md and USER.md, Hermes ships with seven external memory provider plugins: Honcho, OpenViking, Mem0, Holographic, RetainDB, ByteRover, and Supermemory. More, such as Hindsight, install from the plugin catalog.12
Hindsight shipped in the core tree through v0.21.4 and moved to the catalog in v0.21.5, maintained by Vectorize. Existing setups migrate on their own: if memory.provider: hindsight is set, hermes update installs the catalog plugin into every profile home that names it, and the first agent start installs it if it is still missing. With security.allow_lazy_installs: false, the agent-start path only logs a message, so run hermes plugins install hindsight yourself. The plugin lands in ~/.hermes/plugins/hindsight/ and is added to plugins.enabled; memory.provider, memory.hindsight.*, HINDSIGHT_API_KEY, and your memory data stay as they were. Check with hermes memory status and hermes plugins list. The hermes-agent[hindsight] pip extra is gone.53
External providers run alongside built-in memory (never replacing it) and add capabilities like knowledge graphs, semantic search, automatic fact extraction, and cross-session user modeling:612
hermes memory setup # Pick a provider and configure it
hermes memory status # Check what's active
hermes memory off # Disable external provider (built-in only)
Only one external provider can be active at a time. Built-in memory is always active.6
Session Auto-Resume (v0.13.0+)
v0.13.0 makes mid-agent interruption survivable. The gateway auto-resumes interrupted sessions after a restart; /update restarts preserve session state through the upgrade; source-file reloads during dev keep the active session alive instead of forcing a new one.18 Practical effect: long-running gateway work and cron-driven jobs no longer reset their context window when the process restarts.
Checkpoints v2 (v0.13.0+)
State persistence is rewritten in v0.13.0 as a single-store design with real pruning, disk guardrails, and no orphan shadow repos.18 The previous checkpoint system accumulated state on disk across long-running profiles; the v2 store puts a hard ceiling on local checkpoint storage and removes the duplicated bookkeeping that drove that growth. No user-facing config change is required; the next checkpoint write uses the v2 path.
Personality & SOUL.md
SOUL.md is the primary identity of a Hermes instance. It occupies slot #1 in the system prompt, replacing the hardcoded default identity.13
Hermes seeds a default SOUL.md automatically at ~/.hermes/SOUL.md (or $HERMES_HOME/SOUL.md for custom profiles). Existing user files are never overwritten. Hermes only loads SOUL.md from HERMES_HOME — it does not look in the current working directory. This makes personality predictable across projects.13
What Belongs in SOUL.md
Use it for durable voice and personality guidance:13
- tone
- communication style
- level of directness
- default interaction style
- what to avoid stylistically
- how Hermes should handle uncertainty, disagreement, ambiguity
Use it less for:13
- one-off project instructions
- file paths
- repo conventions
- temporary workflow details
Those belong in AGENTS.md, not SOUL.md.
SOUL.md vs AGENTS.md
This is the most important distinction in Hermes identity management:13
SOUL.md — identity, tone, style, communication defaults, personality-level behavior.
AGENTS.md — project architecture, coding conventions, tool preferences, repo-specific workflows, commands, ports, paths, deployment notes.
A useful rule: if it should follow you everywhere, it belongs in SOUL.md. If it belongs to a project, it belongs in AGENTS.md.13
Built-in Personalities
Hermes ships with built-in personalities you can switch to with /personality:1333
| Name | Description |
|---|---|
helpful |
Friendly, general-purpose assistant |
concise |
Brief, to-the-point responses |
technical |
Detailed, accurate technical expert |
creative |
Innovative, outside-the-box thinking |
teacher |
Patient educator with clear examples |
kawaii |
Cute expressions, sparkles, enthusiasm |
catgirl |
Neko-chan with cat-like expressions |
pirate |
Captain Hermes, tech-savvy buccaneer |
shakespeare |
Bardic prose with dramatic flair |
surfer |
Chill bro vibes |
noir |
Hard-boiled detective narration |
uwu |
Maximum cute with uwu-speak |
philosopher |
Deep contemplation on every query |
hype |
MAXIMUM ENERGY |
Custom personalities in config.yaml:13
agent:
personalities:
codereviewer: >
You are a meticulous code reviewer. Identify bugs, security issues,
performance concerns, and unclear design choices. Be precise and constructive.
Then switch with /personality codereviewer.
SOUL.md vs /personality
SOUL.md is the baseline voice. /personality is a session-level overlay.13 Keep a pragmatic default SOUL.md, then use /personality teacher for a tutoring conversation or /personality creative for brainstorming.
Nous Tool Gateway (v0.10.0+)
As of Hermes Agent v0.10.0 (2026-04-16), paid Nous Portal subscribers gain managed access to a curated set of tools through their existing Portal credentials — no extra API keys to manage.61 The Hermes CLI itself remains MIT-licensed and fully open source. What changed is that your Portal auth now unlocks more than model inference.
The fastest route in is hermes setup --portal, which the README introduces as “One command from a fresh install”: it logs you in via OAuth, sets Nous as your provider, and turns on the Tool Gateway. From then on hermes portal manages the relationship. hermes portal login (the default when no subcommand is given) runs the same one-shot onboarding; hermes portal info prints the “Portal auth + Tool Gateway routing summary”; hermes portal open opens the subscription page in your default browser; hermes portal tools lists the gateway tools and which are routed via Nous. hermes portal status survives as a hidden back-compat alias for info, and the docs’ CLI reference at the tag still spells the subcommand status, so both work.28
What’s in the gateway
| Tool | Provider | Use case |
|---|---|---|
| Web search | Firecrawl | Retrieval for agents that need fresh information |
| Image generation | FAL / FLUX 2 Pro | Generate images inline without configuring a FAL key |
| Text-to-speech | OpenAI TTS | Spoken output on messaging gateways |
| Browser automation | Browser Use | Headless navigation and scraping |
How it works
The gateway is opt-in per tool via a new use_gateway config field. If you have Portal credentials in hermes auth and enable the gateway for a tool, that tool’s calls route through Portal. Otherwise your direct API key (if present) is used.
# config.yaml — per-tool gateway opt-in
tools:
web_search:
provider: firecrawl
use_gateway: true # route via Nous Portal subscription
image_generation:
provider: fal
use_gateway: true
Runtime precedence: when the gateway is available and a tool has use_gateway: true, Hermes prefers the gateway even if you also have a direct API key configured. This matters for billing — gateway calls draw from your Portal subscription, not from your direct API key’s balance.
Enabling the gateway
hermes model # select Nous Portal (OAuth flow)
hermes tools # per-platform tool picker integrates gateway tools
hermes status # confirms gateway/subscription detection
The subscription is detected automatically from the Portal OAuth credentials you already have in hermes auth — there is no separate login step. Since v0.19.0 you can also manage the subscription itself from inside a session: /subscription shows your plan and remaining allowance, previews exactly what an upgrade costs or when a downgrade takes effect, and applies the change with scheduled-change banners and undo; /topup adds credit. The desktop app has a matching billing settings tab.56
Pricing and access
Pricing and tier names are published on the Nous Portal pricing page (https://portal.nousresearch.com/pricing). This guide does not enumerate tiers because they’re the responsibility of the Portal product, not the Hermes CLI, and they change independently of Hermes releases. Sign up at https://portal.nousresearch.com/ and check the pricing page for current tiers.
Nous free tier and guided first launch (v0.21.2+)
Since v0.21.2 a fresh install does not need a paid plan or an API key to produce a working agent: free inference and connectors come out of the box, with one command to sign in, and /login starts the sign-in from inside a chat. Connector tools (Gmail, Linear, Notion, and the rest) are searchable through tool_search like any other tool. The desktop adds a guided first launch behind HERMES_GUEST_ONBOARDING=1; only the literal value 1 enables it – the desktop’s own test asserts that 'true', '0', and an empty value all leave it off, and the launch decision is stamped into the spawned backend’s environment so an inherited value never leaks through.48
Since the v0.21.4 window, connecting those connectors is one backend-owned operation rather than scattered per-frontend logic: a manage_connections tool call drives a pure-data connection state machine on the backend (with a fixed 300-second operation deadline – deliberately not a config key), and Desktop, TUI, and CLI all render it as the same setup card. The card draws a field per missing credential (name, prompt, required) and holds its verb until every required field has text, and the backend enforces the same secrets-only split on every frontend.51
Deprecation notice
HERMES_ENABLE_NOUS_MANAGED_TOOLSenv var is removed in v0.10.0. Managed tools are now enabled via the per-tooluse_gatewayconfig field and gated on your Portal subscription state.61
Framing: what this release is not
The Hermes Agent CLI is not gated behind a subscription. The project is still MIT-licensed, all core features (CLI, skills, memory, messaging gateway, cron, MCP, local dashboard, BYOK for every provider) work end-to-end without paying anyone. v0.10.0 adds a convenience path for users who already pay for Nous Portal — it doesn’t remove anything from the free path.
Messaging Gateway
Hermes can run as a long-running gateway process that connects to 28 messaging platforms from a single gateway process: Telegram, Discord, Slack, WhatsApp, Signal, SMS, Email, Home Assistant, Mattermost, Matrix, DingTalk, Feishu/Lark, WeCom, Weixin (WeChat), BlueBubbles (iMessage), QQBot, Microsoft Teams, Tencent Yuanbao, Google Chat, LINE, SimpleX Chat, Photon (iMessage), WhatsApp Cloud API, WeCom Callback, Raft, IRC, ntfy, Buzz, and a generic Webhook adapter.360171819 That 28 is the docs’ Platform Comparison table at tag v2026.8.31; underneath it, gateway/config.py defines 24 built-in Platform enum members (including non-chat entries such as local, api_server, webhook, msgraph_webhook, and relay) and resolves any other name on demand to one of the 22 bundled adapter directories under plugins/platforms/, so a bare enum or directory count will not match the docs number.25 v0.9.0 added iMessage via BlueBubbles (auto-webhook registration, setup wizard, crash resilience) and native WeChat support via iLink Bot API with WeCom callback mode for enterprise apps.16 v0.11.0 added QQBot.60 v0.12.0 added Microsoft Teams and Tencent Yuanbao.17 v0.13.0 added Google Chat as the 20th platform, riding the same pluggable adapter architecture; IRC and Microsoft Teams were also migrated onto the new adapter pattern with generic env_enablement_fn / cron_deliver_env_var plugin hooks.18 v0.14.0 adds LINE and SimpleX Chat and completes the Microsoft Teams stack end-to-end with Graph auth, webhook listener, pipeline runtime, and outbound delivery.19 v0.17.0 (June 19, 2026) adds relay-free iMessage via Photon Spectrum (device-code OAuth with hermes photon login — no Mac/BlueBubbles relay required), the official WhatsApp Business Cloud API adapter (replacing the bridge-process requirement), SimpleX groups and native attachments, and Raft as a bundled platform plugin.21 Two more sit in the docs table with no release-note fanfare: ntfy is a lightweight HTTP pub-sub push channel (subscribe to a topic from the ntfy mobile app, message the topic to talk to the agent, and get the reply back on your phone; works with the public ntfy.sh server or a self-hosted instance, no SDK or daemon required), and Buzz connects Hermes to a Buzz community, Block’s open-source human-plus-agent collaboration platform on the Nostr protocol, shelling out to the buzz CLI for outbound messages and using a native Nostr WebSocket subscription inbound. Both are wired through hermes gateway setup.25
Setup
hermes gateway setup # Interactive platform configuration
hermes gateway install # Install as user service (systemd/launchd)
hermes gateway start # Start the installed service
hermes gateway stop
hermes gateway restart
hermes gateway status
hermes gateway run # Run in foreground (debugging)
The interactive setup walks you through connecting each platform: API tokens, bot IDs, channel mappings, allowlists.6
How Messages Flow
From the upstream architecture docs:3
Platform event → Adapter.on_message() → MessageEvent
→ GatewayRunner._handle_message()
→ authorize user
→ resolve session key
→ create AIAgent with session history
→ AIAgent.run_conversation()
→ deliver response back through adapter
Every messaging platform runs through the same AIAgent conversation loop as the CLI. That’s why slash commands work identically in both places and why a cron job scheduled in Telegram can deliver its output to Discord — the platform difference is just at the edge.3
v0.19.0 adds profile-based message routing and durable delivery. A single multiplexed gateway sharing one bot token can route specific guilds, channels, or threads to different profiles — each with fully isolated config, skills, memory, and secrets — with a GATEWAY_MULTIPLEX_PROFILES override, and a hardening wave means one misconfigured profile can no longer take down the whole gateway. Under the hood, the routing index moved into state.db (sessions.json is now an optional legacy mirror), and final responses are recorded in a durable delivery-obligation ledger around the platform send — a finished answer that hits a gateway crash is redelivered on the next boot instead of silently lost.56 The isolation between routed profiles was hardened at v0.21.2, which closed a cluster of cross-profile leaks in exactly this multiplex setup: inherited allow-lists, credentials sent to the default profile’s host, the default profile’s vault secrets reaching stdio MCP servers, cross-profile MEDIA: attachments, and a sibling’s Nous bearer surviving in per-process memos (#107609-#107630).48
v0.21.1 makes conversation boundaries explicit-only. The session-lifecycle doc at tag v2026.9.7 states the contract in four sentences: “Inactivity and wall-clock time never rotate a conversation. /new and /reset create an explicit boundary; context compression continues to manage long histories. Legacy timer configuration is ignored. The existing SessionResetPolicy datatype is inert compatibility data, not a runtime policy.” Explicit suspension still creates a boundary on the next inbound turn, recovery respects finalized boundaries rather than reopening them, and resource-only eviction leaves conversations resumable. If a session on your gateway seems to “never expire,” that is the design now; rotate it yourself with /new.43
Since the v0.21.4 window, a second hermes gateway run attaches or refuses instead of double-binding. The rule is one hermes serve and one hermes gateway run per host per OS user, each multiplexing every profile. Starting a gateway for a profile the running multiplexer already serves just attaches and exits 0; if it does not serve that profile yet, Hermes asks it to rescan profiles/ and attaches once it does; if it cannot be made to serve the profile, the command refuses rather than silently starting a second gateway. --replace now targets that host process, whichever profile launched it, and --force skips the question entirely when the owner is wedged. Standalone gateways still coexist: if the running gateway is another profile’s standalone (non-multiplexing) one, your profile starts its own beside it as before, until that migration is forced (#109417). If two gateways launch at the same moment, the one that loses the host lock exits 75, which every supervisor Hermes generates retries; by then the winner’s record exists, so the retry attaches or refuses by the rules above. The locks and the rendezvous record live under $HERMES_GATEWAY_LOCK_DIR, else $XDG_STATE_HOME/hermes/gateway-locks (default ~/.local/state/hermes/gateway-locks), scoped to the OS user; a record whose PID is dead, or was recycled by another process, is treated as stale and ignored, so a crashed gateway does not block the next start. The Desktop app follows the same rule, attaching to the running host backend instead of spawning a second one.51
Since v0.21.5, multiplexing is no longer optional. gateway.multiplex_profiles has one valid value, true: an unset key resolves on and is written into the default profile’s config.yaml, and an explicit false is rewritten to true in place with a one-time boxed notice at that gateway start and again in the next hermes update summary. Two controls replace the old opt-out. To take one profile offline without stopping everyone’s bots, run hermes -p <name> gateway stop: the host parks it (a profiles/<name>/gateway.parked marker), and hermes -p <name> gateway start unparks it; the dashboard and Desktop Stop/Start buttons do the same. A named profile that still needs its own gateway sets gateway.standalone: true in its own config.yaml; the host never serves it, and its stop/start act on its own process. The docs call this key “a temporary compatibility shim”, not a supported topology, and it is ignored with a warning on the default profile. A split across OS users, or a HERMES_HOME outside the default home’s profiles/, still takes --force.53
User Authorization & Pairing
hermes pairing list # Show pending and approved users
hermes pairing approve <platform> <code>
hermes pairing revoke <platform> <user-id>
hermes pairing clear-pending
Pairing codes prevent random strangers from talking to your gateway. A user sends a pairing code from their messaging platform; you approve it with hermes pairing approve; from then on they’re authorized.6
unauthorized_dm_behavior decides what a stranger’s DM gets before pairing: pair DMs a pairing code, ignore drops the message silently, and decline, the value added in the v0.21.4 window, sends one polite refusal and then goes silent toward that sender for 24 hours (#88028). The refusal is deduplicated per platform and sender, across aliases, and its text comes from unauthorized_dm_decline_message (a global key; empty means the built-in reply: “Hi! I’m a personal assistant and can only chat with my owner, so I can’t help you directly. Sorry!”). Set the behavior globally, or per platform where the setup wizard writes it:
# ~/.hermes/config.yaml
unauthorized_dm_behavior: decline # global; gateway.unauthorized_dm_behavior also works
unauthorized_dm_decline_message: "" # empty = built-in reply
platforms:
telegram:
unauthorized_dm_behavior: pair # per-platform value always wins
When nothing is set, the effective default depends on your allowlists. With none configured it is pair; once any allowlist is set (GATEWAY_ALLOWED_USERS, or a platform’s allowed-users, group allowed-users, or group allowed-chats variable), it becomes ignore, because the allowlist signals a deliberately restricted gateway, and sending codes to unknown contacts is noisy and a potential information leak (#9337). A global ignore or decline overrides that rule, but a global pair cannot, since it reads the same as the default; to keep pairing alongside an allowlist, set pair per platform. (A platform adapter’s own dm_policy, where one is set, is consulted before the allowlist rule.) Email is inbox-shaped and defaults to ignore unless its own per-platform key opts in; a global value does not reach it. hermes gateway setup offers decline as one of its answers when you configure a platform without an allowlist.51
Scheduled Tasks (Cron)
Hermes has a first-class cron system where jobs are agent tasks, not shell commands. Each scheduled job runs through a fresh AIAgent with the configured prompt, optional attached skills, and delivers results to any platform:36
hermes cron list
hermes cron create --prompt "Check HN for AI news and summarize" --schedule "0 9 * * *" --deliver telegram
hermes cron edit <id>
hermes cron pause <id>
hermes cron resume <id>
hermes cron run <id> # Trigger now on the next tick
hermes cron remove <id>
hermes cron status # Check if scheduler is running
hermes cron tick # Run due jobs once and exit
Or create one conversationally inside a messaging chat:
Every morning at 9am, check Hacker News for AI news and send me a summary on Telegram.
The agent will set up the cron job via its tools. Jobs persist in JSON and survive restarts.3
v0.21.0 gave scheduled jobs memory and judgment. Four mechanisms, all verified at tag v2026.8.31:3537
- Continuity.
continuity=trueinjects a job’s own most recent output into each run, so a scout or monitor “wakes up seeing what it reported last time and can dedupe and continue where it left off” – the injected framing is “avoid repeating what was already reported”, the first run is unchanged, and internally the flag is stored as the reservedselfentry incontext_from. Toggle it withhermes cron create ... --continuityandhermes cron edit <job_id> --continuity/--no-continuity.37 - Durable notepads. Every job gets a small KV scratchpad for cursors, watermarks, and watchlists (16 KB per value, 64 KB per job – it is prompt-injected each run, so the caps are deliberate), written via
hermes cron notepad <job_id> set <key> <value>, which the running agent invokes through its terminal tool.37 - Monitor mode. A job can attach a cheap
monitor_script/monitor_urlsource that runs first each tick: unchanged output (compared as exact bytes) suppresses the agent run entirely – no LLM call, no delivery, a silentno_changerun – while a change injects a diff block and runs the agent normally. Emit stable output from monitor scripts, or every tick looks like a change.37 - Per-job reasoning effort and Bot Chat delivery.
--reasoning-effortpins a job’s thinking level (nonethroughultra), overriding the global and per-model settings for that job’s runs; anddeliver=bot-chatlands the output in a profile’s canonical Bot Chat session as a real incoming message, where the bot “acts on anything that needs action, and responds in its chat” instead of a human just reading a channel.37
MCP Integration
Hermes supports the Model Context Protocol as both a client and a server:6
As a client — connect Hermes to external MCP servers to extend its tool surface:
hermes mcp add <name> --url https://example.com/mcp
hermes mcp add <name> --command npx --args "-y,@modelcontextprotocol/server-github"
hermes mcp list
hermes mcp test <name>
hermes mcp remove <name>
hermes mcp configure <name> # Toggle individual tool selection
hermes mcp login <name> # Force re-auth for an OAuth server (--flow browser|device)
hermes mcp reauth [--all] # Re-authenticate one OAuth server, or every one
Or manually in config.yaml:14
mcp_servers:
github:
command: npx
args: ["-y", "@modelcontextprotocol/server-github"]
env:
GITHUB_PERSONAL_ACCESS_TOKEN: "ghp_xxx"
As of v0.19.0, MCP tools are exposed to the model under the mcp__server__tool naming convention — every tool name carries its server name, so two servers exposing the same tool no longer collide — and MCP server log notifications are surfaced in agent.log.56
v0.21.1 adds a device-code path to MCP OAuth: hermes mcp login <name> takes --flow {browser,device} – browser is the standing PKCE flow, device is an RFC 8628 device-code login for headless or remote machines – and the flag overrides the server’s oauth.flow config. The same window enforces profile ownership throughout OAuth sessions, ignores malformed OAuth metadata caches instead of wedging a server, and relays desktop MCP OAuth through client-local callbacks; -t/--toolsets now also filters which configured MCP servers are spawned, so a scoped invocation skips cold-starting servers it does not need.44
v0.21.0 turns the desktop’s MCP surface into a command center: servers and the catalog merge into one page with drag-in “paste anything” import, background health checks that surface expiring auth before a tool call fails, a fleet cost/usage overlay showing schema token estimates and 30-day usage per server, and hermes:// deep links that install an MCP server with explicit confirmation.35
The v0.21.4 window adds mcp.discovery_concurrency (default 4; 0 = unlimited, #117373): a cap on how many configured MCP servers the discovery pass connects to simultaneously. Every server still connects – the cap only stops them all arriving at once – and a non-integer or negative value logs a warning and falls back to the default.51
As a server — expose Hermes conversations to other agents:
hermes mcp serve
hermes mcp serve -v # Verbose
Context Compression
Hermes automatically compresses long conversations to stay within your model’s context window. The compression summarizer is a separate LLM call – you can point it at any provider or endpoint.4 As of v0.20.6, the retained-tail policy defaults to lean (compression.tail_mode: lean), and at the tag the summarizer’s model/provider/endpoint knobs live under auxiliary.compression.* rather than the older compression.summary_* – legacy keys are migrated automatically on first load (config version 17).3031
compression:
enabled: true
threshold: 0.50 # Compress at this % of context limit
threshold_tokens: null # Optional absolute token cap -- trigger fires at the lower of ratio vs cap
target_ratio: 0.20 # Fraction of threshold to preserve as recent tail (legacy tail mode)
tail_mode: lean # Tail retention: "lean" (default) or "legacy"
protect_last_n: 20 # Min recent messages to keep uncompressed
protect_first_n: 3 # Non-system head messages pinned across compactions
auxiliary:
compression:
model: "" # Empty = main chat model; e.g. "google/gemini-3-flash-preview"
provider: "auto" # "auto", "openrouter", "nous", "codex", "main", etc.
base_url: null # Custom OpenAI-compatible endpoint (overrides provider)
What tail_mode decides. legacy keeps a target_ratio-sized verbatim tail – on big-window or raised-threshold setups that hoards 100-240K tokens per compaction. lean keeps a clamped verbatim tail of 2.5% of the context window (10K floor, 25K cap) and carries continuity in the summary instead: a detailed identifier-preserving session log of the compacted region (one auxiliary summarizer call per attempt), a mechanically extracted anchor index (PR numbers, SHAs, paths, error strings – regex, never paraphrased), every real user message quoted verbatim, and a session_search recovery pointer so the agent can re-access anything summarized away. The docs’ measured result on 500K-token real sessions: ~49K tokens retained instead of ~162K. Old tool results inside the lean tail are demoted to one-line stubs carrying a recovery pointer, and unknown tail_mode values fall back to lean.31
auxiliary.compression.provider |
auxiliary.compression.base_url |
Result |
|---|---|---|
auto (default) |
not set | Auto-detect best available provider |
nous / openrouter / etc. |
not set | Force that provider, use its auth |
| any | set | Use the custom endpoint directly (provider ignored) |
The summary model must support a context length at least as large as your main model’s, since it receives the full middle section of the conversation in a single call – if its window is smaller, the call fails and the middle turns are dropped without a summary.431
Budget Pressure Warnings
When the agent works on a complex task with many tool calls, it can burn through its iteration budget (default: 500 turns as of v0.20.0, up from 90) without realizing it. Budget pressure automatically warns the model:4
| Threshold | Level | What the model sees |
|---|---|---|
| 70% | Caution | [BUDGET: 350/500. 150 iterations left. Start consolidating.] |
| 90% | Warning | [BUDGET WARNING: 450/500. Only 50 left. Respond NOW.] |
Stream Timeouts
The LLM streaming connection has two timeout layers that auto-adjust for local providers (localhost, LAN IPs):4
| Timeout | Default | Local providers | Env var |
|---|---|---|---|
| Socket read timeout | 120s | Auto-raised to 1800s | HERMES_STREAM_READ_TIMEOUT |
| Stale stream detection | 180s | Auto-disabled | HERMES_STREAM_STALE_TIMEOUT |
| API call (non-streaming) | 1800s | Unchanged | HERMES_API_TIMEOUT |
The socket read timeout is raised to 30 minutes for local endpoints because local LLMs can take minutes for prefill on large contexts before producing the first token.4
Local Web Dashboard (v0.9.0+)
A browser-based dashboard for managing your Hermes Agent locally. Configure settings, monitor sessions, browse skills, and manage your gateway without touching config files or the terminal.16 Launch with hermes dashboard. This is the easiest onboarding path for new users who prefer a GUI.
Background Process Monitoring (v0.9.0+)
watch_patterns lets you set patterns to monitor in background process output and get notified in real-time when they match.16 Monitor for errors, wait for specific events (“listening on port”), or watch build logs — all without polling. Combined with notify_on_complete from v0.8.0 (which notifies on background task completion), Hermes now has a full background process observability layer.15
Pluggable Context Engine (v0.9.0+)
Context management is now a pluggable slot via hermes plugins. Swap in custom context engines that control what the agent sees each turn — filtering, summarization, or domain-specific context injection.16 This decouples context strategy from the core agent loop, allowing per-project or per-domain context customization.
Backup & Restore (v0.9.0+)
hermes backup creates a full archive of your config, sessions, skills, and memory. hermes import restores from a backup archive.16 Use this to migrate between machines, create snapshots before major changes, or share a known-good configuration with teammates.
Termux / Android Support (v0.9.0+)
Hermes runs natively on Android via Termux. Adapted install paths, TUI optimizations for mobile screens, voice backend support, and /image command work on-device.16
Security Hardening (v0.13.0+)
v0.13.0 closed 8 P0 security issues and changed one default in the user’s favor.18 v0.14.0 follows with another 12 P0 and 50 P1 closures, including sudo brute-force / sudo-stdin hardening, dangerous-command bypass fixes, tool-error sanitization before model reinjection, dashboard plugin API auth, skills-hub SSRF coverage, and supply-chain advisory scanning during install.19
| Fix | What changed |
|---|---|
| Secret redaction default-on | Previously opt-in. Logs and hermes debug share uploads redact secrets unless explicitly disabled. v0.12.0 had disabled redaction by default after payload-corruption reports; v0.13.0 re-enables it as the safer baseline. |
| Discord cross-guild DM bypass (CVSS 8.1) | Discord role allowlists are now guild-scoped, closing a path where a user role on one guild authorized DMs across all of them. |
| WhatsApp default restrictions | The WhatsApp adapter rejects strangers by default and never responds in self-chat. |
| MCP OAuth TOCTOU window | Closed a race condition during credential save in MCP OAuth flows. |
CLI auth.json TOCTOU |
Closed an analogous TOCTOU window in the credential writer for the CLI auth store. |
| Browser SSRF floor | Hybrid routing enforces a cloud-metadata SSRF floor against requests that try to reach 169.254.169.254 and equivalents. |
| Cron prompt-injection scanning | Assembled prompts (including loaded skill content) are scanned for prompt injection before the cron job runs. |
hermes debug share redaction |
Debug share uploads redact log content at upload time, not just at write time. |
If you maintain a Hermes deployment, treat v0.13.0 and v0.14.0 as security-relevant upgrades, not just feature drops. v0.13.0 closes the Discord cross-guild bypass and two TOCTOU windows; v0.14.0 adds another hardening pass across sudo handling, tool-error reinjection, plugin APIs, skills-hub SSRF, and dependency advisories.
v0.21.0 adds a fourth wave. Writes to protected agent-instruction files – AGENTS.md, CLAUDE.md, SOUL.md, .cursorrules, skills, and memory stores – now always require approval, so a prompt-injected agent cannot quietly rewrite its own standing orders. The gate ships on by default (security.protected_instruction_files: true, with an fnmatch basename extension list in protected_instruction_extra_patterns), and the source names the exact vector it closes: “an injected instruction that edits AGENTS.md / CLAUDE.md / SOUL.md”, noting that instruction files are loaded from cwd trees, so “an AGENTS.md anywhere the agent might later run from is a live target.”3539 The same release closes secret-leak gaps across terminal errors, .env file reads, checkpoints, and ACP logs; teaches the approval system Windows destructive commands and paths; makes macOS permission grants survive updates via a stable TCC signing identity (one-shot setup: hermes desktop --setup-tcc-identity, macOS-only, requires openssl/security/codesign); removes the Blender MCP catalog entry and skill after an upstream compromise; and adds Tier-1 security scanning to plugin installs.3539
Architecture for Practitioners
This section is for people who want to understand what’s happening under the hood so they can debug it, extend it, or reason about performance. It’s a synthesis of the upstream architecture docs.3
Entry Points → AIAgent
Every entry point in Hermes ultimately calls AIAgent.run_conversation():
┌──────────────────────────────────────────────────────────────────┐
│ Entry Points │
│ │
│ CLI (cli.py) Gateway (gateway/run.py) ACP (acp_adapter/) │
│ Batch Runner API Server Python Library │
└──────────┬──────────────┬───────────────────────┬────────────────┘
│ │ │
▼ ▼ ▼
┌──────────────────────────────────────────────────────────────────┐
│ AIAgent (run_agent.py) │
│ │
│ ┌─────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Prompt │ │ Provider │ │ Tool │ │
│ │ Builder │ │ Resolution │ │ Dispatch │ │
│ └──────┬──────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ ┌──────┴───────┐ ┌──────┴───────┐ ┌──────┴───────┐ │
│ │ Compression │ │ 3 API Modes │ │ Tool Registry│ │
│ │ & Caching │ │ chat_compl │ │ 47 tools │ │
│ │ │ │ codex_resp │ │ 20 toolsets │ │
│ │ │ │ anthropic │ │ │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└──────────────────────────────────────────────────────────────────┘
Diagram adapted from the upstream architecture docs.3
“47 tools / 20 toolsets” vs “28 tools” in your banner. The “47 tools” count is the upstream repository’s total tool registry — every tool Hermes ships with source code for, across every toolset. Your actual running CLI will show a smaller number in its startup banner (the installation I verified this guide against reports 28 tools / 89 skills). That’s not a bug. Many toolsets are opt-in and have to be explicitly enabled in config.yaml under toolsets: — messaging platform adapters, browser automation, heavier scraping tools, etc. The registry total is “what’s available”; the banner number is “what’s enabled in your current profile.” Check which toolsets are active with hermes tools --list and enable or disable individual toolsets with the toolsets: block in ~/.hermes/config.yaml (or /tools list / /tools enable <name> / /tools disable <name> inside a running session — removing a tool triggers a session reset so the agent rebuilds its tool manifest).
The Three API Modes
Hermes abstracts provider differences into three API modes, selected automatically at runtime:3
| API mode | Used by |
|---|---|
chat_completions |
OpenRouter, z.ai, Kimi, MiniMax, DeepSeek, Alibaba, most custom endpoints, any OpenAI-compatible server |
codex_responses |
OpenAI Codex (via ChatGPT OAuth) |
anthropic_messages |
Anthropic API (native), Anthropic OAuth, Anthropic-compatible proxies |
The runtime_provider.py resolver maps (provider, model) tuples to (api_mode, api_key, base_url) for 18+ providers, handling OAuth flows, credential pools, and alias resolution.3
Data Flow Through a CLI Session
User input → HermesCLI.process_input()
→ AIAgent.run_conversation()
→ agent.prompt_builder.build_system_prompt()
→ runtime_provider.resolve_runtime_provider()
→ API call (chat_completions / codex_responses / anthropic_messages)
→ tool_calls? → model_tools.handle_function_call() → loop
→ final response → display → save to SessionDB
From the upstream architecture page.3
Prompt Assembly Order
The prompt stack includes:13
SOUL.md(agent identity — or built-in fallback if unavailable)- Tool-aware behavior guidance
- Memory/user context (
MEMORY.md,USER.md) - Skills guidance
- Context files (
AGENTS.md,.cursorrules) - Timestamp
- Platform-specific formatting hints
- Optional system-prompt overlays such as
/personality
SOUL.md is the foundation — everything else builds on top of it.13
Session Storage
SQLite-based session storage with FTS5 full-text search. Sessions have lineage tracking (parent/child across compressions), per-platform isolation, and atomic writes with contention handling.3
If the store gives you trouble on a v0.21.x install, be on v0.21.2 or later: it closed the class of v0.21.0-era state.db fragility (second writers cancelling each other’s locks, healthy databases reported as corrupt, one bad row killing sessions list) and moved hosted-room coordination out of the root store into a dedicated shared-state.db so profile gateways never open the master session store writable. hermes doctor now names structural damage and FTS-index damage separately, FTS damage degrades search instead of failing the turn, and hermes sessions recover --inspect-only (offline, non-destructive, profile-pinned) reports canonical-table readability without creating an output database – the subcommand predates the window but had not been documented here. See What’s New in v0.21.2.48 For a store stuck in the wrong SQLite journal mode, the v0.21.4 window added the offline converter hermes sessions set-journal-mode delete|wal (see the hermes sessions row in Top-Level Commands).51
Plugin System
Three discovery sources: ~/.hermes/plugins/ (user), .hermes/plugins/ (project), and pip entry points. Plugins register tools, hooks, and CLI commands through a context API. Memory providers are a specialized plugin type under plugins/memory/.3 Since v0.21.2 there is also a curated, SHA-pinned plugin catalog you can browse and install from by name, and hermes plugins pack for “declarative, shareable plugin sets”: one hermes-pack.yaml pinning a set of plugins to exact commit SHAs, with pack installs fanning out to ordinary pinned installs and capability consent staying per-plugin.48
The v0.21.4 window built that catalog out from a CLI surface into a shipped artifact. The repo’s plugin-catalog/ directory grew from 9 entries at v2026.9.14 to 228 at v2026.9.21 – one reviewed YAML per plugin (name, repo, maintainer, tier, category, capabilities), pinned to an exact 40-character commit SHA, with presence in the directory constituting admission – published as plugin-catalog.json for the CLI (fetched live, cached under ~/.hermes/cache/). The docs site now builds a page per plugin (/docs/plugins/<name>) and per author (/docs/plugins/by/<slug>) from that same data, each page rendering the plugin’s README fetched from the pinned commit, never a branch tip, through a build-time allowlist that drops raw HTML, with added/updated sorting stamped from committer dates. The ten community plugins the v0.21.4 release names are all catalog entries at the tag, under their actual catalog slugs: hermes-tailscale, hermes-ssh, shodan, hermes-terminal, hermes-rss (plus a separate rss-reader), hermes-resetwatch, done-bell, kiwi, cognee, and web-octen. On the desktop, the Plugins hub can now uninstall a plugin behind a confirm dialog – catalog plugins through plugins.manage remove, standalone desktop plugins through the Electron loader.51
hermes plugins # Interactive enable/disable UI
hermes plugins browse # List every curated plugin catalog entry (v0.21.2+)
hermes plugins search <query> # Search the curated plugin catalog (v0.21.2+)
hermes plugins install <name|repo> # Install from the curated catalog, a Git URL, or owner/repo
hermes plugins enable <name>
hermes plugins disable <name>
hermes plugins list
hermes plugins pack install <src> # Shareable SHA-pinned plugin sets; also: pack export, pack show (v0.21.2+)
Compat window (v0.21.1), now expired: external plugins had to move off pre-decomposition import paths by 2026-09-14. The September 2026 decomposition (PR #102117) relocated the internals that plugins commonly imported, and a temporary
COMPAT_MANIFEST.mdlayer re-exported 1,148 moved public names from their old modules, warning once per name per process (HermesPluginCompatWarning). The removal took effect on schedule on 2026-09-14 via a date gate in the shipped code (COMPAT_REMOVAL_DATEinhermes_cli/plugin_compat.py; no code revert was needed): since that date an affected plugin is disabled – not loaded, with a red notice in the CLI banner,hermes doctor, andhermes update, a one-time desktop modal, and the reason shown inhermes plugins list. Authors:hermes plugins compat <path>still prints everyfile:linewith old path -> new path and exits 1 while anything remains (add--jsonfor machine-readable output). Users stuck on an unmaintained plugin can still setplugins.allow_deprecated_imports: trueinconfig.yaml– it must be a literal YAML boolean, not a quoted string – and it still works because the revert that deletes the old paths has not landed: the manifest and shims are present atv2026.9.14, atv2026.9.21, atv2026.9.24, and onmainas of September 24 (the v0.21.4 window touched the compat module only to cache plugin scans and normalize Windows paths; the v0.21.5 window lefthermes_cli/plugin_compat.pyunchanged and removed only the manifest entries for the deleted Hindsight module). The hatch stops working the moment that revert lands, because the old paths themselves disappear. Only public top-level names were covered; private names and test monkeypatch seams were never a surface and are not restored.4253
Design Principles
From the upstream architecture page:3
| Principle | What it means in practice |
|---|---|
| Prompt stability | System prompt doesn’t change mid-conversation. No cache-breaking mutations except explicit user actions (/model) |
| Observable execution | Every tool call is visible to the user via callbacks. Progress updates in CLI (spinner) and gateway (chat messages) |
| Interruptible | API calls and tool execution can be cancelled mid-flight by user input or signals |
| Platform-agnostic core | One AIAgent class serves CLI, gateway, ACP, batch, and API server. Platform differences live in the entry point |
| Loose coupling | Optional subsystems (MCP, plugins, memory providers, RL environments) use registry patterns and check_fn gating, not hard dependencies |
| Profile isolation | Each profile gets its own HERMES_HOME, config, memory, sessions, and gateway PID. Multiple profiles run concurrently |
Migration from OpenClaw
Hermes Agent is the successor to OpenClaw. If you’re migrating from an existing OpenClaw installation:65
hermes claw migrate --dry-run # Preview what would be migrated
hermes claw migrate --preset full # Full migration including API keys
hermes claw migrate --preset user-data --overwrite # User data only, no secrets
hermes claw migrate --source /custom/path # Non-default OpenClaw location
hermes claw migrate reads from ~/.openclaw by default (also auto-detects legacy ~/.clawdbot and ~/.moldbot directories) and writes to ~/.hermes.6
Directly imported (30+ categories): SOUL.md, MEMORY.md, USER.md, AGENTS.md, skills from 4 source directories, default model, custom providers, MCP servers, messaging platform tokens and allowlists (Telegram, Discord, Slack, WhatsApp, Signal, Matrix, Mattermost), agent defaults (reasoning effort, compression, human delay, timezone, sandbox), session reset policies (now inert: since v0.21.1, timers never rotate a conversation43), approval rules, TTS config, browser settings, tool settings, exec timeout, command allowlist, gateway config, and API keys from 3 sources.6
Archived for manual review: cron jobs, plugins, hooks/webhooks, memory backend (QMD), skills registry config, UI/identity, logging, multi-agent setup, channel bindings, IDENTITY.md, TOOLS.md, HEARTBEAT.md, BOOTSTRAP.md.6
API key resolution checks three sources in priority order: config values → ~/.openclaw/.env → auth-profiles.json.6
Troubleshooting
“No inference provider configured. Run ‘hermes model’ to choose a provider and model”
The first error every fresh install hits: Hermes has no resolved provider yet. It means exactly what it says — none of the three auth paths has produced a usable provider. Run:
hermes model
The interactive picker walks through every supported provider, including the OAuth device-code flows (Nous Portal, GitHub Copilot, Anthropic, OpenAI Codex) and custom endpoints for self-hosted servers. If you expected a provider to already be configured, hermes doctor shows which credentials Hermes can actually see; the usual causes are an API key set in the wrong place (it belongs in .env or via hermes config set, not your shell profile), an expired OAuth credential in ~/.hermes/auth.json, or a config.yaml custom endpoint that lost its base_url. Auth paths are covered in depth in Authentication & Providers.27
“API key not set”
Run hermes model to configure your provider interactively, or hermes config set OPENROUTER_API_KEY your_key. The hermes doctor command will tell you exactly which keys are missing.7
“Context limit: 2048 tokens” at startup (local models)
Hermes auto-detects context length from your server’s /v1/models endpoint, but many local servers report low defaults. Set it explicitly in config.yaml:2
model:
default: your-model
provider: custom
base_url: http://localhost:11434/v1
context_length: 32768
Tool calls appear as text instead of executing
Your server doesn’t have tool calling enabled, or the model doesn’t support it through the server’s implementation.2
| Server | Fix |
|---|---|
| llama.cpp | Add --jinja to the startup command |
| vLLM | Add --enable-auto-tool-choice --tool-call-parser hermes |
| SGLang | Add --tool-call-parser qwen (or appropriate parser) |
| Ollama | Tool calling is enabled by default — check your model supports it with ollama show <model> |
| LM Studio | Update to 0.3.6+ and use a model with native tool support |
Responses get cut off mid-sentence
Two possible causes:2
- Low output cap (
max_tokens) on the server — SGLang defaults to 128 tokens per response. Set--default-max-tokenson the server or configuremodel.max_tokensinconfig.yaml. - Context exhaustion — The model filled its context window. Increase
model.context_lengthor enable context compression in Hermes.
“Connection refused” from WSL2 to a Windows-hosted model server
WSL2 uses a virtual network adapter with its own subnet — localhost inside WSL2 refers to the Linux VM, not the Windows host. Two options:2
Mirrored networking (Windows 11 22H2+): edit %USERPROFILE%\.wslconfig:
[wsl2]
networkingMode=mirrored
Then wsl --shutdown and restart. localhost now works bidirectionally.
Host IP fallback (older Windows): get the Windows host IP from inside WSL2 and use it instead of localhost:
ip route show | grep -i default | awk '{ print $3 }'
# Use that IP as the base_url host
You also need the model server to bind to 0.0.0.0, not 127.0.0.1 — set OLLAMA_HOST=0.0.0.0 for Ollama, add --host 0.0.0.0 for llama-server/SGLang, or enable “Serve on Network” in LM Studio.2
The iteration budget ignores agent.max_turns
If the activity line reads N/90 (or another stale ceiling) while config.yaml says agent.max_turns: 500, the likely cause is a stale HERMES_MAX_ITERATIONS line in ~/.hermes/.env: the setup wizard used to dual-write the budget to both stores, and if the startup bridge bails on an earlier config-parse error, the .env ghost silently wins. As of v0.21.1, hermes doctor detects the shadowing pair and hermes doctor --fix deletes the .env line, leaving config.yaml authoritative.41
Where is everything?
hermes status and hermes dump are your friends here. hermes logs list shows all log files with sizes. hermes config path prints the config file location. hermes config env-path prints the .env location.6
FAQ
What’s the difference between Hermes Agent and Claude Code?
Claude Code is Anthropic’s official CLI, locked to Anthropic models. Hermes Agent is an open-source agent framework from Nous Research that works with any OpenAI-compatible provider — Nous Portal, OpenRouter, Anthropic, GitHub Copilot, z.ai, Kimi, MiniMax, DeepSeek, Hugging Face, Google, or your own self-hosted endpoint.12 Hermes also ships a messaging gateway for Telegram/Discord/Slack/WhatsApp/Signal that Claude Code does not have.
Can I use Hermes with an Anthropic API key?
Yes. Three ways:2
- Set
ANTHROPIC_API_KEYin~/.hermes/.envand runhermes chat --provider anthropic --model claude-sonnet-4-6 - Run
hermes modeland select Anthropic — Hermes will use Claude Code’s credential store when available - Set a manual
ANTHROPIC_TOKEN(setup-token or OAuth token) as a fallback
Option 2 is preferred if you already use Claude Code on the same machine — it keeps refreshable Claude credentials refreshable.
How do I switch providers without losing my conversation?
Use /model provider:model inside a session. The conversation history, memory, and skills all carry over:9
/model zai:glm-5
/model openrouter:anthropic/claude-sonnet-4
/model custom:local:qwen-2.5
I configured Anthropic but vision/web/compression don’t work
On current builds this mostly cannot happen the old way. By default (auxiliary.*.provider: "auto") every auxiliary task – vision, approval classification, compression, session titles – runs on your main chat model, so an Anthropic-only setup serves them with the OAuth it already has. The old default (Gemini Flash via OpenRouter → Nous → Codex auto-detection, degrading silently when none were configured) is gone: “auto now uses the main model for everyone, and per-task overrides in config.yaml still win.”34
If an auxiliary task still fails, look for an explicit per-task override pointing at a provider you never configured (auxiliary.<task>.provider / .model in config.yaml), or stale legacy keys: as of tag v2026.8.31 the compression summarizer is configured like every other auxiliary slot – auxiliary.compression.provider – and legacy compression.summary_* keys are migrated there automatically (config version 17).31 Web extraction is no longer an LLM task at all (“no LLM involved”), so a web-summarization failure is not an auxiliary-model problem on current builds.34 To pin a task back to your main provider explicitly:
auxiliary:
vision: { provider: "main" }
compression: { provider: "main" }
What is the difference between SOUL.md and AGENTS.md?
SOUL.md is your agent’s identity — tone, style, communication defaults. It lives in ~/.hermes/SOUL.md and follows you everywhere. AGENTS.md is project-specific — architecture, conventions, commands, paths — and lives in your project directory.13 If it should follow you everywhere, SOUL.md. If it belongs to a project, AGENTS.md.
How do I run multiple Hermes instances side-by-side?
Profiles. Each profile gets its own HERMES_HOME, config, memory, sessions, and gateway PID:6
hermes profile create work --clone
hermes profile use work # Sticky default
hermes -p work chat -q "..." # One-off without switching
hermes profile alias work --name h-work # Wrapper script
Does Hermes support local LLMs?
Yes, through the custom endpoint path. Hermes works with any OpenAI-compatible server: Ollama, vLLM, SGLang, llama.cpp/llama-server, LM Studio, LocalAI, Jan, or your own.2 See Custom & Self-Hosted Endpoints for per-server setup.
Why does my startup banner show fewer tools than the guide says Hermes has?
The guide cites 47 tools / 20 toolsets from the upstream architecture registry — that’s the full count of tools Hermes ships source code for across every toolset. Your running install shows a smaller number in the banner (the reference install used for this guide reports 28 tools) because Hermes only enables the default toolset set at startup. Many toolsets are opt-in: messaging gateway adapters, browser automation, heavier scraping stacks, and several specialized integrations have to be explicitly listed under toolsets: in ~/.hermes/config.yaml before they load. Registry total = “what’s available if you enable it.” Banner total = “what your current profile actually loaded.” Use hermes tools --list to see which toolsets are active and which are available but disabled. Toggle individual toolsets at runtime with /tools enable <name> and /tools disable <name> (disabling triggers a session reset so the agent rebuilds its tool manifest with the new shape).
How does Hermes handle model fallback when my primary provider fails?
Configure a fallback_model block in config.yaml:2
fallback_model:
provider: openrouter
model: anthropic/claude-sonnet-4
When the primary fails (rate limit, server error, auth failure), Hermes swaps to the fallback mid-session without losing conversation history. Fires at most once per session.
Can the agent improve its own skills over time?
Yes — that’s the “self-improving” part of Hermes Agent. The agent can create, update, and delete skills via the skill_manage tool. When it figures out a non-trivial workflow, it saves the approach as a skill for future reuse.11 The agent creates skills after complex tasks (5+ tool calls), when it hits errors and finds the working path, when you correct its approach, or when it discovers a non-trivial workflow.
Is there an IDE integration?
Yes — Hermes can run as an ACP (Agent Client Protocol) server for VS Code, Zed, and JetBrains:6
pip install -e '.[acp]'
hermes acp
Changelog
| Date | Change | Source |
|---|---|---|
| 2026-09-24 | Guide v1.23: the v0.14.0 orientation block, which had sat above the current release since the guide’s first version, now closes the What’s-New history as the oldest section, in the newest-first order the rest of the history uses. No content change; the TL;DR now runs straight from Key Takeaways into the source note and Choose Your Path. | Guide structure |
| 2026-09-24 | Guide v1.22: Hermes v0.21.5 (tag v2026.9.24, Sep 24) – third rollup patch, curated notes deferred to v0.22.0. New What’s-New section at the top; TL;DR points to it. Covered at tag depth: Hindsight moved from the core tree to the plugin catalog (absent from the release notes), with its automatic migration; gateway.multiplex_profiles: false retired, per-profile parking, and gateway.standalone; GPT-6 Sol/Luna and Claude Opus 5.5 in the Nous and OpenRouter pickers. External Memory Providers now lists seven bundled providers plus Hindsight from the catalog. Messaging Gateway gains a paragraph on the retired opt-out. Compat box re-checked at the tag: plugin_compat.py unchanged, escape hatch still works. Corrected: the bundled provider-plugin count is 38, not 39, and the OpenCode Free row is marked removed; both changes date from September 18. |
5253 |
| 2026-09-23 | Guide v1.21: correctness and readability pass, still at Hermes v0.21.4 (tag v2026.9.21; no newer release). What’s New now runs newest first, v0.21.3 has its own section, and the TL;DR leads with the current release. Pairing: decline is the only new value; the effective default is ignore once any allowlist is set; the YAML shape is platforms.<name>.unauthorized_dm_behavior or the top-level key. stream-json accepts --query-file, which gains an options row. set-journal-mode notes the stop-everything rule and Windows --force. The host-singleton paragraph now leads with operator behavior. Writer-process sentences removed. |
5051 |
| 2026-09-22 | Guide v1.20: Hermes v0.21.4 (tag v2026.9.21, Sep 21) – the second rollup patch: “5,071 non-merge commits”, “5,169 changed files”, “1,812 merged PRs”, “2,116 closed issues” since v0.21.3 under a thin note that defers curated coverage to v0.22.0. New What’s-New section below the v0.21.3 block. All five headline figures reproduced exactly in the local clone at the stated measurement commit 4b8a8134 (the tag adds one release commit: 5,072 non-merge; compare total 5,173 with merges, matching the GitHub compare API); second-largest tag-to-tag window ever behind v0.21.1’s 5,139 (that claim re-checked, still true). The release’s undocumented-on-purpose list covered ONLY after at-tag source verification, each item placed in its standing section: host singleton (gateway/host_rendezvous.py: one hermes serve + one hermes gateway run per host per OS user, host lock + rendezvous record with (pid, createTime) liveness proof; host_attach.py five outcomes ATTACH/RESCAN/REPLACE_HOST/REFUSE/START; Desktop half host-backend-attach.ts ladder ledger -> HTTP -> token -> WS with a host-level spawn gate – Messaging Gateway gains the paragraph); one backend-owned connector operation (tools/connectors/operation.py “Pure data, no I/O”, 300 s deadline deliberately not a config key, manage_connections setup card on Desktop/TUI/CLI per the at-tag test – Nous free-tier subsection extended); --format stream-json (_parser.py:247-249 + hermes_cli/stream_json.py: system/init -> text/tool_use/tool_result -> one result envelope; requires -q, implies --quiet, refuses --tui, 5,000-char tool cap – chat options table gains the row); skills.auto_load (config_defaults.py:1435, “pinned as fully loaded in every new session (CLI, TUI, gateway, cron, API)”, resolved once at prompt build, missing names warn and skip, HERMES_IGNORE_RULES suppresses – new Pinned Skills subsection); gateway decline (gateway/config.py:139, one polite refusal then 24 h silence per DECLINE_DEDUPE_SECONDS #88028, unauthorized_dm_decline_message, per-platform via platforms.<name>.extra, Email defaults ignore – Pairing section extended); mcp.discovery_concurrency (config_defaults.py:526 default 4, 0 = unlimited #117373, invalid values warn to default, every server still connects – MCP section extended); session_search after/before + OR-relaxed retry (tool schema :708-725, inclusive/exclusive bounds, ISO or 7d/24h/2w; hermes_state_search.py:1151-1163 zero-result ANY-term retry on the unicode61 index, exact-hit semantics kept, OR/NOT/single-term/CJK exempt – Session Search extended, and that section’s stale “Gemini Flash summarization” claim FIXED to the no-LLM single-shape design of v0.15.0 per 34); hermes sessions set-journal-mode delete\|wal (subcommands/sessions.py:177 + sessions_cmd_journal_mode.py, offline self-service for #100896, refuses foreign holders, verifies header bytes 18/19, doctor points at it – sessions row + Session Storage extended); Desktop wave (font field desktop.font_family overriding the theme’s --dt-font-sans, accessibility-first suggestions; “Update engine” one-click runtime update with visible-retry failure; Plugins-hub uninstall behind confirm via plugins.manage remove / Electron loader); video catalogs (plugins/video_gen/fal ltx-2.5 + kling-o3 with at-tag capability strings; toolsets-reference roster row); plugin catalog buildout (plugin-catalog/ 9 -> 228 entries in-window, SHA-pinned YAML admission, per-plugin/per-author website pages with pinned-commit READMEs through an allowlist renderer, added/updated sorting from committer dates; the release’s ten community plugins verified present under actual slugs hermes-tailscale/hermes-ssh/shodan/hermes-terminal/hermes-rss/hermes-resetwatch/done-bell/kiwi/cognee/web-octen – Plugin System extended). The fix-run category (profile/multiplex isolation, cron, kanban, Desktop, state.db) HELD for the v0.22.0 scan per the release’s own deferral. Compat state re-verified: the path-deleting revert STILL not landed – manifest + shims present at v2026.9.21 AND on main at a53b42ddea (2026-09-22, fetched same day); plugin_compat.py window changes are a plugin-scan cache + POSIX-form Windows paths only (#112576), the literal-boolean guard moved :261-268 -> :296-303 still is True; currency refreshed in the v0.21.1 bullet, the v0.21.3 block, the Plugin System box, and 42. Re-stamps at v2026.9.21: static providers 39 (AST count at models_catalog_static.py:311) and provider-plugin dirs 39, both unchanged; the TL;DR provider-count stamp moved v2026.9.14 -> v2026.9.21; v0.21.3’s “current release” wording retired. |
505142 |
| 2026-09-15 | Guide v1.19: Hermes v0.21.2 (tag v2026.9.11, Sep 11) “The state.db Patch Release” + v0.21.3 (tag v2026.9.14, Sep 14), and the plugin-compat deadline aftermath. New What’s-New section below the v0.21.1 section. v0.21.2 (“947 non-merge commits”, “312 merged PRs”, “140 contributors”) leads with the state.db reliability campaign (six PRs, 44 issues: hosted-room coordination out of the root store into shared-state.db – gateway/hosted_rooms.py:398-426 verified at the tag; dashboard opens read-only first; cron guard through the tracked connection registry; doctor --fix refuses unprovably-safe checkpoints; FTS-index damage degrades search instead of fail-closing the turn; corrupt rows render ? instead of killing sessions list; profile-pinned binding; read-only opens stop taking the write lock, 4-20 s -> 0.01 s; the release’s operator framing kept: hermes doctor first, then hermes sessions recover --inspect-only – parser at hermes_cli/subcommands/sessions.py:185-196, and the subcommand PREDATES the window (present at v2026.8.31), so it is documented as newly-covered, not new); multi-profile isolation hardening (#107609-#107630 closed real holes in the “fully isolated” promise carried since v0.19.0 – inherited allow-lists, credentials to the default profile’s host, vault secrets to stdio MCP servers, cross-profile MEDIA: attachments, sibling Nous bearers; honest-framing notes added at Profiles and the Messaging Gateway multiplex paragraph); the password-blind credential vault (sign in / pay / autofill from 1Password, Bitwarden, or the local vault via metadata-only namespaced handles with the password resolved at fill time; master password never a tool argument; TOTP from a saved authenticator key – agent/vault_backends/ + agent/vault_store.py:74-105 verified at the tag); the plugin catalog (hermes plugins browse/search + catalog-aware install + pack install/export/show, SHA-pinned – subparsers verified at the tag; Plugin System command block updated); the Nous free tier + guided first launch (free inference and connectors, /login from a chat, HERMES_GUEST_ONBOARDING=1 where ONLY the literal 1 enables – guest-onboarding-flag.test.ts verified; new subsection under Nous Tool Gateway); and desktop spawn-storm fixes. Small v0.21.3 block (“1,036 non-merge commits”, “338 merged PRs”; tag cut so auto-updating Cloud agents receive it): single-flight token refresh ends refresh-burst session revocation (commit 5dea46d13d, #110061), duplicate state.db writer handles stopped (commit 939a2f64b4, #110934), both verified in the v2026.9.11..v2026.9.14 window only; the release’s own deferral quoted (“Full curated release notes for this window ship with v0.22.0, which will document everything from v0.21.0 onward” … “Nothing in this window is skipped”) and its undocumented-on-purpose list (reasoning-effort pickers, OpenRouter PKCE, HEIF/AVIF, the FAL wave, Slack Agent Sessions API, cross-VM WAL refusal) named in the row and HELD for the v0.22.0 scan. Compat-deadline aftermath rewritten past-tense in the v0.21.1 decomposition bullet, the Plugin System box, and 42: the removal activated ON SCHEDULE via a date gate, not a code revert (hermes_cli/plugin_compat.py:32 at v2026.9.14 sets COMPAT_REMOVAL_DATE; removal_in_effect() at :86-90 true from 2026-09-14), affected plugins now disabled with the red notice – but the revert deleting the old import paths has NOT landed (COMPAT_MANIFEST.md + compat_manifest.json + shims present at the tag AND on main at 5d59366010, fetched 2026-09-15 12:55 PT), so plugins.allow_deprecated_imports: true still keeps affected plugins loading; literal boolean only (plugin_compat.py:261-268, is True – a quoted string never opens the bypass). Changelog-only, release- or source-verified: Telegram bots_require_mention makes bots require an @mention, breaking bot-to-bot loops; hermes -z --resume continues the session (-z = --oneshot, hermes_cli/_parser.py:113); passive update checks hit the GitHub API at most once a day instead of git-fetching every 30 min (banner.py:129-131 at v2026.9.11, :136-139 at v2026.9.14); hermes backup -k/--keep prunes to the newest 3 zips by default (subcommands/backup.py:24-26) and config.yaml backups live in one bounded backups/config/ dir; model_thresholds keys can be provider-scoped as provider:substr (agent/context_compressor.py:1558-1567); --clone-all no longer copies cron jobs (the flag lives on hermes profile, subcommands/profile.py:24, NOT on hermes cron); kanban promote refuses undone parents and kanban_request_review rejects unknown reviewer profiles; /model and auxiliary auto never bill a provider you did not select and never auto-switch to one without credentials; pickers add DeepSeek V4.1 Flash (Nous Portal + OpenRouter), GPT Image 2.5, and Opus 5 + Fable 5.1 on the native Anthropic picker; debug share retention shrunk to 1 day on the dpaste fallback. Re-stamps at v2026.9.14: static providers 39 (AST count of CANONICAL_PROVIDERS at models_catalog_static.py:314) and provider-plugin dirs 39, both unchanged; --max-turns default 500 now at cli.py:404 (41’s cli.py:400 stays correct for its tag); model_catalog.ttl_minutes 20 (config_defaults.py:1866); v0.21.1’s largest-window claim (5,139) still true against the two new windows (959 and 1,037 local non-merge counts); the TL;DR provider-count stamp moved v2026.9.7 -> v2026.9.14. |
484942 |
| 2026-09-08 | Guide v1.18: Hermes v0.21.1 (tag v2026.9.7, Sep 7) – the rollup patch: the largest single tag-to-tag window yet (“5,139 non-merge commits”, “632 merged PRs”) under a deliberately thin patch note that defers curated notes to v0.22.0. New What’s-New section above the Pantheon section, six source-verified clusters: the September decomposition + the 2026-09-14 plugin-compat deadline (COMPAT_MANIFEST.md: 1,148 moved-lazy names, HermesPluginCompatWarning once per name, the hermes plugins compat checker, plugins.allow_deprecated_imports escape hatch; a compat-window box added to Plugin System), explicit-only gateway conversation boundaries (SessionResetPolicy now inert – Messaging Gateway updated, claw-migration list annotated), MCP device-code OAuth (hermes mcp login --flow device, RFC 8628; login and reauth --all added to the MCP command block; profile ownership through OAuth sessions, malformed metadata caches ignored, desktop client-local callbacks, -t filters MCP server spawning), delegation reliability (completion units via delegation.independent_completions with per-task group, one completion per call by default; background-process handoff with orphaned_processes and unread_completions named on results; delegation.fallback_providers; normalizer-validated child chains; crash-durable partials; children never inherit the 1h cache tier – all read from the delegate tool source at the tag), providers/models (GPT-6 Astra + Astra Pro with fast/flex tiers, account-gated on Codex OAuth with an opt-in -900k variant; claude-fable-5.1; gemini-3.7/3.8-flash; qwen3.8-max-0902; Muse Spark 1.3 + the muse-image image_gen plugin; Tavily search/extract; the managed llama.cpp runtime; external-process providers; 20-minute catalog refresh via model_catalog.ttl_minutes), and the desktop wave (in-app browser comment mode carrying selector/markup/styles per annotation with region-grouped batches; structured session + automation controls; drag-to-create sessions; foreign-transcript session import; display.resume_last_session; first-open real-profile consent; built-in optional-skills catalog; Russian desktop locale). Changelog-only, commit- or source-verified: cron reliability (restart handoff hardened across three verified-fix commits, delivery dedup serialized with terminal retention, paused-job creation race closed, Discord cron media routed to its target with upload failures reported, continuity preserved across silent audit ticks), approvals shell parsing (GNU env split escapes and argv0 operands, env argv and shell-comment boundaries, quoted command-substitution bodies keep their command boundaries, approvals.deny inside isolated containers), the gateway startup-liveness watchdog (hermes_startup_watchdog.py; gateway.startup_watchdog with a 300s deadline, hard-exit 75), state.db resilience (sqlite3-.recover lost-and-found salvage, like_scan FTS routing, new docs/state-db-recovery.md), perf (search_files runs ripgrep natively on local POSIX hosts, delegate-child transcripts excluded from the trigram FTS index at schema v30, a shared OpenAI httpx client across main and auxiliary paths), gateway.trust_env, the Slack Block Kit /model picker, and remote-sandbox media delivery (credentials and symlinks to them never leave the sandbox); Troubleshooting gains the stale-HERMES_MAX_ITERATIONS-ghost doctor check. At-tag re-stamps v2026.8.31 -> v2026.9.7: static providers 39 and provider-plugin dirs 39, both entry-identical to v0.21.0 – the pre-scan claim of “38 static, xai absorbed into the plugin” was refuted at source (the xai tuple is still in the static list; only the file moved, hermes_cli/models.py -> hermes_cli/models_catalog_static.py:311); prompt_builder.py re-pathed to agent/prompt_builder.py in the mental model and the CLI data flow; v0.16’s hourly catalog-refresh claim kept version-scoped with the 20-minute cadence added beside it. Standing claims re-verified, unchanged: the hermes approvals test naming note, --max-turns default 500 (cli.py:400 at the tag; the stale docstring saying 60 was not propagated), read_file 2000 lines, delegation caps 250/10, compression.tail_mode: lean, Node 26, 17 CLI locales. |
41424344454647 |
| 2026-07-28 | Guide v1.12: Search-demand-led coverage pass — two well-converting topics had no heading to land on. No new release. GSC shows hermes swarm and hermes agent swarm converting at 4.4–6.3% from position ~8, and hermes smart model routing at 4.9% from position 6.7, while neither term had a section: swarm existed only as prose inside Multi-Agent Kanban and in changelog rows, and “smart model routing” appeared nowhere but a footnote. Added What Is a Hermes Swarm? under Multi-Agent Kanban — defines a swarm as parallel workers over one durable board, documents v0.15.0’s swarm topology, auto-decomposition, per-task model overrides, scheduled tasks, and worktree management, and tables the failure each mechanism prevents. Retitled Provider Rotation & Fallback to Smart Model Routing: Provider Rotation & Fallback with a lead that ties credential pools, fallback model, and auxiliary routing together as one system. No anchor was linked internally, so the retitle breaks nothing. |
59 2 6 |
| 2026-08-31 | Guide v1.17: Hermes v0.21.0 “The Pantheon Release” (tag v2026.8.31, Aug 31) – the curated rollup lands. New What’s-New section above the Herald section, organized by the release’s own feature areas: Bot Mode bundled and default-on in desktop (named profiles, deterministic avatars, group chats with @-mentions), hermes peer bot-to-bot DMs (replies land in each agent’s canonical Bot Chat), cron continuity (continuity=true, durable notepads, monitor-mode LLM skip, per-job effort pinning – the Scheduled Tasks section gains the at-tag mechanics of all four), live subagent orchestration (delegate_task list/steer/stop, JSON-schema child outputs, defaults raised to 250 iterations / 10 concurrent children – source-verified in config_defaults.py, against which the delegation docs page’s 3/50 is stale), the MCP command center (+hermes:// install links – MCP section updated), the CLI power wave (/palette + Ctrl+P and a shared-registry /status added to the slash tables, filter-as-you-type /model noted; the release-notes name hermes approval-check shown to be hermes approvals test at the tag – no such subcommand exists), the agent-driven in-app browser, the providers/models wave (six new providers; the matrix gains Tencent TokenPlan, Nebius Token Factory, Ramp Router, and Alibaba Token Plan rows with docs-quoted env vars; model_overrides; data-training-tier warnings; pip provider plugins), the security wave (AGENTS.md/skills/memory writes always require approval – security.protected_instruction_files: true and the tools/file_tools.py gate verified at the tag; redaction sweep; Windows approval coverage; hermes desktop --setup-tcc-identity; Blender MCP removed – Security Hardening gains a fourth wave), gateway maturation, the 8-skill wave, and the REVERTED list (Model Council /council, DCP context engine, WS-only gateway server – seq-stamped replay #94219 DID ship; Electron back to 40.10.2; grep confirms the guide never documented any of the reverted features). The rollup subsection’s deferred-notes framing is RESOLVED and now points at the curated section; window blocks kept as the per-tag record. At-tag re-stamps v2026.8.27 -> v2026.8.31: static providers 38 -> 39 (+tencent-tokenplan), provider-plugin dirs 37 -> 39 (+nebius-token-factory, +router, both auth_type="api_key"), docs provider rows 41 -> 45; unchanged and re-verified: 28 docs platform rows / 24 Platform enum members (identical list) / 22 adapter dirs, _BUILTIN_SUBCOMMANDS identical (73, peer present since the rollup windows, no version), terminal backends 7 + plugin registry (registry byte-identical), 14 personalities and 17 locales (byte-identical/identical lists), compression lean default and the auxiliary slot roster, _startup_fast.py/portal_cli.py byte-identical. Drift fixed while verifying: hermes chat -q on a TTY now seeds an interactive session (new --oneshot restores answer-and-exit) and the chat table’s --max-turns default corrected 90 -> 500 per the at-tag parser help. |
35363738394033 |
| 2026-08-27 | Guide v1.16: Hermes v0.20.6 (tag v2026.8.27, Aug 27) – fourth rollup. The rollup subsection gains its fourth window, per the release’s own description (~1,313 commits across ~1,557 files, +177,113 / −21,682 – ~525 merged PRs since v0.20.5): consent-gated real-profile browsing (default Chromium profile, Windows close-with-approval flow), the desktop Browser in its own OS window, a managed SSH remote-update engine and fleet profile rail, remote MCP catalog expansion (50+ live-verified vendor-hosted servers incl. Cloudflare, Grafana Cloud, Better Stack, Railway), opt-in OS-keychain encryption for stored secrets, new picker models (GLM-5.3-Flash, MiniMax M3 free, MiniMax H3 Max video), TTL result caching for web_search/web_extract, multi-query tool_search with stemming, updaters pausing gateways over the control socket, image/package-managed installs refusing unsafe in-place updates, cron durable-incident acks, Slack link-unfurl controls, and shared Docker container identities. Two release claims contradicted standing sections; both confirmed at the tag and fixed: lean-tail compression is the default (Context Compression rewritten with at-tag keys – tail_mode: lean, threshold_tokens, protect_first_n, summarizer knobs under auxiliary.compression.* with the config-version-17 auto-migration; the inverted troubleshooting note fixed) and terminal backends are plugin-extensible (restyled as seven built-in backends plus a plugin registry, provider-picker style; built-in names reserved). At-tag stamps re-verified and moved v2026.8.19 → v2026.8.27: 38 static providers + 37 provider-plugin dirs, 24 Platform enum members / 22 adapter dirs / 28 docs-table rows, no version subcommand (73 _BUILTIN_SUBCOMMANDS entries, worktree present), parser options and quoted help strings unchanged, BUILTIN_PERSONALITIES still 14 (helpful through hype); locales/ now 17 catalogs (en + 16 translations; at-tag count added). main sits AT the tag (ahead_by: 0, status identical). Also fixed, drift found while verifying: the auxiliary system’s stale default – auto now routes every auxiliary task to the main chat model, not Gemini Flash via OpenRouter → Nous → Codex detection; Auxiliary Models rewritten with the at-tag slot roster (title_generation, tts_audio_tags, triage_specifier, kanban_decomposer, profile_describer added with sourced one-liners; web_extract and session_search slots removed upstream – neither uses an LLM any more – and flush_memories gone from the defaults), and the Anthropic-only troubleshooting entry re-grounded. The release’s fresh-install line uses the raw GitHub script URL; the guide keeps the canonical hermes-agent.nousresearch.com/install.sh. Curated notes still promised for v0.21.0. |
293031323334 |
| 2026-08-26 | Guide v1.15: CORRECTION – no new release (newest tag v2026.8.19 / v0.20.5, Aug 21); standing claims re-verified at the tag. Terminal backends: six → seven, vercel_sandbox added to the table and the config comment. Messaging platforms: the stray “22” counts reconciled to the docs comparison table’s 28 at the tag, with the counting basis stated (24 built-in Platform enum members, 22 bundled adapter directories) and ntfy and Buzz added to the gateway list. Providers: “~20 / ~22 first-class” replaced with the sourced count (38 static CANONICAL_PROVIDERS entries plus auto-extension from 37 bundled provider plugins; 41 cloud/subscription rows on the docs page), “complete list” dropped, the Qwen OAuth row corrected, and 17 matrix rows added (OpenCode Free, OpenAI API direct, Vertex AI, Azure Foundry, Bedrock, NVIDIA NIM, Ollama Cloud, StepFun, MiniMax OAuth, Meta AI, NovitaAI, Arcee AI, GMI Cloud, Actual Computer, Tencent TokenHub, CommandCode, Alibaba Coding Plan). hermes version → hermes --version (not a subcommand at the tag). hermes honcho annotated as plugin-conditional. “fleet --plan” → hermes update --plan. Global Options gains --in, --tui, --cli, --dev, --ignore-rules, --ignore-user-config; Top-Level Commands gains 34 rows from _BUILTIN_SUBCOMMANDS, including hermes worktree list\|prune with its flags; /worktree slash command added; hermes setup --portal and hermes portal login\|info\|open\|tools documented. main sat 1,104 commits past the tag at verification time (all since rolled up into v2026.8.27); curated notes are promised for v0.21.0. |
23242526272829 |
| 2026-08-24 | Guide v1.14: v0.20.5 (tag v2026.8.19, published Aug 21). The rollup subsection gains its third window (~746 commits / ~323 PRs since v0.20.4): the keyless web tier (5-vendor free rotation with ring failover — web search on fresh installs with zero keys), CLI polish (fuzzy /model picker, Ctrl+P command palette, richer /status), Bot Mode group-room threads + foldable summaries + PDF/file drag-and-drop, hermes worktree list/prune, hermes update receipts and fleet --plan verification, cron persistent memory with per-job reasoning effort, execution-discipline stall guards from the Composio eval findings, the opencode-free zero-auth provider, and desktop perf (paint-first hydration, React Compiler in both renderers). v0.21.0 with full curated notes is still pending - summaries sourced from the release’s own window description. Header lineage updated. |
23 |
| 2026-08-20 | Guide v1.13: v0.20.3 (tag v2026.8.16.2, published Aug 17) and v0.20.4 (tag v2026.8.18, Aug 18). New subsection under the Herald Release: the rollup train is now feature-bearing, not just stabilization. v0.20.3 (~250 commits / ~125 PRs): MCP 2.x SDK migration with 2026-07-28 stateless-protocol support, bundled Bot Mode plugin (hermes-bots) with the core teammate protocol, CommandCode provider plugin, Cua Driver 0.20 computer-use runtime contracts, Python runtime ownership hardening, cron scheduler self-heal, session-handoff data-loss fixes, ecosystem ports (/worktree, /rollback hand-edit preservation, plugin install security scanning). v0.20.4 (~146 commits / ~74 PRs): desktop glass/translucency surface with frost picker, tabbed SESSIONS|BOTS sidebar with per-bot hide/unhide, NVIDIA SkillEvaluator Tier 1 advisory scanning on skill installs (license + security), cron media-send hardening, hermes update parked-branch honesty. Both releases state full curated notes ship with v0.21.0 — summaries sourced from the releases’ own window descriptions. Header sentence and tag lineage updated. |
54 |
| 2026-08-16 | Guide v1.12: v0.20.0 “The Herald Release” (Aug 3, tag v2026.8.3), plus v0.20.1 (Aug 13) and v0.20.2 (Aug 16) stabilization tags. Three corrections fix instructions that no longer work: Node 26 is required (installer pins NODE_VERSION="26" and rejects older runtimes — the docs site’s Node v22 line is stale, so the installer and release notes win), pip and Homebrew are retired rather than deprecated (“shell installer / Docker / Nix are the supported channels”), and the default iteration limit moved 90 → 500, which invalidated every number in the budget-pressure table. Install command corrected to the canonical https://hermes-agent.nousresearch.com/install.sh. Removed surface: the claude-marketplace skill source is gone, replaced in the source list by browse-sh; default GitHub taps are now openai, anthropics, huggingface, NVIDIA, and gstack. Messaging gateway recounted 22 → 28 platforms by enumerating the docs comparison table directly (the docs publish no official total). Windows (native) is Tier 1, not early beta; macOS is Apple Silicon only. New section covers the release itself: conversational voice with barge-in, A2A v1.0, signed outbound lifecycle webhooks, grounded-citations skill, the !//init//diff//context//focus CLI wave and hermes import-agent, command-helper secret source, hermes -w cold start ~14s → ~1.8s, and desktop artifacts plus a Plugin SDK. Verified unchanged: the three auth paths, ~/.hermes/ layout, hermes update, and the documented tool list. |
55 |
| 2026-07-21 | Guide v1.11: v0.19.0 “The Quicksilver Release” (July 20, 2026, tag v2026.7.20). Added the “What’s New in v0.19.0” section: ~80% first-turn TTFT cut (cold submit→dispatch ~4.3s → ~0.9s across CLI/gateway/TUI/desktop/cron), reasoning streamed live by default (display.show_reasoning ON), desktop ~20-PR speed wave (14× faster streaming markdown) + TUI incremental markdown; pip/Homebrew installs deprecated (warn-only “unsupported legacy,” PyPI/Homebrew publishing removal planned) — corrected the Installation section and TL;DR to the one-line installer; pluggable SecretSource with Bitwarden + 1Password providers (op:// refs, multi-vault, deterministic precedence, per-variable provenance); smart approvals default (independent LLM reviewer per flagged command) + user-defined deny rules that hold under YOLO + /deny <reason> + plugin pre_tool_call approve escalation re-landed; terminal billing /subscription + /topup + desktop billing tab (retired the “no separate subscription command” claim); live subagent transcripts + durable background delegation + delivery-obligation ledger in state.db; max_async_children deprecated for unified delegation concurrency caps; gateway profile-based message routing (one multiplexed bot token → isolated profiles, GATEWAY_MULTIPLEX_PROFILES, routing index in state.db, sessions.json optional legacy mirror); providers/models: Fireworks AI (#2 picker slot), DeepInfra, Upstage Solar, GPT-5.6 (Sol/Terra/Luna + Pro) end-to-end, grok-4.5 GA, kimi-k3 (kimi-k2.x retired), Claude Sonnet 5 fully wired, per-provider enabled: false + excluded_providers, reasoning effort max/ultra tiers with per-model/per-MoA-slot overrides and session-scoped /reasoning; CLI/MCP: hermes sessions export (Markdown/Quarto/HTML/prompt-only/HF-trace, --redact), /model --once, stacked slash-skill invocations, --safe-mode, hermes config get/unset, hermes serve true headless, MCP mcp__server__tool naming. Also logs the missed patch tags v0.18.1 (tag v2026.7.7) and v0.18.2 (tag v2026.7.7.2), July 7–8, 2026 — infrastructure patch rollups; v0.18.2’s substantive fix unpins WhatsApp Baileys to 7.0.0-rc13 for reliable Docker builds. |
56 57 |
| 2026-07-16 | Added a first troubleshooting entry for the verbatim startup error “No inference provider configured. Run ‘hermes model’ to choose a provider and model” — search-demand-led; points to the interactive picker, hermes doctor, and the three auth paths. No product changes. |
2 7 |
| 2026-07-01 | Guide v1.10: v0.18.0 “The Judgment Release” (July 1, 2026, tag v2026.7.1). Added the “What’s New in v0.18.0” section: full P0/P1 backlog closed (~692 items); Mixture-of-Agents first-class with labelled per-model ensemble output and live streaming; completion contracts — /goal verifies its own work by running project checks; /learn (describe a workflow → reusable skill, CONTRIBUTING.md-compliant); /journey memory/skill timeline + desktop memory graph; background subagent fan-out (concurrent delegated tasks); Desktop Projects (project/repo/lane); scale-to-zero gateway with drain coordination; Google Vertex AI (Gemini via GCP service accounts, auto OAuth2 refresh); /prompt \$EDITOR composer. Source: hermes-agent releases. |
22 |
| 2026-06-21 | Guide v1.9: v0.17.0 “The Reach Release” (June 19, 2026, tag v2026.6.19). Added the “What’s New in v0.17.0” section. Messaging: relay-free iMessage via Photon Spectrum (hermes photon login, device-code OAuth), official WhatsApp Business Cloud API adapter (no bridge), SimpleX groups + attachments, Raft platform plugin. Models: z-ai/glm-5.2 (1M), anthropic/claude-fable-5, laguna-m.1, nemotron-3-ultra, grok-composer-2.5-fast (xAI OAuth, 200k); xAI default → grok-build-0.1; Anthropic adaptive models drop the reasoning field. Desktop/dashboard: background subagents with live watch-windows (delegate_task(background=true)), full profile builder, reworked Skills Hub, Automation Blueprints, secure 401 login, VS Code Marketplace themes, Japanese + Traditional Chinese UI. Skills/tools: image_generate image-to-image editing, memory atomic operations batch, simplify-code skill, boolean write_approval (replaces write_mode). Architecture: MCP elicitation handler, pluggable CronScheduler + Chronos, Managed scope (/etc/hermes), Gateway-Gateway relay. Commands: /version, /billing, hermes curator run --consolidate (opt-in). Security: shell-escape denylist bypass closed, fail-closed approval/gateway adapters, cron env sanitized, secrets redacted in debug dumps, MCP stdio exfil screening, urllib3 + PyJWT CVE bumps. |
21 |
| 2026-06-08 | Guide v1.8: v0.16.0 “The Surface Release” (June 5, 2026, tag v2026.6.5). Retitled the guide to v0.16 and added the “What’s New in v0.16.0” section. Headline: Hermes is no longer terminal-only. Native Hermes Desktop app (Electron, macOS/Linux/Windows) with one-click install, in-app self-update, streaming chat, drag-and-drop + clipboard image paste, Cmd+K palette, session archive/search, status-bar model picker, remote-gateway connect over secure WebSocket (OAuth or user/pass, per-profile hosts, cross-profile @session links), and a full Simplified Chinese translation via typed i18n. Browser admin panel (web dashboard → full admin): MCP catalog enable/disable, credential management, webhook/hook creation, memory config, gateway controls, System page with check-before-update + Debug Share, new Channels page, and pluggable auth (user/pass, self-hosted OIDC, hermes dashboard register). New commands: /undo [N] (CLI/TUI/messaging), configurable default interface (cli/tui, --cli), TUI unified /model + Sessions overlay, hermes portal, hermes prompt-size, hermes sessions optimize. New models: deepseek-v4-flash, MiniMax-M3 (1M context), qwen3.7-plus, gemini-3.5-flash; first-class xAI Grok OAuth in the desktop launcher; fuzzy model picker; hourly catalog refresh. Skills: leaner default set (Spotify → native plugin, Linear → hermes mcp install linear, dead skills removed), environments: relevance gate (kanban/docker/s6), NVIDIA/skills default trusted tap, progressive (scoped) MCP/plugin tool disclosure. Security: CVE-2026-48710 (Starlette BadHost) pinned ≥1.0.1; SSRF checks off the event loop; Bedrock bearer token stripped from subprocess env; bws_cache.json read-guarded; docker restart/stop/kill added to dangerous patterns; invisible-unicode sanitization. Closed 2 P0 + 62 P1 (16 security-tagged). |
20 |
| 2026-05-31 | Guide v1.7.1: v0.15.1 (May 29, 2026, 01:12 UTC) — Velocity patch. Same-day post-Velocity hotfix; pinned tag v2026.5.29 line. Fixes the dashboard 401 reload loop affecting loopback mode deployments. Docker no longer treats --insecure as implicit — set HERMES_DASHBOARD_INSECURE=1 explicitly to opt back in. MCP bare commands (npx, npm, node) resolve correctly inside Docker containers again. Skills page source pills and category sidebar render. Kanban workers respond to SIGTERM cleanly instead of orphaning processes. Skills.sh catalog expanded from 858 to 19,932 entries via sitemap discovery. 28 commits, 21 merged PRs, 9 contributors. v0.15.2 (May 29, 2026, 13:37 UTC) — Velocity packaging patch. Fixes wheel and sdist distributions to bundle plugin.yaml manifests so installs from PyPI work without sideloading the source tree. Packaging-only hotfix, 4 contributors. |
58 |
| 2026-05-28 | Guide v1.7: Added v0.15.0 (May 28, 2026) — The Velocity release (tag v2026.5.28). Headline: a massive refactor pass + new orchestration primitives. Codebase refactoring: run_agent.py reduced 76% (16,083 → 3,821 lines), distributed across 14 cohesive modules. Multi-agent Kanban v2: auto-decomposition of high-level goals into subtasks, swarm topology for parallel worker coordination, per-task model overrides, scheduled tasks, worktree management. Performance: additional second saved on cold start; 47% reduction in per-conversation function calls; session_search redesigned 4,500× faster with the LLM dependency removed (and its API cost eliminated). Security: Promptware defense protects against Brainworm-class prompt injection at three security chokepoints; Bitwarden Secrets Manager integration replaces multiple per-provider API keys with a single bootstrap token. Skill bundles: load multiple skills simultaneously with one slash command. TUI session orchestrator: multi-session management within a single terminal window. New providers: Krea 2 (Medium/Large) and FAL plugin support for image generation; xAI integration round adds a web-search plugin, OAuth upstream, retired-model detection, and natural TTS pauses. Stats: 1,302 commits, 747 merged PRs, 321 community contributors. Per GitHub release notes, a same-day-or-following patch release addresses dashboard 401 reload-loop, Docker --insecure explicit env var, MCP bare command resolution in Docker (npx, npm, node), Skills page restoration, Kanban worker SIGTERM handling, and the full 19,932-entry Skills catalog via sitemap. |
59 |
| 2026-05-21 | Guide v1.6: Added v0.14.0 (May 16, 2026) — The Foundation release. Headline: lighter install/runtime foundation plus broader provider, gateway, media, and verification surfaces. Added SuperGrok OAuth with grok-4.3 1M context, OpenAI-compatible hermes proxy for OAuth providers, first-class x_search, pip install hermes-agent, lazy dependency installs, ~19s faster launch, 180x faster browser CDP calls, LINE + SimpleX Chat for 22 messaging platforms, Microsoft Teams end-to-end, /handoff, /subgoal, native clarify buttons on Telegram/Discord, Discord history backfill, raw-pixel vision_analyze, per-turn file-mutation verifier footer, LSP semantic diagnostics on every write, unified video_generate, computer_use via cua-driver for non-Anthropic providers, OSC8 clickable URLs, Zed ACP Registry support, OpenRouter Pareto Code router, NovitaAI, Codex app-server runtime, huggingface/skills trusted tap, 9 optional skills, plugin ctx.llm / tool_override, Brave/DDGS web search, Qwen Cloud rename, native Windows beta, and 12 P0 / 50 P1 closures. |
19 |
| 2026-05-07 | Guide v1.5: Added v0.13.0 (May 7, 2026) — The Tenacity release. Headline: a durable multi-agent Kanban board (heartbeat, reclaim, zombie detection, hallucination gate, per-task max_retries, multi-project boards) that turns swarms into a first-class primitive instead of a delegation pattern. /goal command locks the agent onto a target across turns (Ralph-loop pattern as a slash command). New video_analyze tool, Gemini-first with extensible compatible-model support. xAI Custom Voices TTS provider with voice cloning. 7-language i18n (zh-Hans, ja, de, es, fr, uk, tr) for CLI and gateway messages; docs zh-Hans only. Google Chat as 20th messaging platform via the pluggable adapter pattern; IRC + Microsoft Teams migrated onto the same pattern. ProviderProfile ABC + plugins/model-providers/ for pluggable third-party providers without core changes. Session auto-resume across gateway restart, /update, and source-file reload. Checkpoints v2 rewrite with single-store design, real pruning, and disk guardrails. Eight P0 security closures: secret redaction default-on, Discord cross-guild DM bypass (CVSS 8.1), WhatsApp stranger-reject + self-chat-mute, MCP OAuth TOCTOU, CLI auth.json TOCTOU, browser SSRF floor, cron prompt-injection scanning, hermes debug share redaction. Post-write linting for Python/JSON/YAML/TOML, cron no_agent script-only mode, platform allowlists across Slack/Telegram/Mattermost/Matrix/DingTalk, MCP enhancements (SSE transport, OAuth forwarding, image MEDIA tags). Stats since v0.12.0: 864 commits, 588 merged PRs, 829 files changed, 295 community contributors, 282 issues closed (13 P0, 36 P1). |
18 |
| 2026-05-06 | Guide v1.4: Added v0.12.0 (April 30, 2026) — The Curator release. Headline: an autonomous background Curator running on the gateway’s cron ticker (7-day default cycle) that grades the skill library on a rubric, prunes dead skills, consolidates related skills, and writes per-run reports — Hermes maintains itself between active sessions. Self-improvement loop upgraded with rubric-based grading, active-update bias, proper runtime inheritance, and scoped toolsets restricted to memory and skills. Four new inference providers: GMI Cloud, Azure AI Foundry, MiniMax OAuth, and Tencent Tokenhub. LM Studio promoted to first-class. Remote model catalog manifests now auto-update without releases. Two new messaging platforms: Microsoft Teams (19th, via pluggable gateway architecture) and Tencent Yuanbao (18th, native text + media). Native Spotify via PKCE OAuth with bundled skill; Google Meet plugin for calls and transcription; Piper local TTS provider. ComfyUI v5 + TouchDesigner-MCP moved from optional to bundled-by-default. New skills: Humanizer, claude-design, design-md, airtable. CLI additions: hermes -z one-shot mode, hermes update --check preflight, /reload-skills slash command, pluggable busy-indicator styles. Visible TUI cold start cut ~57% via lazy agent init and lazy imports. Security: secret redaction disabled by default to prevent payload corruption; hardline blocklist for unrecoverable commands. Stats: 1,096 commits, 550 merged PRs, 213 community contributors. |
17 |
| 2026-04-25 | Guide v1.3: Added v0.11.0 (April 23, 2026) — The Interface release. Full React/Ink rewrite of the interactive TUI with a Python JSON-RPC backend (tui_gateway); sticky composer, live streaming with OSC-52 clipboard support, stable picker keys, status bar with per-turn stopwatch and git branch, /clear confirm, light-theme preset, subagent spawn observability overlay. Pluggable transport architecture — format conversion and HTTP transport extracted into agent/transports/ for cleaner provider plumbing. Native AWS Bedrock via the Converse API. Five new inference paths: NVIDIA NIM, Arcee AI, Step Plan, Google Gemini CLI OAuth, and Vercel ai-gateway. GPT-5.5 via Codex OAuth — the new OpenAI flagship is now reachable through ChatGPT Codex OAuth without a separate API key. QQBot (17th messaging platform) with QR-scan setup and streaming. Plugin surface expansion: slash commands, tool dispatch, execution blocking, result transformation. /steer <prompt> — mid-run agent nudges that inject a note the running agent sees after its next tool call, without interrupting the turn or breaking the prompt cache. Shell hooks wire scripts as lifecycle hooks without Python plugins. Webhook direct-delivery mode forwards payloads straight to a platform chat, bypassing the agent for fan-out. Smarter delegation with orchestrator roles, configurable spawn depth, and file coordination. Dashboard gains a plugin system, live theme switching, i18n, and mobile responsiveness. Stats since v0.9.0: 1,556 commits, 761 merged PRs, 1,314 files changed, 224,174 insertions, 29 community contributors. |
60 |
| 2026-04-16 | Guide v1.2: Added v0.10.0 — Nous Tool Gateway. Paid Nous Portal subscribers now access managed tools (Firecrawl web search, FAL / FLUX 2 Pro image generation, OpenAI TTS, Browser Use browser automation) with no extra API keys. Opt-in per tool via new use_gateway config field. Runtime prefers gateway over direct API keys when both are configured. HERMES_ENABLE_NOUS_MANAGED_TOOLS env var removed. Hermes Agent CLI remains MIT-licensed and fully free. |
61 |
| 2026-04-13 | Guide v1.1: Added v0.8.0 and v0.9.0 features. Local web dashboard, /fast mode, iMessage + WeChat platforms (16 total), background process monitoring (watch_patterns), pluggable context engine, hermes backup/hermes import, Termux/Android, xAI + MiMo + Google AI Studio + Qwen providers, /debug command, comprehensive security hardening. |
15 16 |
| 2026-04-10 | Guide v1.0: Initial release covering Hermes Agent v0.7.0. Provider auth, config, CLI, slash commands, tools, skills, memory, gateway, cron, MCP, compression, architecture, OpenClaw migration, troubleshooting, FAQ. |
References
-
Nous Research, “Hermes Agent” project README on GitHub. Primary source for the product description (self-improving agent, multi-provider, messaging gateway, terminal backends, skill evolution, cron scheduler, delegation) and the “Quick Install” one-liner. ↩↩↩
-
Nous Research, “AI Providers” in the Hermes Agent documentation. Primary source for the full provider list, auth methods per provider (Nous Portal OAuth, Codex device code, GitHub Copilot token types, Anthropic three-method auth, Chinese AI providers, Hugging Face routing, custom endpoints), the three auth paths (API key in
.env, OAuth viahermes model, custom endpoint inconfig.yaml), the/modelslash command syntax (includingcustom:name:model), Ollama/vLLM/SGLang/llama.cpp/LM Studio setup templates, WSL2 networking instructions, context length detection chain, fallback model configuration, smart model routing, and named custom providers. All provider-specific environment variable names, token types, base URL overrides, and model identifiers in this post come from this page. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Architecture” in the Hermes Agent developer guide. Primary source for the system overview diagram, directory structure, data flow through CLI session and gateway message paths, the three API modes (
chat_completions,codex_responses,anthropic_messages), provider resolution viaruntime_provider.py, session persistence via SQLite + FTS5, messaging gateway platform list, plugin system discovery sources, profile isolation, and the six design principles. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Configuration” in the Hermes Agent user guide. Primary source for the configuration directory structure, the
config.yamlvs.envrule (“config.yamlwins for non-secret settings”), the configuration precedence chain (CLI args → env → config.yaml → .env → defaults), context compression settings (compression.*block withthreshold,threshold_tokens,target_ratio,tail_mode,protect_last_n,protect_first_n; the summarizer’s model/provider/endpoint live underauxiliary.compression.*since the config-version-17 migration), budget pressure thresholds (70% caution, 90% warning), streaming timeouts with local provider auto-adjustment, and the full auxiliary model configuration block (auxiliary:withvision,web_extract,approval,compression,session_search,skills_hub,mcp,flush_memoriesslots). The"main"provider restriction to auxiliary/compression/fallback slots is also from this page. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Migrate from OpenClaw” in the Hermes Agent guides. Source for the OpenClaw → Hermes migration flow. ↩↩
-
Nous Research, “CLI Commands Reference” in the Hermes Agent reference documentation. Primary source for every top-level CLI command documented in this post, including
hermes chat,hermes model,hermes gateway,hermes setup,hermes auth,hermes status,hermes cron,hermes webhook,hermes doctor,hermes dump,hermes logs,hermes config,hermes pairing,hermes skills,hermes honcho,hermes memory,hermes acp,hermes mcp,hermes plugins,hermes tools,hermes sessions,hermes insights,hermes claw,hermes profile,hermes completion,hermes update, andhermes uninstall. All subcommand flags, option descriptions, credential pool behavior, log filtering syntax, OpenClaw migration flags, profile management commands, and service installation commands in this post come from this page. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Installation” in the Hermes Agent getting-started guide. Primary source for the one-line installer command, the installer’s behavior (prerequisites, platform support, Termux auto-detection, Windows/WSL2 requirements), the optional extras table, the manual installation steps, and the verification commands. ↩↩↩↩↩↩↩↩↩
-
Nous Research, “CLI Commands Reference” — see specifically the
hermes dumpsection describing the command’s output format (header, environment, identity, model, terminal, API keys, features, services, workload, config overrides) and intended use for sharing diagnostics. ↩ -
Nous Research, “Slash Commands Reference” in the Hermes Agent reference documentation. Primary source for every slash command listed in this post, the
COMMAND_REGISTRYarchitecture, the CLI vs messaging split, dynamic skill slash commands, quick commands inconfig.yaml, prefix matching behavior, and the messaging-only commands (/status,/sethome,/approve,/deny,/update,/commands). ↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Tools & Toolsets” in the Hermes Agent user guide. Primary source for the tool category overview, toolset usage commands, the seven terminal backends (local, docker, ssh, singularity, modal, daytona, vercel_sandbox), container configuration (cpu, memory, disk, persistent), security hardening for containers, background process management API, and sudo support. ↩↩↩↩↩↩↩↩↩↩
-
Nous Research, “Skills System” in the Hermes Agent user guide. Primary source for progressive disclosure,
SKILL.mdformat, platform-specific skills, conditional activation (fallback_for_toolsets,requires_toolsets,fallback_for_tools,requires_tools), agent-managed skills viaskill_manage, the skill hub commands and source list (official,skills-sh,well-known,github,clawhub,claude-marketplace,lobehub), security scanning and trust levels, and external skill directories. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Persistent Memory” in the Hermes Agent user guide. Primary source for the
MEMORY.md/USER.mdcharacter limits, the frozen snapshot pattern, memory tool actions (add,replace,remove), what to save vs skip, the memory vs session search comparison, and the external memory providers. At tagv2026.9.24the companion Memory Providers page (line 9) reads “Hermes Agent ships with 7 external memory provider plugins” and adds that “more (such as Hindsight) are available from the plugin catalog”;plugins/memory/at the tag holdsbyterover,holographic,honcho,mem0,openviking,retaindb, andsupermemory. ↩↩↩↩↩↩↩↩ -
Nous Research, “Personality & SOUL.md” in the Hermes Agent user guide. Primary source for
SOUL.mdbehavior (lives inHERMES_HOME, never overwritten, slot #1 in system prompt, security-scanned before inclusion), SOUL.md vs AGENTS.md distinction, the built-in personality list (14 personalities fromhelpfultohype), custom personalities inconfig.yaml, the/personalityoverlay pattern, and the full prompt stack assembly order. ↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Research, “Use MCP with Hermes” and MCP Config Reference in the Hermes Agent guides and reference. Source for
mcp_servers:configuration format inconfig.yamlwithcommand,args,envfields. ↩ -
Hermes Agent v0.8.0 Release Notes. April 8, 2026. Background process auto-notifications, free MiMo v2 Pro on Nous Portal, live
/modelswitching across platforms, Google AI Studio native provider, Qwen OAuth, inactivity-based timeouts, approval buttons on Slack/Telegram, MCP OAuth 2.1 PKCE, centralized logging, plugin system expansion. ↩↩↩↩↩ -
Hermes Agent v0.9.0 Release Notes. April 13, 2026. Local web dashboard, Fast Mode (
/fast), iMessage via BlueBubbles, WeChat + WeCom, Termux/Android, background process monitoring (watch_patterns), xAI + Xiaomi MiMo native providers, pluggable context engine, unified proxy support, security hardening (path traversal, shell injection, SSRF, RCE fixes),hermes backup/hermes import,/debug+hermes debug share, 16 supported platforms. 487 commits, 269 merged PRs, 24 contributors. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.12.0 Release Notes. April 30, 2026. “The Curator release.” Autonomous background Curator that grades, prunes, and consolidates the skill library on a 7-day default cycle running on the gateway’s cron ticker. Self-improvement loop upgraded: rubric-based grading, active-update bias, proper runtime inheritance, scoped toolsets restricted to memory and skills. Four new inference providers: GMI Cloud, Azure AI Foundry, MiniMax OAuth, Tencent Tokenhub. LM Studio promoted to first-class. Remote model catalog manifests auto-update without releases. Two new messaging platforms: Microsoft Teams (19th, via pluggable gateway architecture) and Tencent Yuanbao (18th, native text + media). Native Spotify via PKCE OAuth with bundled skill; Google Meet plugin for calls and transcription; Piper local TTS provider. ComfyUI v5 + TouchDesigner-MCP bundled by default. New skills: Humanizer, claude-design, design-md, airtable. CLI:
hermes -zone-shot mode,hermes update --checkpreflight,/reload-skillsslash command, pluggable busy-indicator styles. TUI cold start cut ~57% via lazy initialization. Security: secret redaction disabled by default; hardline blocklist for unrecoverable commands. Stats since v0.11.0: 1,096 commits, 550 merged PRs, 213 community contributors. See also: v2026.4.30 release tag. ↩↩↩ -
Hermes Agent v0.13.0 Release Notes. May 7, 2026. “The Tenacity release.” Multi-agent Kanban board with heartbeat, reclaim, zombie detection, hallucination gate, per-task
max_retries, multi-project boards./goalslash command for cross-turn target locking (Ralph loop primitive) with configurable turn budget.video_analyzetool, Gemini-first with compatible multimodal extensibility. xAI Custom Voices TTS provider with voice cloning. 7-language i18n: zh-Hans, ja, de, es, fr, uk, tr (CLI + gateway messages; docs zh-Hans only). Google Chat as 20th messaging platform via pluggable adapter pattern with genericenv_enablement_fn/cron_deliver_env_varplugin hooks; IRC and Microsoft Teams migrated onto the same pattern.ProviderProfileABC +plugins/model-providers/for pluggable third-party providers. Session auto-resume across gateway restart,/update, and source-file reloads. Checkpoints v2 single-store rewrite with real pruning, disk guardrails, no orphan shadow repos. Eight P0 security closures: secret redaction default-on, Discord cross-guild DM bypass (CVSS 8.1, role allowlists guild-scoped), WhatsApp default-rejects-strangers + never-respond-in-self-chat, MCP OAuth credential-save TOCTOU, CLIauth.jsonTOCTOU in credential writers, browser cloud-metadata SSRF floor in hybrid routing, cron assembled-prompt scanning (including skill content) for prompt injection,hermes debug sharelog-content redaction at upload time. Additional notable items: post-write linting for Python/JSON/YAML/TOML, cronno_agentscript-only watchdog mode, platform allowlists across Slack/Telegram/Mattermost/Matrix/DingTalk, MCP enhancements (SSE transport, OAuth forwarding, image results as MEDIA tags). Stats since v0.12.0: 864 commits, 588 merged PRs, 829 files changed, 295 community contributors, 282 issues closed (13 P0, 36 P1). ↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.14.0 Release Notes. May 16, 2026. “The Foundation release.” Since v0.13.0: 808 commits, 633 merged PRs, 1,393 files changed, 165,061 insertions, 545 issues closed (12 P0, 50 P1), and 215 community contributors. Adds SuperGrok OAuth with grok-4.3 1M context,
hermes proxy,x_search, PyPI packaging, lazy dependencies, cross-session 1h Claude prompt cache, ~19s faster launch, 180x faster browser CDP calls, LINE and SimpleX Chat for 22 messaging platforms,/handoff, native clarify buttons, Discord history backfill, raw-pixelvision_analyze, per-turn file-mutation verifier footer, LSP semantic diagnostics, unifiedvideo_generate, cua-drivercomputer_use, OSC8 links, Zed ACP Registry support, OpenRouter Pareto Code router, NovitaAI, Codex app-server runtime,huggingface/skills, pluginctx.llm,tool_override, Brave/DDGS search, dangerous-command hardening,/subgoal, Qwen Cloud rename, native Windows beta, 16 total locales, and broad documentation/test updates. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.16.0 release notes, “The Surface Release,” tag
v2026.6.5, published 2026-06-06T00:55:58Z (release-tag date June 5, 2026); latest as of 2026-06-08. New native Hermes Desktop (Electron, macOS/Linux/Windows; remote-gateway connect over secure WebSocket with OAuth or user/pass; per-profile remote hosts; cross-profile@sessionlinks; Simplified Chinese UI via typed i18n,display.language). Web dashboard expanded to a full admin panel (MCP catalog toggles, credential management, webhook/hook creation, memory config, gateway controls, System page with check-before-update + Debug Share, Channels page; pluggable auth incl. self-hosted OIDC andhermes dashboard register). New commands:/undo [N], configurable default interface (cli/tui,--cli), TUI/model+ Sessions overlay,hermes portal,hermes prompt-size,hermes sessions optimize. New models:deepseek-v4-flash,MiniMax-M3(1M context),qwen3.7-plus,gemini-3.5-flash; xAI Grok OAuth; fuzzy picker; hourly catalog refresh. Skills: leaner default set,environments:relevance gate,NVIDIA/skillsdefault trusted tap, progressive tool disclosure, MCP false-OAuth-success fix. Security: CVE-2026-48710 (Starlette BadHost) pinned ≥1.0.1, SSRF checks off the event loop, Bedrock bearer token stripped from subprocess env,bws_cache.jsonread-guarded,docker restart/stop/killdangerous-pattern additions, invisible-unicode sanitization; 2 P0 + 62 P1 closed (16 security-tagged). Release-note marketing framing (PR/commit counts, “none of this existed a week ago”) excluded; only concrete feature/version facts tied to the tag are recorded. Current-session verification June 8, 2026. ↩↩↩↩↩↩↩↩ -
Hermes Agent v0.17.0 release notes, “The Reach Release,” tag
v2026.6.19, June 19, 2026; latest as of 2026-06-21. Messaging: iMessage via Photon Spectrum (device-code OAuth,hermes photon login, no Mac relay); official WhatsApp Business Cloud API adapter (replaces bridge process); SimpleX groups, native attachments, text batching, auto-accept; Raft bundled platform plugin. Models/providers:z-ai/glm-5.2(1M context),anthropic/claude-fable-5,laguna-m.1,nemotron-3-ultra,grok-composer-2.5-fast(xAI OAuth, 200k context); xAI default →grok-build-0.1; Anthropic adaptive models use the modern thinking contract (noreasoningfield). CLI/slash:/version,/billing,hermes photon login,hermes curator run --consolidate(opt-in),hermes modelGUI, profile cloning. Desktop: background subagent watch-windows (delegate_task(background=true)), Composer model selector, rebindable shortcuts, native OS notifications, per-thread drafts, VS Code Marketplace themes, Japanese + Traditional Chinese UI. Dashboard: full profile builder, global profile switcher, Skills Hub rehaul with security scan, Automation Blueprints, secure login (401 behind OAuth). Skills/tools:image_generateimage-to-image editing across providers,memoryoperationsatomic batch,simplify-codeparallel-review skill, booleanwrite_approvalreplaceswrite_mode. Architecture: background subagents (handle returned immediately, result re-enters as a turn), MCP elicitation handler for mid-tool-call confirmation, late-connecting MCP tools exposed between turns, pluggable CronScheduler + Chronos managed-cron, Managed scope (/etc/hermesadmin-pinned), Gateway-Gateway relay. Security: shell-escape denylist bypass closed, fail-closed on missing approval module and own-policy gateway adapters, cron job-script env sanitized, secrets redacted in debug dumps, host metadata withheld from public status, MCP stdio exfil-pattern screening, urllib3 + PyJWT CVE bumps. Release marketing framing (commit/PR counts) excluded. Current-session verification June 21, 2026. ↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.18.0 release notes (tag
v2026.7.1), July 1, 2026 — “The Judgment Release.” Priority backlog sweep (every P0/P1 closed, ~692 items in twelve days); Mixture-of-Agents selectable as a first-class model across all interfaces with each reference model’s full output rendered as its own labelled block and live answer streaming; completion contracts for/goal(agent verifies its own work by running project checks);/learncommand (turn anything into a reusable skill by describing it, with automatic CONTRIBUTING.md compliance);/journeyvisual memory/skill timeline with editing and a desktop memory graph; background subagent fan-out (multiple concurrent delegated tasks); Desktop Projects (project/repo/lane model); scale-to-zero gateway with drain coordination; Google Vertex AI support (Gemini via GCP service accounts, automatic OAuth2 token refresh);/prompt$EDITOR command. Current-session verification July 1, 2026 (PST) against the GitHub releases page; v0.18.0 is the latest release. ↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.20.5 release notes (tag
v2026.8.19, stated release date August 19, published August 21, 2026; fetched via the GitHub API August 24, 2026 - prerelease: false). Verbatim window description: “~746 commits across ~1,250 files (+111,500 / -20,701) - ~323 merged PRs including Bot Mode group-room threads, foldable conversation summaries, blob-face avatars, and PDF/file attachments with drag & drop; the keyless web tier (5-vendor free rotation with ring failover, web search on fresh installs with zero keys); a CLI polish wave (fuzzy /model picker, Ctrl+P command palette, richer /status); execution-discipline and runtime stall guards from the Composio eval findings;hermes updatereceipts and fleet--planverification;hermes worktree list/prune; the opencode-free zero-auth provider; multi-question clarify; desktop perf work (paint-first Bot Mode hydration, compositor spinners, React Compiler in both renderers); and cron jobs gaining persistent memory and per-job reasoning effort.” The release repeats: “Full curated release notes for this window will ship with v0.21.0.” ↩↩↩↩↩↩↩↩↩↩↩ -
Terminal backends at tag
v2026.8.31(re-verified for guide v1.17;agent/terminal_env_registry.pyis byte-identical to itsv2026.8.31version, thetools/terminal_tool.pydocstring facts are unchanged, andtools/environments/still carries the same seven backend modules, now alongside apath_utils.pyhelper):tools/terminal_tool.pymodule docstring, verbatim: “A terminal tool that executes commands in local, Docker, Modal, SSH, Singularity, Daytona, and Vercel Sandbox environments” and, in the environment selection list, thevercel_sandboxentry “Execute in Vercel Sandbox cloud sandboxes”; the backend moduletools/environments/vercel_sandbox.pyexists at the tag alongsidedaytona.py,docker.py,local.py,modal.py,singularity.py, andssh.py. The README at the tag: “Seven terminal backends – local, Docker, SSH, Singularity, Modal, Daytona, and Vercel Sandbox.” The Tools & Toolsets docs page (sourcewebsite/docs/user-guide/features/tools.mdat the tag) tablesvercel_sandboxas “Vercel Sandbox cloud microVM” for “Cloud execution with snapshot-backed filesystem persistence”, shows the config comment# or: docker, ssh, singularity, modal, daytona, vercel_sandbox, and states: “Authenticate with all three ofVERCEL_TOKEN,VERCEL_PROJECT_ID, andVERCEL_TEAM_ID. … Supported runtimes arenode24,node22, andpython3.13; Hermes defaults to/vercel/sandboxas the remote workspace root.” Atv2026.8.31thetools/environments/package also carries shared plumbing (base.pywith theBaseEnvironmentABC,file_sync.py,modal_utils.py,managed_modal.py); the package docstring counts managed Modal as a mode of Modal, not an eighth backend: “Modal additionally has direct and Nous-managed modes, selected via terminal.modal_mode.” ↩↩↩ -
Messaging platform count at tag
v2026.8.31(re-verified for guide v1.17; identical tov2026.8.31– 28 docs-table rows, the same 24Platformenum members, the same 22 adapter directories). The Messaging Gateway docs page (sourcewebsite/docs/user-guide/messaging/index.mdat the tag) has a “Platform Comparison” table with 28 rows: Telegram, Discord, Slack, Google Chat, WhatsApp, WhatsApp Cloud API, Signal, SMS, Email, Home Assistant, Mattermost, Matrix, DingTalk, Feishu/Lark, WeCom, WeCom Callback, Weixin, BlueBubbles, Photon (iMessage), QQ, Yuanbao, Microsoft Teams, LINE, ntfy, Raft, IRC, Buzz, SimpleX.gateway/config.pydefinesclass Platform(Enum)with 24 explicit members (local,telegram,discord,whatsapp,whatsapp_cloud,slack,signal,mattermost,matrix,homeassistant,email,sms,dingtalk,api_server,webhook,msgraph_webhook,feishu,wecom,wecom_callback,weixin,bluebubbles,qqbot,yuanbao,relay) and documents: “Plugin platforms use dynamic members created on-demand by_missing_()so thatPlatform("irc")works without modifying this enum.” The tag’splugins/platforms/tree holds 22 adapter directories:a2a buzz dingtalk discord email feishu google_chat homeassistant irc line matrix mattermost ntfy photon raft simplex slack sms teams telegram wecom whatsapp. Platform descriptions: the ntfy page, verbatim: “ntfy is a simple HTTP-based pub-sub notification service. It works with the free public server atntfy.shor any self-hosted instance … subscribe to a topic from the ntfy mobile app, send messages to the topic to talk to the agent, get the response back on your phone.” The Buzz page, verbatim: “The Buzz adapter connects Hermes to a Buzz community – Block’s open-source human+agent collaboration platform built on the Nostr protocol – and relays messages between Buzz channels (or DMs) and the agent. Outbound traffic shells out to thebuzzCLI binary … inbound uses a native Nostr WebSocket subscription.” Both pages: “Runhermes gateway setupand pick … for a guided walk-through.” ↩↩↩ -
Provider count at tag
v2026.8.31(re-verified for guide v1.17; v0.21.0 CHANGED these counts fromv2026.8.31’s 38 static entries / 37 plugin directories / 41 docs rows).hermes_cli/models.pydeclaresCANONICAL_PROVIDERS: list[ProviderEntry]with 39 static entries (nous,fireworks,openrouter,moa,novita,lmstudio,anthropic,openai-codex,openai-api,alibaba,xai-oauth,xiaomi,tencent-tokenhub,tencent-tokenplan,nvidia,copilot,copilot-acp,huggingface,gemini,vertex,deepseek,xai,zai,kimi-coding,kimi-coding-cn,stepfun,minimax,minimax-oauth,minimax-cn,ollama-cloud,arcee,gmi,kilocode,opencode-zen,opencode-go,bedrock,azure-foundry,ai-gateway,qwen-oauth), followed by the comment “Auto-extend CANONICAL_PROVIDERS with any provider registered in providers/ that is not already in the list above. Adding plugins/model-providers// is sufficient to expose a new provider in the model picker”; the loop skips only the oauth_device_code,oauth_external,external_process,aws_sdk,copilot, andvertexauth types (unchanged). The tag’splugins/model-providers/tree holds 39 provider directories (nebius-token-factoryandrouteradded in the v0.21.0 window); the nine without a static entry (actual,alibaba-coding-plan,commandcode,deepinfra,meta-ai,nebius-token-factory,opencode-free,router,upstage) all resolve to theapi_keyauth type – seven declareauth_type="api_key"explicitly (including both new plugins), andcommandcodeandopencode-freeinheritauth_type: str = "api_key"fromproviders/base.py. The AI Providers docs page (sourcewebsite/docs/integrations/providers.mdat the tag) tables 45 named providers plus a “Custom Endpoint” row (rows added sincev2026.8.31: Ramp Router, Nebius Token Factory, Tencent TokenPlan, Alibaba Cloud (Token Plan) – quoted in 36); every env-var name, provider slug, alias, and auth note in the new matrix rows is quoted from that table, including “OpenCode Free | Keyless – no API key or account needed (provider:opencode-free, aliases:free,opencode_free). Select viahermes modelor/model free; requests are sent anonymously”, “Google Vertex AI | … OAuth2 via service-account JSON or ADC, GCP billing”, “AWS Bedrock | … standard AWS credentials chain via boto3”, and “CommandCode | … Works with GOAT/Pro/Max/Provider plans (not the $1 Go plan – no API access)”. The Meta AI row comes from the same page’s first-class API-key block: “Meta Model API (Muse Spark family) …hermes chat --provider meta-ai --model muse-spark-1.2… Requires: MODEL_API_KEY”. Cerebras appears only in the page’s “Other Compatible Providers” table (https://api.cerebras.ai/v1, “Wafer-scale chip inference”), so it is listed under the custom-endpoint row rather than as a first-class provider. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
CLI surface at tag
v2026.8.31(re-verified for guide v1.17; the quoted global-option help strings and every_BUILTIN_SUBCOMMANDSfact are unchanged fromv2026.8.31– the frozenset is identical, 73 entries – whilehermes_cli/_parser.pychanged only insidehermes chat, documented in 40). Global options and their help strings are fromhermes_cli/_parser.py:--in(“Change into DIR before starting or resuming. Combined with ‘–resume latest’ or -c, the most recent session for DIR’s workspace is picked, and the session stays in DIR (skips the recorded-cwd restore).”),--ignore-user-config(“Ignore ~/.hermes/config.yaml and fall back to built-in defaults (credentials in .env are still loaded)”),--ignore-rules(“Skip auto-injection of AGENTS.md, SOUL.md, .cursorrules, memory, and preloaded skills”),--tui(“Launch the modern TUI instead of the classic REPL”),--cli(“Force the classic prompt_toolkit REPL (overrides display.interface=tui)”),--dev(“With –tui: run TypeScript sources via tsx (skip dist build)”), and--version/-V(“Show version and exit”).hermes_cli/_startup_fast.pyfast-paths onlyargv in (["--version"], ["-V"]), and the_BUILTIN_SUBCOMMANDSfrozenset inhermes_cli/main.pycontains noversionentry; the CLI Commands Reference documentshermes --version(“Show version information”) and nohermes version. Command descriptions are the parserhelp=strings fromhermes_cli/main.pyandhermes_cli/subcommands/*.pyat the tag (for exampleapprovals: “Approval-prompt tools (mine history into allowlist proposals)”;pause: “Emergency stop: pause cron/kanban dispatch and new gateway turns”;sync: “Skill Sync – sync your skills across devices and with your team”;verify: “Detect a project’s run recipe and smoke-test it”;login: “Deprecated. Usehermes authto manage credentials,hermes modelto select a provider, orhermes setupfor full setup.”), cross-checked against the docs’ Top-level commands table.hermes update --planis inhermes_cli/subcommands/update.py: “Show the update plan and exit without changing anything: install kind (git/docker/nix), every running Hermes service across all profiles with its supervisor and running code version, and how each will be restarted. Read-only; safe on a live fleet.”; nofleetcommand exists in_BUILTIN_SUBCOMMANDS.hermes worktreeis registered inhermes_cli/main.py(help: “Audit and reclaim accumulated git worktrees and merged branches”) withlist(aliasesls,audit; “Classify every tree: age, size, verdict, reason (default action)”),prune(“Remove safe trees and delete fully-merged local branches”),--repo,--dry-run(“Show the plan without changing anything”),--trees-only(“Only remove worktrees; leave local branches alone”), and--branches-only(“Only delete merged local branches; leave worktrees alone”). The/worktreeslash command is_handle_worktree_commandinhermes_cli/cli_commands_mixin.py(syntax block:/worktree,/worktree new [name],/worktree list,/worktree prune [--dry-run]); the CLI user guide “Worktree cleanup” section adds: “Inside a session,/worktree prune [--dry-run]does the same (and never touches the tree the session is running in).” ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Nous Portal commands at tag
v2026.8.31(re-verified for guide v1.17;hermes_cli/portal_cli.pyis byte-identical to itsv2026.8.31version). The README at the tag, section “Skip the API-key collection – Nous Portal”: “One command from a fresh install:hermes setup --portal… That logs you in via OAuth, sets Nous as your provider, and turns on the Tool Gateway. Check what’s wired up any time withhermes portal info.”hermes_cli/portal_cli.pyregistersportal(help: “Set up Nous Portal (login, model pick, Tool Gateway); see alsoportal info”) withlogin(“Log in to Nous Portal + set it up (default; one-shot onboarding)”),info(“Show Portal auth + Tool Gateway routing summary”),open(“Open the Portal subscription page in your default browser”),tools(“List Tool Gateway tools and which are routed via Nous”), and the code comment “statusretained as a hidden back-compat alias forinfo.” The Tool Gateway docs page showshermes setup --portal # Fresh install: Nous OAuth + set Nous as provider + turn on the Tool Gateway in one goandhermes portal info # Portal auth + Tool Gateway routing summary; the CLI Commands Reference at the same tag still documentshermes portal [status|open|tools]. ↩↩↩ -
GitHub compare API,
NousResearch/hermes-agent,v2026.8.19...main, fetched August 26, 2026:ahead_by: 1104,behind_by: 0. The releases API listsv2026.8.19(Hermes Agent v0.20.5, published 2026-08-21T12:16:39Z) as the newest tag, whose notes state: “Full curated release notes for this window will ship with v0.21.0, which will document everything from v0.20.0 onward.” Re-checked August 27, 2026 against the new tag:v2026.8.27...mainreturnsahead_by: 0,behind_by: 0,status: identical–mainnow sits exactly at the v0.20.6 tag. ↩↩ -
Hermes Agent v0.20.6 release notes (tag
v2026.8.27, stated release date August 27, published 2026-08-27T12:06:53Z; fetched via the GitHub API August 27, 2026 - prerelease: false). Verbatim framing: “Patch release. This tag rolls up the ~525 PRs merged since v0.20.5 into a stable tagged release for downstream consumers (Docker images, hosted deployments, fresh installs).” Verbatim window description: “Since v0.20.5 (v2026.8.19, tagged August 21), this window landed ~1,313 commits across ~1,557 files (+177,113 / -21,682) - ~525 merged PRs including consent-gated real-profile browsing (use your default Chromium profile for local browsing, with Windows close-with-approval flow); the desktop Browser getting its own OS window plus a managed SSH remote-update engine and fleet profile rail; a large remote MCP catalog expansion (50+ live-verified vendor-hosted servers incl. Cloudflare, Grafana Cloud, Better Stack, Railway); TTL result caching for web_search/web_extract; lean-tail compression as the default; multi-query tool_search with stemming; opt-in OS-keychain encryption for stored secrets (no more per-launch macOS Keychain prompts); updaters pausing gateways over the control socket instead of tree-killing them; image/package-managed installs refusing unsafe in-place updates (#91277 Phase 3); cron durable-incident acks and clearer code-skew failures; Slack link-unfurl controls; shared Docker container identities; pluggable terminal environment backends; and new models across the pickers (GLM-5.3-Flash, MiniMax M3 free, MiniMax H3 Max video).” The release repeats: “Full curated release notes for this window will ship with v0.21.0, which will document everything from v0.20.0 onward - highlights, feature areas, and complete contributor credits. Nothing in this window is skipped.” ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Compression at tag
v2026.8.31, re-verified atv2026.8.31for guide v1.17:tail_mode: str = "lean", the("legacy", "lean")guard, agent_init’s"lean"default, the config-version-17 migration, and the stale auxiliary-comment note below all hold at the new tag (context_compressor.pygained a pinned-summary-route retry path that changes none of the quoted defaults).agent/context_compressor.pydefaultstail_mode: str = "lean"inContextCompressor.__init__, with the comment “Lean tail mode (#compaction-v2): ‘lean’ = small clamped recency tail + verbatim-user-message summary section + recovery pointers; ‘legacy’ = 0.20window tail (shipping behavior)” and the guardself.tail_mode = tail_mode if tail_mode in ("legacy", "lean") else "lean".agent/agent_init.pyreadscompression.tail_modewith default"lean"; its comment quantifies both modes: “‘lean’ (default) keeps a clamped 2.5%/10K-25K verbatim tail with recovery-pointer machinery … ‘legacy’ restores the pre-#87326 0.20threshold verbatim tail, which on big-window/raised-threshold setups hoards 100-240K tokens per compaction.” The Configuration docs page (sourcewebsite/docs/user-guide/configuration.mdat the tag) shows thecompression:block withenabled,threshold: 0.50,threshold_tokens: null,target_ratio: 0.20,tail_mode: lean(annotated atv2026.8.31: “‘lean’ (default - clamped 2.5% tail, 10K-25K, with a detailed session log + anchor index + session_search recovery pointers in the summary, all from ONE auxiliary summarizer call; ~3x fewer retained tokens after compaction) or ‘legacy’ (0.20×threshold verbatim tail)”),protect_last_n: 20, andprotect_first_n: 3, puts the summarizer knobs underauxiliary.compression(model,provider,base_url), and notes: “Older configs withcompression.summary_model,compression.summary_provider, andcompression.summary_base_urlare automatically migrated toauxiliary.compression.*on first load (config version 17). No manual action needed.” The migration is_migrate_to_17inhermes_cli/config_migrations.py(“Version 16 -> 17: remove legacy compression.summary_ keys”). The context-compression developer page (sourcewebsite/docs/developer-guide/context-compression-and-caching.mdat the tag) tablestail_modewith defaultlean, states “Result on 500K-token real sessions: ~49K retained vs ~162K” and “Old tool results inside the lean tail are demoted to one-line stubs carrying a recovery pointer”, and carries the summary-model warning (atv2026.8.31): “The summary model must have a context window at least as large as the main agent model’s. … The compressor then drops the middle turns without a summary, silently losing conversation context.” (One stale spot upstream: the Configuration page’s “Full auxiliary config reference” still comments theauxiliary.compressionslot as “Context compression timeout (separate from compression. config)”; the migration code, the interactivehermes modelauxiliary picker, and the page’s own compression section are the newer, agreeing sources.) ↩↩↩↩↩↩↩ -
Pluggable terminal backends at tag
v2026.8.31(re-verified for guide v1.17;agent/terminal_env_registry.pyis byte-identical to itsv2026.8.31version).agent/terminal_env_registry.pymodule docstring, verbatim: “Central map of registered pluggable terminal backends. Populated by plugins at load time via PluginContext.register_terminal_environment_provider; consumed by tools.terminal_tool._create_environment” and “Built-in backend names are reserved - register_provider rejects a provider whose name collides with one, so a plugin can never shadow the in-tree docker/modal/… implementations”; there is deliberately “no active-provider resolution here: the active backend is whatever TERMINAL_ENV / terminal.backend names, exactly as for built-in backends.” ItsBUILTIN_BACKEND_NAMESfrozenset holds the seven backends plus the internal-mode aliasmanaged_modal: local, docker, singularity, modal, managed_modal, daytona, vercel_sandbox, ssh. The new docs page Terminal Environment Provider Plugins (sourcewebsite/docs/developer-guide/terminal-environment-plugin.mdat the tag), verbatim: “Hermes runs shell commands through a pluggable set of terminal backends. The built-in backends (local, Docker, Singularity, Modal, Daytona, Vercel Sandbox, SSH) live in the core repo under tools/environments/. Third-party sandbox vendors integrate as plugins instead - a standalone plugin repo installed under ~/.hermes/plugins/, registering a backend the user selects exactly like a built-in one via terminal.backend in config.yaml”; it tables the surfaces a registered provider drives (command dispatch, thehermes setupbackend picker, dashboard probe status,hermes status/hermes doctorchecks, system-prompt environment hints, container path/cwd handling, secret stripping, per-session sandbox isolation) and states the design goal: “Declaring these flags on the provider closes the classic ‘new backend missed classification site N’ bug class - the core consults the registry at each site instead of a hardcoded list of names.” ↩↩↩ -
Re-verification sweep at tag
v2026.8.31(guide v1.17). Locales: thelocales/tree holds 17 message catalogs –en.yamlplus 16 translations (af,ar,de,es,fr,ga,hu,it,ja,ko,pt,ru,tr,uk,zh-hant,zh) – identical tov2026.8.27. Personalities:BUILTIN_PERSONALITIESinhermes_cli/personality.pycarries the same 14 entries,helpfulthroughhype(helpful, concise, technical, creative, teacher, kawaii, catgirl, pirate, shakespeare, surfer, noir, uwu, philosopher, hype); the file is byte-identical to itsv2026.8.27version. Counts that CHANGED at the tag: 39 staticCANONICAL_PROVIDERSentries (was 38) and 39plugins/model-providers/directories (was 37) – see 26 and 36. Counts unchanged: 24Platformenum members (identical member list ingateway/config.py, which gained only aroom_link_urlgateway field), 22plugins/platforms/adapter directories, 28 rows in the docs’ Platform Comparison table;_BUILTIN_SUBCOMMANDSholds the identical 73 entries (peerandworktreepresent, still noversion, and noapproval-check– the release-notes name forhermes approvals test);hermes_cli/_startup_fast.pyandhermes_cli/portal_cli.pyare byte-identical to theirv2026.8.27versions. ↩↩↩↩ -
Auxiliary model routing at tag
v2026.8.31(re-verified for guide v1.17; the"auxiliary"defaults block at the new tag carries the identical slot roster, removal notes, and non-slot settings, and the docs quotes below hold verbatim). The Configuration docs page (sourcewebsite/docs/user-guide/configuration.mdat the tag), verbatim: “By default (auxiliary.*.provider: "auto"), Hermes routes every auxiliary task to your main chat model - the same provider/model you picked inhermes model. You don’t need to configure anything to get started, but be aware that on expensive reasoning models (Opus, MiniMax M2.7, etc.) auxiliary tasks add meaningful cost.”; its “Why ‘auto’ uses your main model” note: “Earlier builds split aggregator users (OpenRouter, Nous Portal) onto a cheap provider-side default. That was surprising - users who paid for an aggregator subscription would see a different model handling their auxiliary traffic.autonow uses the main model for everyone, and per-task overrides inconfig.yamlstill win.”; and on web extraction: “(Web extraction is not an auxiliary task:web_extractand browser snapshots truncate long content deterministically and store the full text forread_filepaging - no LLM involved.)” The authoritative slot roster is the"auxiliary"defaults block inhermes_cli/config_defaults.pyat the tag: vision, compression, skills_hub, approval, review, mcp, title_generation, memory_query_rewrite, tts_audio_tags, triage_specifier, kanban_decomposer, profile_describer, goal_judge, curator, monitor, background_review, moa_reference, and moa_aggregator (plus the non-slot settings transient_retries, free_only, openrouter_model, stream_only_base_urls); slots carryprovider,model,base_url,api_key,timeout,extra_body, and a per-taskreasoning_effort. Removal notes in the same file, verbatim: “web_extract no longer uses an auxiliary LLM - pages are truncate-and-stored with a read_file pointer (no summarization), and browser snapshots follow the same pattern. The oldauxiliary.web_extract.*block was removed here. Existing values in user config.yaml files are harmless leftovers and ignored.” and “session_search no longer uses an auxiliary LLM (PR #27590 - single-shape tool returns DB content directly)”; noflush_memorieskey exists in the block. Slot one-liners are the same file’s comments (“Triage specifier - flesh out a rough one-liner in the Kanban Triage column into a concrete spec, then promote it totodo. Invoked byhermes kanban specify”; “Kanban decomposer - decomposes a triage task into a graph of child tasks routed to specialist profiles by description. Invoked byhermes kanban decomposeand the kanban auto-decompose dispatcher”; “Profile describer - auto-generates a 1-2 sentence description of what a profile is good at. Invoked byhermes profile describe <name> --autoand the dashboard’s auto-generate button”; “Goal judge - evaluates whether a /goal run’s latest response satisfies the goal/contract”; “Curator - skill-usage review fork”; “Background review - the post-turn self-improvement fork that decides whether to save a memory / patch a skill”), the Configuration page’s full-reference comments (“Gemini 3.1 TTS hidden audio-tag insertion”; “Auto-generated session titles. Empty language follows the conversation”;auxiliary.title_generation.enabled: falsedisables auto-titles), and the Kanban docs page config table (“auxiliary.kanban_decomposer| Model that produces the task graph (called by Decompose)”; “auxiliary.profile_describer| Model that auto-generates profile descriptions (called byhermes profile describe --auto)”). Interactive path: runhermes modeland pick “Configure auxiliary models” for a per-task picker (vision, title_generation, tts_audio_tags, compression, approval, triage_specifier, kanban_decomposer, profile_describer, delegation); the Delegation entry persists to top-leveldelegation.*because subagents “are full child agents, not side-LLM calls.” ↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.21.0 release notes, “The Pantheon Release,” tag
v2026.8.31, stated release date August 31, published 2026-08-31T19:29:49Z. Verbatim stats line: “Since v0.20.0: ~5,800 commits · ~2,475 merged PRs · ~5,680 files changed · ~869,000 insertions · ~135,000 deletions · ~2,100 issues closed · 760+ contributors”. Verbatim framing: “The Pantheon Release. v0.20.0 made Hermes the herald - he spoke, and he carried word to other agents. In v0.21.0 the gods assemble.” and “This release rolls up everything from the v0.20.1-v0.20.6 infrastructure patch tags - those windows are fully documented here.” Feature-area quotes used above, verbatim: Bot Mode - “Bot Mode is now a bundled, default-on part of the desktop app: every agent profile gets a name, a deterministic avatar face (with randomize/lock controls), and a place in a shared roster” and “Before, ‘multi-agent’ meant plumbing; now it looks like a chat app full of coworkers” (#87886, #88243, #89386, #96726);hermes peer- “Replies land in each agent’s canonical Bot Chat, so conversations between agents are durable and inspectable, not fire-and-forget” (#88725, #88178, #91487); cron - “continuity=truecarries each run’s output into the next (so a monitor can dedupe against what it already reported), every job gets a durable notepad scratchpad, monitor-mode jobs skip the LLM entirely when nothing changed” (#91447, #80774, #81139, #81138); delegation - “list running children, steer one mid-flight with a course correction, or stop it early and keep the partial result. Add optional JSON-schema validation on child outputs, per-delegation cost surfaced in results, and raised defaults (250 iterations, 10 concurrent children)” (#85232, #81144, #81142, #86506, #86745); MCP - “hermes://deep links that install an MCP server with explicit confirmation” (#87525-#87581); CLI - “Ctrl+P opens a fuzzy command palette, the/modelpicker filters as you type,/statusshows reasoning mode, pending approvals, and context usage, and the status bar can display live cache-hit %, latency, and tokens/sec with per-field toggles” (#90730, #90717, #90745, #98250, #98282, #97666); browser - “Hermes now navigates, clicks, and reads it directly” (#90197, #89366); providers - “Meta Model API (Muse Spark) arrives as a built-in provider, alongside CommandCode, Tencent TokenPlan, Nebius Token Factory, Ramp Router, and Actual Computer” (#88565, #88308, #97917, #97916, #97915, and #79644/#26491 salvage); security - “Protected agent-instruction files (AGENTS.md, skills, memory stores) now always require write approval so a prompt-injected agent can’t quietly rewrite its own standing orders” (#81152) plus the redaction sweep (#80965), Windows approval coverage (#84428), TCC identity (#95091), and the Blender MCP removal (#83404). Reverted section, verbatim: “Model Council mode (/council) - landed then reverted; not in this release.”; “DCP context engine - landed then reverted; not in this release.”; “WS-only gateway server (#94245) - merged then reverted (#96118); FastAPI remains on the desktop boot path. The seq-stamped event replay (#94219) DID ship.”; “Electron rolled back to 40.10.2; TCC interpreter anchor removed (superseded by the signing-identity approach that shipped).” The bug-fix section closes: “…and roughly two thousand more closed issues’ worth - this window averaged ~85 merged PRs per day.” ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
New providers at tag
v2026.8.31.hermes_cli/models.pygainsProviderEntry("tencent-tokenplan", "Tencent TokenPlan", "Tencent TokenPlan (Hy4 preview via api.lkeap.cloud.tencent.com, Anthropic Messages)")(static entry #39), catalog modelshy4-preview,hy3,hy3-preview, and a display-only picker group"tencent": ("Tencent Hy", "Hy4 / Hy3 via TokenHub & TokenPlan", ["tencent-tokenhub", "tencent-tokenplan"]).plugins/model-providers/gainsnebius-token-factory/androuter/, both declaringauth_type="api_key", so the auto-extension absorbs them. The router plugin docstring, verbatim: “Provider profile for Ramp Router, Ramp’s LLM gateway: one OpenAI Responses-compatible endpoint athttps://api.router.com/v1that routes each request across upstream providers (OpenAI, Anthropic, xAI, Fireworks, …) and handles fallbacks and spend controls server-side” … “Responses API is the native wire.” … “Valid model IDs are whatever the key’sGET /v1/modelsreturns (BYOK accounts see extra entries), so this profile ships nofallback_models- the picker relies on the live fetch, per Router’s own guidance to never hardcode model names.” The AI Providers docs page at the tag tables 45 named providers plus the Custom Endpoint row; the four rows added sincev2026.8.27, quoted: “Ramp Router |RAMP_ROUTER_API_KEYin~/.hermes/.env(provider:router; aliases:ramp-router,ramp,router.com; Responses-native gateway, live account-scoped catalog)”; “Nebius Token Factory |NEBIUS_API_KEYin~/.hermes/.env(provider:nebius-token-factory; aliases:nebius,nebius-tf,tokenfactory)”; “Tencent TokenPlan |TOKENPLAN_API_KEYin~/.hermes/.env(provider:tencent-tokenplan, aliases:tokenplan,tencent-lkeap; Anthropic Messages endpoint)”; “Alibaba Cloud (Token Plan) |ALIBABA_TOKEN_PLAN_API_KEYin~/.hermes/.env(provider:alibaba-token-plan; mainland-China endpoint:alibaba-token-plan-cn) - Model Studio flat-token tier”.model_overridesis inhermes_cli/config_defaults.pyat the tag ("model_overrides": {}; explicitmodel_overrides.<provider>.<model_id>entries win over the catalog, and_defaultentries fill gaps only for models the catalog does not cover). ↩↩↩↩↩↩↩↩ -
Cron continuity at tag
v2026.8.31. The cron docs page (sourcewebsite/docs/user-guide/features/cron.mdat the tag), verbatim: “Setcontinuity=trueand the job injects its own most recent output into each run. Recurring jobs normally start every run with amnesia - a news scout re-reports the same stories, a monitor re-alerts on the same condition. With continuity on, the job wakes up seeing what it reported last time and can dedupe and continue where it left off”; “On later runs the previous output is prepended with continuity framing (‘avoid repeating what was already reported’) … Internally the flag is stored as the reservedselfentry incontext_from.”; “From the CLI:hermes cron create "every 6h" "Scan for news" --continuity, andhermes cron edit <job_id> --continuity/--no-continuityto toggle it on an existing job. The same toggle appears in the dashboard’s cron editor and the desktop Bot Mode routine dialog.” Per-job effort, verbatim: “A job can pin its own thinking level, independent of the model pin: one ofnone,minimal,low,medium,high,xhigh,max,ultra. When set, it overrides both the globalagent.reasoning_effortand per-modelagent.reasoning_overridesfor that job’s runs”; set viahermes cron create/edit --reasoning-effort high. Bot Chat delivery, verbatim: “bot-chatdelivers the output into a profile’s canonical “Bot Chat” session as a real message … the recipient here is the bot itself: it receives the output as an incoming message, acts on anything that needs action, and responds in its chat.” Notepads: thecron/notepad.pymodule docstring - “A tiny KV scratchpad each cron job can use to carry state across scheduled wake-ups (cursors, watermarks, watchlists)”, caps “MAX_VALUE_BYTES(16 KB)” and “MAX_JOB_TOTAL_BYTES(64 KB)” (“the notepad is prompt-injected each run, so unbounded growth would bloat every wake-up’s prompt”), and “Write path is the CLI (hermes cron notepad <job_id> set <key> <value>), which the running agent invokes via its terminal tool; no model tool is added.” Monitor mode: thecron/monitor.pydocstring - “Monitor-mode cron support - hash-suppressed change detection” attaching a “cheap monitor source (monitor_scriptormonitor_url)”; “unchanged -> the agent run is suppressed entirely (no LLM, no delivery); the tick is recorded as a silentno_changerun”; “Output is compared as EXACT BYTES - no timestamp stripping or whitespace normalization. Monitor scripts should emit stable output … or every tick will look like a change.”; “enabler: #80774.” ↩↩↩↩↩↩↩ -
Live subagent orchestration at tag
v2026.8.31. The delegation docs page (sourcewebsite/docs/user-guide/features/delegation.mdat the tag), “Steering a Running Subagent”: the control surface is{"action": "list"},{"action": "steer", "subagent_id": "sa-0-1a2b3c4d", "message": "focus on pricing instead"},{"action": "stop", "subagent_id": "sa-0-1a2b3c4d"}; verbatim: “listreturns the conversation’s live children:subagent_id, goal, status,running_seconds,accepting_steer, and the live transcript path”; “stopends a child early at its next iteration boundary; the partial result still re-enters the conversation as a normal completion message”; control actions are “scoped to the caller’s own spawn tree - a conversation can never see or control another session’s children - and never consume the per-turn subagent spawn cap, sostopkeeps working even after the cap is hit”; steer delivery is honest about the race (“Queued is not delivered, but it is never synthetic success”, withpending_steerdrained into results andmissed_steermarked when a child finished first). Defaults:hermes_cli/config_defaults.pysetsdelegation.max_iterations: 250(“per-subagent iteration cap (each subagent gets its own budget, independent of the parent’s max_iterations)”) anddelegation.max_concurrent_children: 10(“unified concurrency cap: max parallel children per batch AND max concurrent background (background=true) delegation units … (Replaces the deprecated max_async_children.)”). One stale spot upstream: the delegation docs page itself still says “3 tasks by default” and shows “max_iterations: 50 … (default: 50)” in its config reference; the shippedconfig_defaults.pyand the release notes (“raised defaults (250 iterations, 10 concurrent children)”) are the newer, agreeing sources. ↩↩ -
Security wave at tag
v2026.8.31. Protected instruction files:hermes_cli/config_defaults.pydefaultssecurity.protected_instruction_files: True(comment: “Writes to agent-instruction files (AGENTS.md/CLAUDE.md/SOUL.md/…)”) alongsideprotected_instruction_extra_patterns: [](fnmatch on basename);tools/file_tools.pyimplements the gate - “Protected agent-instruction files (always-ask approval gate)” over_PROTECTED_INSTRUCTION_BASENAMES = frozenset({"agents.md", "claude.md", "soul.md", ".cursorrules", ...})- and names the threat model, verbatim: “vector: an injected instruction that edits AGENTS.md / CLAUDE.md / SOUL.md” and “project-context instruction files are loaded from cwd trees - an AGENTS.md anywhere the agent might later run from is a live target.” TCC identity:hermes desktop --setup-tcc-identityis implemented inhermes_cli/main.py-_desktop_macos_setup_tcc_identity(identity: str = "Hermes Local Signing"), “One-shot setup forhermes desktop --setup-tcc-identity”, guarded “(–setup-tcc-identity is macOS-only; skipping)” and “(–setup-tcc-identity requires openssl, security, and codesign …)”. The redaction sweep (#80965, #80964, #81675, #81686, #88232), Windows approval coverage (#84428), Blender MCP removal (#83404), and Tier-1 plugin-install scanning (#80728) are per the release notes.35 ↩↩↩↩ -
CLI surface changes at tag
v2026.8.31.COMMAND_REGISTRYinhermes_cli/commands.pyregistersCommandDef("palette", "Open the fuzzy command palette (also Ctrl+P)", "Info", ...)andCommandDef("status", "Show session, model, token, and context info", "Session", ...)–/statusis a shared-registry session command, not messaging-only. There is noapproval-checkentry in_BUILTIN_SUBCOMMANDS(73 entries, identical tov2026.8.27); the dry-run the v0.21.0 release notes callhermes approval-check(#81137) ishermes approvals test– parser help inhermes_cli/subcommands/approvals.pyat the tag: “Dry-run the approval verdict for a command (never executes it)”, with--backend(“Terminal backend type to evaluate against (default: local …”) and--json.hermes_cli/_parser.pychanged only insidehermes chat:-q/--queryhelp is now “Query to run. On a real TTY the prompt seeds an interactive session (submitted literally as the first turn); combined with –oneshot or -Q, or on a non-TTY, it answers and exits.”; a new--oneshotflag documents “With -q/–query-file: answer the query and exit (legacy single-query behavior) instead of seeding an interactive session. Implied on non-TTY stdio and by -Q/–quiet.”; and--max-turnsdocuments “Maximum tool-calling iterations per conversation turn (default: 500, or agent.max_turns in config)”. ↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.21.1 release notes, tag
v2026.9.7, stated release date September 7, published 2026-09-07T22:17Z. The body is deliberately thin; verbatim: a “Patch release” that “rolls up current main since v0.21.0 for tagged deployments and downstream consumers”; window stats “5,139 non-merge commits across 4,364 changed files (+601,014 / -768,419)” and “632 merged PRs” at preparation time; and “Full curated release notes for this window will ship with v0.22.0.” Local-clone check:git rev-list --count --no-merges v2026.8.31..v2026.9.7= 5,140 andgit diff --shortstat= +601,018 / -768,423 – the body’s numbers are the pre-release-commit snapshot, offset by exactly the release commit; the largest previous adjacent-tag window isv2026.7.20..v2026.7.30at 2,790 non-merge commits (every adjacent pair fromv2026.3.12forward measured in the local clone). The Troubleshooting entry’s doctor check ishermes_cli/doctor_config.py:330-362at the tag (_drift_max_iterations_ghost): a staleHERMES_MAX_ITERATIONSin.envshadowsagent.max_turnswhen the startup bridge bails on an earlier config-parse error (issue #17534), andhermes doctor --fixremoves the.envline. Current-session verification September 8, 2026. ↩↩↩↩↩ -
COMPAT_MANIFEST.mdat tagv2026.9.7(repo root, 3,869 lines). Verbatim: “The September 2026 decomposition (PR #102117) split the large modules of Hermes Agent into focused files”; “Internal import paths are not a stable API”; “This layer is temporary and removed on 2026-09-14. It was added as a single commit and is removed by reverting that commit.”; the what-happens table (before 2026-09-14: yellow notice naming the plugin, the date, andhermes plugins compat; the plugin “loads; each old-path resolution emitsHermesPluginCompatWarningonce”; from 2026-09-14: “red notice: plugin DISABLED” and “not loaded;hermes plugins listshows the reason”; Desktop: “one-time modal”); the escape hatch, verbatim: “plugins.allow_deprecated_imports: trueinconfig.yamlkeeps affected plugins loading after the date, until the revert actually removes the paths.”; kind counts: moved-lazy 1148, import 592, restored-def 290, restored-helper 41, restored-import 17, module-stub 3, unrestorable 34; scope: only public top-level names, “Test monkeypatch seams are likewise not preserved.” Checker parser athermes_cli/subcommands/plugins.py:104-112:hermes plugins compat [path] [--json], description verbatim “Statically scans every enabled external plugin for imports of pre-decomposition module paths (see COMPAT_MANIFEST.md) and prints file:line, old path -> new path. Exits 1 when any plugin is affected.” Decomposition shape verified at the tag:agent/= 214 top-level modules + 7 subpackages (lsp,monitoring,pet,proxy_sources,secret_sources,transports,verify),hermes_cli/subcommands/= 61 modules,CANONICAL_PROVIDERSathermes_cli/models_catalog_static.py:311with 39 entries whose slugs are identical to thev2026.8.31list inhermes_cli/models.py(thexaistatic tuple included;plugins/model-providers/still 39 directories), the old top-levelprompt_builder.pygone (nowagent/prompt_builder.py), andrun_agent.pystillAIAgent’s home at the root. Aftermath, re-verified for guide v1.19 (September 15, 2026): the removal activated on schedule as a date gate, not a revert – at tagv2026.9.14,hermes_cli/plugin_compat.py:32setsCOMPAT_REMOVAL_DATE = _dt.date(2026, 9, 14),removal_in_effect()(lines 86-90) returns true from that date or when the manifest file is missing, andallow_deprecated_imports()(lines 261-268) honours only a literal boolean (the guard is... is True; comment in source: “Literal boolean only”, so a YAML string such as"false"or"no"can never open the post-removal bypass). The revert deleting the old paths has NOT landed:COMPAT_MANIFEST.md,compat_manifest.json, andhermes_cli/plugin_compat.pyare all present atv2026.9.14and onmainat commit5d59366010(2026-09-15 12:55 PT, fetched same day), so the escape hatch still resolves the old paths; it stops working the moment the revert lands. Re-verified for guide v1.20 (September 22, 2026): still no revert – all three files are present at tagv2026.9.21and onmainat commita53b42ddea(committed 2026-09-22, fetched same day). Atv2026.9.21the gate and guard are substantively unchanged:COMPAT_REMOVAL_DATEstill atplugin_compat.py:32,removal_in_effect()at :86, andallow_deprecated_imports()now at :296-303 with the identical literal-boolean guard (... is True; the comment now reads “Literal boolean only: YAML"false"/"no"must not open the post-removal bypass.”). The v0.21.4 window’s only changes to the module are performance and portability: a process-wide scan cache keyed on each plugin directory’s(relpath, mtime_ns, size)file signature (a multiplex gateway discovers plugins once per served profile, and re-parsing every plugin’s source cost ~0.4 s per profile on the boot path) plus POSIX-form hit paths on native Windows (#112576), and the manifest itself dropped two opencode-related rows (_OPENCODE_KEYLESS_EXTRA_SLUGS,is_opencode_zen_free_model) – nothing structural. ↩↩↩↩↩↩↩↩↩ -
docs/session-lifecycle.mdat tagv2026.9.7, section “6. Explicit Conversation Boundaries”, verbatim and in full: “Inactivity and wall-clock time never rotate a conversation./newand/resetcreate an explicit boundary; context compression continues to manage long histories. Legacy timer configuration is ignored. The existingSessionResetPolicydatatype is inert compatibility data, not a runtime policy. Explicit suspension still creates a boundary on the next inbound turn. Recovery respects explicit and historical finalized boundaries rather than reopening them. Resource-only eviction and WebSocket orphan reaping leave conversations resumable.” ↩↩↩↩ -
MCP authorization at tag
v2026.9.7. Parser:hermes_cli/subcommands/mcp.py:55-66–login(“Force re-authentication for an OAuth-based MCP server”) takes--flowwith choicesbrowser/device, help verbatim: “OAuth flow (overrides oauth.flow): browser PKCE or RFC 8628 device code”;reauth(“Re-authenticate one OAuth MCP server, or all of them (–all)”) takes an optional name plus--all. The device-code flow landed in-window: commitf5afe8bd40“feat: authorize MCP servers with device codes from the CLI”. Supporting hardening, by in-window commit subject:f914c9b070“fix(mcp): enforce profile ownership throughout OAuth sessions”;f94307a7f7“fix: ignore malformed MCP OAuth metadata caches”;e3ba651b6d“fix(desktop): relay MCP OAuth through client-local callbacks”. Toolset-filtered spawning:tools/mcp_tool_discovery.py:412-426– the filter exists so “hermes -z -t <toolsets>” can “skip cold-starting servers the caller doesn’t need”, and an empty filter skips MCP load entirely. ↩↩↩ -
Delegation reliability at tag
v2026.9.7, read from the delegate tool source. Completion units:tools/delegate_tool_dispatch.py:326-341(_units_of) – verbatim: “Off by default (delegation.independent_completions): the whole call is ONE unit and returns as one message. A per-task flurry of completions (one new turn each) fragmented orchestrators that had no plan for it.”; units are one per distinct taskgroup(first-appearance order) plus one per ungrouped task, each re-entering the conversation on its own; in-window commitc89f3b8800“fix(delegation): one completion per call by default; queued units no longer stalled”. Background-process handoff: commit3c0d90e8ef“feat(delegation): subagents hand background processes to the parent; leftovers are named, not trusted”; at the tag,tools/delegate_tool_child_run.py:744-763(account_background_processes) records handed-off processes on the result, lists still-running un-handed ones asorphaned_processesand exited-but-never-read ones asunread_completions(with an output tail) beforecleanupkills them, docstring: the parent “must hear that from the runtime” rather than trust a child’s claimed “watcher running”; the handoff verb isprocess_manage(action="handoff")(children only), flippingProcessSession.owner_task_idunder the registry lock viaprocess_registry.transfer_ownership(tools/AGENTS.md, Delegation section, at the tag). Fallback surface:delegation.fallback_providersinhermes_cli/config_defaults.py, comment verbatim: “For an unpinned child, null = inherit the parent chain; [] = disable fallback. A child pinned by provider, endpoint, or model gets no fallback unless this setting declares one explicitly.”; chain validation intools/delegate_tool_config.py:417-425(_resolve_child_fallback_chain): “Malformed entries are dropped by the canonical normalizer.” Crash durability:tools/async_delegation.py:222-246durably records each finished child of a still-running multi-child unit on the unit’s own row ("partial": True), so a crash before the unit completes keeps finished children. Cache tier:tools/delegate_tool.py:106-112(_apply_child_cache_ttl), verbatim: “A delegated child never uses the 1h cache tier.”; a child carrying_cache_ttl == "1h"is set to"5m". ↩↩ -
Providers and models at tag
v2026.9.7. Astra tiers:hermes_cli/models_catalog_static.py:22-25–openai/gpt-6-astra-fast“2x price, priority tier”,-flex“0.5x price, flex tier”, plus-pro-fast/-pro-flex;gpt-6-astraand-prosit inOPENROUTER_MODELSand are not in the_OPENROUTER_ONLYexclusion set, so the Nous Portal carries them too. Astra gating and 900K:hermes_cli/codex_models.py:96-101(“Astra is account-gated: only the live account-scoped catalog may advertise it”) andagent/model_metadata.py:1447-1462– Codex OAuth advertises 272K,gpt-6-astrais 900K-eligible with the comment “advertised 272K; 920,043 input OK, 1,000,043 rejected (live 2026-09-04)”, andCODEX_CONTEXT_VARIANT_SUFFIX = "-900k"is a “picker-only opt-in suffix; never sent on the wire” (OpenRouter-side Astra context is 1,050,000 permodel_metadata.py:334). New catalog entries in the same static file (lines 30-42):anthropic/claude-fable-5.1,google/gemini-3.8-flashandgemini-3.7-flash,qwen/qwen3.8-max-0902andqwen/qwen3.8-flash,meta/muse-spark-1.3and-contributor(1M context peragent/model_metadata.py:349).muse-image:plugins/image_gen/meta-ai/__init__.py(“Meta Model API (muse-image): OpenAI-compatible (https://api.meta.ai/v1)”, models includingmuse-image-1.0). Tavily:hermes_cli/config_defaults.py:2522-2525, verbatim “Tavily API key for AI-native web search and extract (optional – keyless works when Tavily is selected)”, toolsweb_searchandweb_extract; thewebblock notes Tavily “is opt-in keyless viahermes tools, not a ring member”. Managed llama.cpp runtime: thehermes_cli/local_runtime/package (“Managed llama.cpp runtime”), the config block atconfig_defaults.py:2327(“official binaries, one supervised” server; docs pointeruser-guide/local-models), and the desktop surfaceapps/desktop/src/api/local-models.ts. External-process providers:agent/auxiliary_client.py:4740-4799(_resolve_external_process_branch, “PROVIDER_REGISTRYexternal_processproviders, served via their registered profile”, keyed on the registered profile “so an out-of-tree ACP provider” resolves). Catalog cadence: config migration 39 -> 40 inhermes_cli/config_migrations.py:622-627(“model_catalog.ttl_hours -> ttl_minutes (default 20)”, user message “Model catalog now refreshes every 20 minutes (model_catalog.ttl_minutes)”), withhermes_cli/model_catalog.pyhonoring legacyttl_hours“only whenttl_minutesis still at its default”. ↩↩↩ -
Desktop wave at tag
v2026.9.7. Comment mode:website/docs/user-guide/desktop.mdat the tag, verbatim spine: “click Annotate in the preview browser bar, then click any element (or drag a box) on the live page and type a note; each saved comment stays as a numbered pin on the page”; “Saving a pin never sends a turn”; “Add N comments attaches a cropped screenshot per pin and a short prompt naming each comment to the composer”; “Each element comment carries its CSS selector, its markup, and the computed styles that matter for layout, so the agent can find the element in your source instead of guessing from the picture”; “Password and hidden field values, and any attribute that looks like a key or token, are redacted on the page before the markup leaves it”; “Larger batches arrive grouped by which part of the page each comment sits in, so twenty-odd comments become a handful of pieces of work rather than one task each”; “because the groups are separate DOM subtrees they usually touch separate files, which is what makes handing them to parallel workers safe”. In-window commits:10f2a20966“feat(desktop): add comment mode to the in-app browser”;e4bda3ff77“feat(desktop): browser comments carry the element’s selector, markup, and styles”; session controls8cb2bcc8c1“expose structured session controls” +bfddf556bf“hydrate structured session controls” +dffd8d62c2“add session automation controls”;6b1e12c7f4“drag to create sessions from New session, project + controls, and profile groups”;9186e3ebc5“session import view for foreign coding-agent transcripts”;a1c25d393a“built-in optional-skills catalog in Capabilities → Skills with one-click install”; Russian localea922dad9d8“feat(desktop): add Russian (ru) locale” +269e5bde33(registersruin locale tests/docs;apps/desktop/src/i18n/ru.tsis new in the window, while the CLI’slocales/remains 17 catalogs).display.resume_last_session:hermes_cli/config_defaults.py:777, defaultTrue, comment verbatim “Desktop reopens the last chat/page on cold start (also in Settings → Appearance).” Real-profile consent:apps/desktop/src/app/chat/right-rail/real-profile-consent-dialog.tsx(“First-open consent prompt for real-profile browsing”, shown when a Browser pane opens whilebrowser.use_real_profileis off; accepting writes the same config key the Capabilities toggle uses, “Not now” mutes for the app run, “Don’t show again” persists across launches). ↩↩ -
Hermes Agent v0.21.2 release notes, “The state.db Patch Release”, tag
v2026.9.11, stated release date September 11, published 2026-09-11T19:20Z. Verbatim framing: “v0.21.0 shipped a large rewrite of the session store’s connection handling, and for some installs it madestate.dbfragile: second writers cancelling each other’s locks, healthy databases reported as corrupt, one bad row killingsessions list.” Stats measured at commit04dd80a977: “947 non-merge commits”, “1,869 changed files”, “312 merged PRs”, “140 contributors” (local-clone check:git rev-list --count --no-merges v2026.9.7..v2026.9.11= 959 – the body’s number is a pre-release snapshot, the same pattern as v0.21.1’s). Campaign heading verbatim: “state.db reliability campaign (six PRs, 44 issues closed)” (PRs #108076, #108082, #108130, #108086, #108074, #108067); updating guidance verbatim: “runhermes doctorfirst; it now names structural vs index damage correctly and points athermes sessions recover --inspect-only(profile-pinned) when a rebuild isn’t enough.” Source verification at the tag: hosted-room state out of the root store atgateway/hosted_rooms.py:398-426–default_db_pathresolves profile gateways to “the shared ROOTshared-state.dbinstead of the masterstate.db”, and its docstring names the recurring multi-writer corruption vector observed across a 6-gateway fleet (2026-09-03) as the reason profile gateways must never open the master session store writable.hermes sessions recoverparser athermes_cli/subcommands/sessions.py:185-196with--inspect-onlyhelp verbatim “Only report canonical table readability; do not create an output database”; the subcommand predates the window (present inhermes_cli/main.pyatv2026.8.31) but had not been documented in this guide. Credential vault:agent/vault_backends/__init__.pydocstring (“Login backends for the browser credential vault”; handles namespaced by backend so browser tools route without schema changes; external managers locked until a per-session unlock; the master password “is never a tool argument, never argv, never persisted”),agent/vault_backends/base.py(aLoginBackend“lists login metadata (never secrets) and resolves ONE password at fill time”), andagent/vault_store.py:74-105(authenticator keys as base32 seeds orotpauth://totpURIs only, counter-based HOTP rejected, codes generated bytotp_now); backendslocal.py,onepassword.py,bitwarden.pyplusagent/secret_sources/{onepassword,bitwarden,command}.pyall present at the tag. Plugin catalog subparsers athermes_cli/subcommands/plugins.py:installhelp “Install a plugin from the curated catalog, a Git URL, or owner/repo”,search“Search the curated Hermes plugin catalog”,browse“List every curated plugin catalog entry”,pack“Declarative, shareable plugin sets (hermes-pack.yaml)” withinstall/export/show. Guest onboarding atapps/desktop/electron/guest-onboarding-flag.test.ts: test title verbatim ‘guestOnboardingEnabled: exactly “1” in env or –guest-onboarding on argv turns the free tier on’; the suite asserts'true','0', and empty are all off, and thatdesktopBackendSpawnEnv“stamps the launch decision last and never lets an inherited value leak”. Multi-profile hardening issue cluster #107609-#107630 as listed in the release. Current-session verification September 15, 2026. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.21.3 release notes, tag
v2026.9.14, stated release date September 14, published 2026-09-14T16:04Z. Purpose verbatim: the tag “exists so the remote-gateway sign-in fixes below reach Cloud agents, which auto-update to the newest release tag.” Stats measured at commit9b419a2d3c: “1,036 non-merge commits”, “2,642 changed files”, “338 merged PRs” (local-clone check:git rev-list --count --no-merges v2026.9.11..v2026.9.14= 1,037, the release commit again). Both documented items verified in the local clone and present only in this window: commit5dea46d13d“fix(dashboard-auth): one refresh single-flight for the cookie gate and the native route, off the event loop” (#110061, fixes #55712; pairs with Portal-side hermes-portal#1209, a sliding 30-day idle horizon with a 5-minute rotated-token grace) and commit939a2f64b4“fix: long-lived processes stop minting duplicate state.db writer handles” (#110934, fixes #100896 and #103339). Deferral verbatim: “Full curated release notes for this window ship with v0.22.0, which will document everything from v0.21.0 onward” and “Nothing in this window is skipped”; the release’s undocumented-on-purpose list includes server-to-client JSON-RPC requests plus a Pydantic wire-contract registry, reasoning-effort selection on every model picker, OpenRouter OAuth PKCE, HEIF/HEIC/AVIF decoding, the Honcho peer-model rework, FAL catalog additions (Wan 3.0, Kling 3.0 / Kling Image v3, MiniMax H3 Max Turbo, Gemini Omni Flash 1.1, Meta Muse), Slack pasted tables and the Agent Sessions API, multiplexed-profile isolation and gateway-liveness fixes, and the state.db WAL refusal on cross-VM filesystems – all held for the v0.22.0 scan per the release’s own framing. Current-session verification September 15, 2026. ↩↩↩ -
Hermes Agent v0.21.4 release notes, tag
v2026.9.21, stated release date September 21, published 2026-09-21T18:10:55Z. Purpose verbatim: “Patch release. This tag rolls up the ~1,800 PRs merged since v0.21.3 into a stable tagged release for downstream consumers (Docker images, Hermes Cloud, hosted deployments). Full curated notes for this window are deferred to v0.22.0.” Stats measured at commit4b8a8134009a: “5,071 non-merge commits” across “5,169 changed files” (+312,961 / -62,855), “1,812 merged PRs” and “2,116 closed issues” – local-clone checks reproduce all five exactly (git rev-list --count --no-merges v2026.9.14..4b8a8134= 5,071;git diff --shortstat= 5,169 files, +312,961 / -62,855; the tag commitd337b736aa“chore(release): v0.21.4 (v2026.9.21)” is4b8a8134’s child, giving 5,072 non-merge and 5,173 total with merges, and the GitHub compare API forv2026.9.14...v2026.9.21reportstotal_commits: 5173, ahead_by: 5173, behind_by: 0). Window ranking (re-checked for guide v1.21, September 23): by non-merge commits across every adjacentv2026.*tag pair,v2026.8.31..v2026.9.7(5,140 with its release commit) leads andv2026.9.14..v2026.9.21(5,072) is second, ahead ofv2026.7.20..v2026.7.30(2,790); by release-stated merged PRs, 1,812 is the largest of any single-window rollup note (v0.21.1 “632 merged PRs”, v0.19.1 “~1,000+ PRs”; v0.21.0’s “~2,475” spans the six v0.20.x tags since v0.20.0, so it is not one tag-to-tag window). Deferral verbatim: “Full curated release notes for this window ship with v0.22.0, which will document everything from v0.21.0 onward” and “Nothing in this window is skipped.” The undocumented-on-purpose list names the gateway singleton lock + rendezvous record with Desktop attaching to the running host backend, the backend-owned connector operation with its Desktop/TUI/CLI setup card,--format stream-json,skills.auto_load, the Desktop chat/UI font picker + one-click local engine updates + Plugins-hub uninstall, thedeclineunauthorized-DM behavior,mcp.discovery_concurrency,session_searchafter/before bounds + OR-relaxed recall retry,hermes sessions set-journal-mode, LTX 2.5 and Kling O3 in the video catalogs, catalog website pages per plugin/author with pinned-commit READMEs and added/updated sorting, “a dozen new community plugins in the catalog (tailscale, ssh, shodan, terminal, rss, resetwatch, done-bell, kiwi, cognee, Octen)” (release shorthand; the catalog entries at the tag arehermes-tailscale,hermes-ssh,shodan,hermes-terminal,hermes-rss,hermes-resetwatch,done-bell,kiwi,cognee,web-octen), and “a large run of profile/multiplex isolation, cron, kanban, Desktop and state.db fixes”. Updating:hermes update(git installs) or the installer one-liner; “Docker / Hermes Cloud: images build from this tag (nousresearch/hermes-agent:v2026.9.21)”. Current-session verification September 22, 2026. ↩↩↩↩↩↩ -
v0.21.4 disclosed items verified in source at tag
v2026.9.21(guide v1.20, September 22, 2026; operator-rule additions in guide v1.21, September 23). Host singleton:gateway/host_rendezvous.pydocstring – “Host-wide singleton rendezvous: one lock + one record per ROLE per OS user”; “exactly ONEhermes serveand ONEhermes gateway runper host, each multiplexing every profile”; a host flock/msvcrtlock “held for the lifetime of the winning process” plus a rendezvous record so a second invocation can “prove it is the same live process, and ATTACH instead of binding a second port”; “Staleness is proved, never assumed” via(pid, createTime)(“an attaching client must never dial a recycled PID’s port”); lock root$HERMES_GATEWAY_LOCK_DIRelse$XDG_STATE_HOME/hermes/gateway-locks, scoped to the OS user (gateway/status.py:308-324: a relativeXDG_STATE_HOMEis ignored and the fallback is~/.local/state). Operator path:gateway/run.py:5466-5499_host_attach_or_noneprints the attach message and exits 0 onATTACH, refuses onREFUSE, sends--replaceto “the HOST process, whichever home launched it” onREPLACE_HOST, and skips the question under--force(“the operator’s escape hatch when the owner is wedged or lying”);_claim_host_gateway_role(gateway/run.py:5331-5346) has the lock-race loser exitEX_TEMPFAIL(75) because “Every supervisor we generate retries 75, and on the retry the owner’s record exists”. Five outcomes ingateway/host_attach.py:ATTACH,RESCAN->ATTACH(control-socketrescan-profiles),REPLACE_HOST,REFUSE(“Never start a second one silently”),START(standalone per-profile gateways coexist “until that migration is forced (#109417)”). Desktop half:apps/desktop/electron/host-backend-attach.ts(“Attach to the host’s running Hermes backend (multiplex-only, Desktop half)”; the IO ladder reads the machine-root spawn ledger, validates a candidate at “HTTP readiness -> served session token -> WebSocket auth”, and holds “a host-level gate so two apps starting at once produce one backend instead of two”;HOST_SPAWN_GATE_STALE_MS = 60_000). Connector operation:tools/connectors/operation.py:1verbatim “One backend-owned connection operation permanage_connectionscall. Pure data, no I/O.”;OPERATION_DEADLINE_SECONDS = 300.0with the comment “Not a config key: a user-tunable wait with clamp rails was a foot-gun (PR1 shipped one, unmerged)”; the tool registered asmanage_connections(tools/connectors/tool.py:16,48,113); per-targetrequired_env“({name, prompt, required}); the card draws a field per entry and holds its verb until every required one has text”; the three-frontend card named intests/hermes_cli/test_mcp_catalog_env_boundary.py:330(“The connector-card backend (Desktop/TUI/CLI setup card) makes the same secrets-only split”). stream-json:hermes_cli/_parser.py:247-249on the chat parser –--formatwith choicestext/stream-json, defaulttext, help verbatim “‘stream-json’ emits newline-delimited JSON events (JSONL), implies –quiet, and cannot be combined with –tui”;hermes_cli/stream_json.pymodule docstring “one JSON object per stdout line …system/init->textdeltas /tool_use/tool_result-> one terminalresultenvelope (exit code, final text, token stats). Diagnostics andsession_idstay on stderr”,_TOOL_OUTPUT_CAP = 5000, exit 2 on the forbidden combinations,stream_json_requestedathermes_cli/stream_json.py:23accepting eitherqueryorquery_filebefore exiting 2 (thecli.py:1703-1705Fire entry point checks the already-resolved query);--query-file PATHathermes_cli/_parser.py:218-222in a mutually exclusive group with-q(help verbatim “Read the single query from a file instead of the command line (‘-’ reads stdin). Safe for arbitrary text: nothing is shell-interpreted”), read by_read_query_file(hermes_cli/main.py:1735-1762), present since tagv2026.8.19; contract tests intests/hermes_cli/test_stream_json.py(including “must never reach stdout under stream-json”). skills.auto_load:hermes_cli/config_defaults.py:1435, default[], comment verbatim “Skill names pinned as fully loaded in every new session (CLI, TUI, gateway, cron, API). Resolved once when the agent’s prompt is first built; missing/disabled names warn and skip; HERMES_IGNORE_RULES suppresses the list like the other auto-injected context.” decline:gateway/config.py:137-139– comment verbatim “‘pair’ DMs a pairing code, ‘ignore’ drops silently, ‘decline’ sends one polite refusal then goes silent toward that sender for gateway.pairing.DECLINE_DEDUPE_SECONDS (#88028)”,UNAUTHORIZED_DM_BEHAVIORS = {"pair", "ignore", "decline"}, field default"pair"(:626),unauthorized_dm_decline_messageempty ->DEFAULT_UNAUTHORIZED_DM_DECLINE_MESSAGE(the reply quoted in the Pairing section);DECLINE_DEDUPE_SECONDS = 24 * 3600and alias-aware decline stamps ingateway/pairing.py:37,565-576; per-platform resolution and the Email default inget_unauthorized_dm_behavior(gateway/config.py:803-809, “Email is inbox-shaped so it defaults to"ignore"unless its ownunauthorized_dm_behavioropts in (a global default does not)”). The key itself, withpair/ignoreand the Email rule, predates the window (present atv2026.9.14,gateway/config.py:564,736-742; introduced by #1919 in March 2026); onlydeclineis new. Effective default:gateway/authz_mixin.py:699-740_get_unauthorized_dm_behavior, order per its docstring “explicit per-platform config; Email -> “ignore”; explicit non-default global; adapter dm_policy (pairing -> “pair”, allowlist/disabled -> “ignore”); any configured allowlist -> “ignore” …; else “pair”” (#9337), with the global consulted only when!= "pair". YAML shape: per-platform keys underplatforms.<name>are promoted intoPlatformConfig.extra(gateway/config.py:456-460, #10206), the top-level key or nestedgateway.unauthorized_dm_behavioris bridged bygateway/config_loader.py:39-46,102, andhermes gateway setupwritesplatforms.<name>.unauthorized_dm_behaviorviawrite_platform_config_field(hermes_cli/config.py:2061-2069;hermes_cli/gateway_setup_wizard.py:182-184,219-249, choice “Politely decline unknown senders (one-time message, then silence)”). The decline message is global only (gateway/run_inbound.py:139). discovery_concurrency:hermes_cli/config_defaults.py:526default 4;tools/mcp_tool_discovery.py:27-40(“mcp.discovery_concurrencyoverrides it, 0 = unlimited (#117373)”; a non-integer or negative value logs “is not a non-negative integer; using %d” and uses the default); cap-not-a-gate semantics pinned bytests/tools/test_mcp_tool.py:2813(“caps simultaneous connects (still concurrent, every server” connects). session_search bounds + retry:tools/session_search_tool.py:578,621-635,708-725–after“Inclusive lower bound on session start time. ISO date/datetime (e.g. 2026-06-01) or relative duration (7d, 24h, 2w = within the last N)”,before“Exclusive upper bound on session start time. ISO date/datetime (a date-only value is midnight UTC that day) or relative duration (7d = older than a week)”, both “Discovery shape only” with “sort is a ranking bias, not a bound”, and the new parameters “appended afterdetail” for schema stability; the OR-relaxed retry inhermes_state_search.py:1151-1163, comment verbatim: “the implicit AND between terms means a paraphrased multi-word query misses a stored sentence that lacks even ONE word … Once the exact query and the substring fallbacks all miss, retry the unicode61 index matching ANY term. … Gated on a zero-result miss so hits keep exact-match semantics; explicit OR/NOT, single-term and CJK-routed queries are left alone”; dedicated suitetests/hermes_state/test_search_or_relaxed_fallback.py. set-journal-mode: parser athermes_cli/subcommands/sessions.py:177, help verbatim “Convert state.db between journal_mode=WAL and DELETE offline (every holder stopped)”;hermes_cli/sessions_cmd_journal_mode.pydocstring – the offline self-service path for #100896 (before it, “the only escape hatch was an undocumented hand-runPRAGMA journal_mode=DELETE”), “refuse while ANY foreign process holds the file or a sidecar (foreign_state_db_holders), flip without waiting out openers (_set_journal_mode_no_wait), then verify the header bytes 18/19 SQLite writes for the mode”; dispatch marked “offline: must not open the store it converts” (hermes_cli/sessions_cmd.py:982);hermes doctorrecommends it with--dbfor non-default stores (hermes_cli/doctor_platform.py:148,163). Parser description verbatim “Run this with the gateway, dashboard and every CLI stopped: it refuses while any process holds the file, switches the mode, and verifies the file header”;--forcehelp “Windows only: proceed without the holder scan (which does not exist there) after stopping every Hermes process yourself” (hermes_cli/subcommands/sessions.py:176-189), enforced in_refusal(sessions_cmd_journal_mode.py:40-44: “cannot prove the database is quiet on Windows – no holder scan”). Desktop wave: font fieldapps/desktop/src/app/settings/chat-font-setting.tsx(CONFIG_PATH = 'desktop.font_family', 550 ms autosave) withthemes/chat-font.ts(suggestions OpenDyslexic, Atkinson Hyperlegible, Lexend, Inter, IBM Plex Sans, Source Sans 3, Noto Sans, Segoe UI, SF Pro Text; empty means the theme’s face) applied throughthemes/context.tsx:279, which overrides the theme token--dt-font-sans; engine updates inapps/desktop/src/app/settings/local-models-settings.tsx+ test (an “Update engine” button when the managed local runtime reportsupdate_available, drivinginstallLocalRuntimeas a progress-reportingruntime-installjob; the test “keeps a failed explicit update visible with a direct retry and no staged models”); Plugins-hub uninstall inapps/desktop/src/app/capabilities/plugins/plugins-tab.tsx+ test (“uninstalls through plugins.manage remove only after the confirm dialog is accepted”; standalone desktop plugins viauninstallDiskPluginon the Electron loader). Video catalogs:plugins/video_gen/fal/__init__.py:43-44,99-101–ltx-2.5(“LTX 2.5”, “Lightricks open-source audio-video model. Native audio, up to 20s / 4K (i2v), camera-motion presets.”, cheap tier,lightricks/ltx-2.5/text-to-video/fastand/image-to-video/fast, aspect ratios 16:9/9:16, integer durations;:128“fal rejects LTX 2.5 >10s at 1440p/2160p”) andkling-o3(“Kling O3 (Standard)”, “Kuaishou frontier. Multi-shot native storytelling, optional audio, 3-15s.”, premium tier, string duration 3-15, i2v derives aspect ratio from the image,generate_audioa real toggle); payload tests attests/plugins/video_gen/test_fal_plugin.py:641,681; roster row inwebsite/docs/reference/toolsets-reference.md:72. Catalog buildout:plugin-catalog/atv2026.9.14= 9 entries +removed.yaml, atv2026.9.21= 228 entries +removed.yaml; entry shape perplugin-catalog/hermes-tailscale.yaml(name, repo, 40-hexsha, description, maintainer,tier: community, category, docs_url, capabilities); admission and publishing model inhermes_cli/plugin_catalog.py:4-10,33(“pinned to an exact 40-character commit SHA. Presence in the directory IS” admission;website/scripts/extract-plugins.pypublishes/docs/api/plugin-catalog.json;LIVE_CATALOG_URLfetched and cached under~/.hermes/cache/plugin-catalog.json); website pages inwebsite/plugins/plugin-catalog-pages/index.js(“/docs/plugins/one page per entry”, “/docs/plugins/by/ one page per maintainer”, “a merged catalog PR is the only way a page appears, changes or disappears”, and the site “degrades, never fails”) and readme.js(every entry’s README rendered “from the PINNED commit” via a raw URL at the sha, “never a branch tip”, so the page shows exactly the README the catalog reviewer read; build-time allowlisted rendering that drops raw HTML, 512 KB cap, hosts limited to raw.githubusercontent.com and gitlab.com; opt out withreadme: false); author slugs and committer-date added/updated stamps inwebsite/scripts/extract-plugins.py:80,142; all ten release-named community plugins present as entries at the tag under the slugs listed in 50. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.21.5 release notes, tag
v2026.9.24, stated release date September 24, published 2026-09-24T10:09:38Z. Purpose verbatim: “Patch release. This tag rolls up the ~460 PRs merged since v0.21.4 into a stable tagged release for downstream consumers (Docker images, Hermes Cloud, hosted deployments). Full curated notes for this window are deferred to v0.22.0.” Stats measured at the tag commitf97608f178(“chore: release v0.21.5 (2026.9.24)”): “1,610 non-merge commits” across “4,828 changed files” (+164,132 / -149,440), “460 merged PRs” and “475 closed issues”. Local-clone checks reproduce the commit and diff figures exactly (git rev-list --count --no-merges v2026.9.21..v2026.9.24= 1,610; 1,638 with merges;git diff --shortstat= 4,828 files, +164,132 / -149,440); the PR and issue counts were not independently checked. The body’s undocumented-on-purpose list names GPT-6 “Sol/Terra/Luna” and Claude Opus 5.5 in the Nous and OpenRouter catalogs, per-profile stop/start/restart withgateway.standalone, and a Desktop plugin-SDK wave; it does not mention Hindsight. Update path verbatim: “hermes update(git installs), or re-run the installer one-liner” and “images build from this tag (nousresearch/hermes-agent:v2026.9.24)”. ↩↩↩↩ -
v0.21.5 items verified in source at tag
v2026.9.24(guide v1.22, September 24, 2026). Hindsight: commit4cbf862abe“chore(memory): remove the bundled hindsight provider (moved to the plugin catalog)” (September 22) is an ancestor of the tag;plugins/memory/at the tag holds seven provider directories (byterover,holographic,honcho,mem0,openviking,retaindb,supermemory) against eight atv2026.9.21; commit73c598e319“build: drop the hermes-agent[hindsight] extra” removed it frompyproject.toml.plugin-catalog/hindsight.yaml:repo: https://github.com/vectorize-io/hindsight,maintainer: vectorize-io,tier: community,requires_hermes: ">=0.21.4". Migration:hermes_cli/memory_provider_migration.pydocstring (lines 1-14) names two hooks, “hermes update” and “agent init”; line 75 prints “Memory provider ‘{name}’ moved out of core – installed its plugin from the catalog”;recover_at_startup()(line 110) “honourssecurity.allow_lazy_installs”; it is called fromhermes_cli/update_cmd_deps.py:535-536andagent/agent_init.py:1315-1316. On-disk effects and the verify commands:memory-providers.md:482-490. Multiplex:hermes_cli/gateway_multiplex_mode.pyline 10, “An explicitfalseis RETIRED”, and lines 43-44, “gateway.multiplex_profiles: false is retired and was rewritten to true” (commitb936546561, September 23).gateway.standalone:hermes_cli/profiles.py:979-982, “The DEFAULT profile is never standalone – it IS the host – and warns once per process if the key is set there” (commit0238c9d740). Parking: commit4c342c05de“stop, start and restart one profile under the host multiplexer”;multi-profile-gateways.md:120-138, 233-244(“a temporary compatibility shim”; “gateway.standalonewins”). Models:hermes_cli/models_catalog_static.pyOPENROUTER_MODELSlines 32 and 36-37 (anthropic/claude-opus-5.5,openai/gpt-6-sol,-sol-pro,gpt-6-luna,-luna-pro); thenouslist at line 162 derives from it. Nogpt-6-terramodel ID appears in the tag’s picker catalogs, so only Sol and Luna are listed here. Providers:CANONICAL_PROVIDERSat line 314 parses to 39 entries (AST count, unchanged);plugins/model-providers/holds 38 directories at bothv2026.9.21andv2026.9.24, against 39 atv2026.9.14. Commit998f614c7f“feat(providers): remove the keyless opencode-free tier” (September 18) deleted the plugin; its message reads “OpenCode’s free tier now returns HTTP 403 for anonymous traffic outside the OpenCode client”.hermes_cli/auth.py:1255-1259keeps an error foropencode-free,free, andopencode_freethat points toopencode-zenandopencode-go. Compat:COMPAT_MANIFEST.md,compat_manifest.json, andhermes_cli/plugin_compat.pyare present at the tag and onmainataa8a33d22d(September 24).git diff v2026.9.21 v2026.9.24 -- hermes_cli/plugin_compat.pyis empty (COMPAT_REMOVAL_DATEat line 32; the “Literal boolean only” check at lines 296-303). The manifest pair lost only the deletedplugins.memory.hindsightentries (-7 / -12 lines). ↩↩↩↩↩↩↩↩ -
Hermes Agent v0.20.3 release notes (tag
v2026.8.16.2, stated release date August 16, published August 17, 2026) and v0.20.4 release notes (tagv2026.8.18, August 18, 2026); both fetched via the GitHub API August 20, 2026 (prerelease: false). v0.20.3 verbatim: “the MCP 2.x SDK migration and 2026-07-28 stateless protocol support, the bundled Bot Mode (hermes-bots) plugin with the core teammate protocol, the CommandCode provider plugin, subprocess Python runtime ownership hardening (PYTHONHOME/PYTHONPATH isolation), Cua Driver 0.20 runtime contracts for computer use.” v0.20.4 verbatim: “the desktop glass/translucency surface work (matte glass, frost picker, macOS pre-select), the tabbed SESSIONS|BOTS sidebar with per-bot hide/unhide, … NVIDIA SkillEvaluator Tier 1 advisory scanning on skill installs (license + security checks).” Both releases: “Full curated release notes for this window will ship with v0.21.0.” ↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.20.0 release notes, “The Herald Release,” tag
v2026.8.3, August 3, 2026, with stabilization tags v2026.8.13 and v2026.8.16. Verbatim from the release: “Node 26 required across installers/heal/upgrade”; “brew + pip/PyPI wheel channels retired (shell installer / Docker / Nix are the supported channels)”; “default iteration limit 90 → 500”; “claude-marketplace source removed”. Node floor independently confirmed in the installer source at scripts/install.sh —NODE_VERSION="26"and the guard “Node.js $(node –version) is too old (Hermes requires Node >=26)” — which also documents the canonical one-linercurl -fsSL https://hermes-agent.nousresearch.com/install.sh | bashin its header comment. Note the conflict: the docs installation page still states Node v22 and is stale against both. Platform tiers from platform support; skills sources and default taps from skills; the 28-platform count derived by enumerating the comparison table at messaging, which publishes no official total. All fetched and verified 2026-08-16. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.19.0 release notes, “The Quicksilver Release,” tag
v2026.7.20, July 20, 2026; latest as of 2026-07-21. Stats since v0.18.0: ~2,245 commits, ~1,065 merged PRs, ~3,300 issues closed, 450+ community contributors. Performance spine: ~80% first-turn TTFT cut, cold submit→dispatch ~4.3s → ~0.9s across CLI/gateway/TUI/desktop/cron (PR #59332); reasoning streamed live by default withdisplay.show_reasoningON and per-token response painting (PR #59389); desktop ~20-PR speed wave incl. 14× faster streaming markdown; TUI incremental markdown. pip/Homebrew installs deprecated to warn-only “unsupported legacy” with PyPI/Homebrew publishing removal planned (PR #57225). PluggableSecretSourceinterface with Bitwarden + 1Password providers,op://refs, multi-vault, deterministic precedence, per-variable provenance (PR #59498). Smart approvals default (independent LLM reviewer per flagged command), user-defined deny rules that hold under YOLO,/deny <reason>(PRs #62661, #59164, #54518); pluginpre_tool_callapprove escalation re-landed (PR #60504). Terminal billing/subscription+/topup+ desktop billing tab (PR #51639). Live subagent transcript files + durable background delegation (PRs #67479, #63494); delivery-obligation ledger instate.db(PR #67181);max_async_childrendeprecated for unified delegation concurrency caps (PR #56955). Gateway profile-based routing +GATEWAY_MULTIPLEX_PROFILES+ routing index moved tostate.db,sessions.jsonoptional legacy mirror (PRs #64835, #65700, #60589, #59203). Providers/models: Fireworks AI first-class at picker slot #2 (PR #62593), DeepInfra, Upstage Solar, GPT-5.6 Sol/Terra/Luna + Pro end-to-end (PR #61616), grok-4.5 GA, kimi-k3 (kimi-k2.x retired), Claude Sonnet 5 fully wired, per-providerenabled: false+excluded_providers(PR #67971); reasoning effortmax/ultratiers with per-model/per-MoA-slot overrides and session-scoped/reasoning(PRs #62650, #64458). CLI/MCP:hermes sessions exportMarkdown/Quarto/HTML/prompt-only/HF-trace with--redact(PR #60186),/model --once(PR #67113), stacked slash-skill invocations (PR #57987),--safe-mode,hermes config get/unset(PR #65540),hermes servetrue headless (PR #55923), MCPmcp__server__toolnaming (PR #52750). Release marketing framing excluded; reverted-in-window items (iron-proxy egress firewall, dynamic-workflow skill, memory provider-actions) deliberately not recorded as shipped. Current-session verification July 21, 2026. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩ -
Hermes Agent v0.18.1 release tag and v0.18.2 release tag, July 7–8, 2026. Infrastructure patch rollups on the v0.18 line; the substantive v0.18.2 fix unpins WhatsApp Baileys to 7.0.0-rc13 for reliable Docker builds. Both patch windows are rolled up into and fully documented by the v0.19.0 release notes. ↩
-
Hermes Agent v0.15.1 release notes and Hermes Agent v0.15.2 release notes. v0.15.1 (May 29, 2026, 01:12 UTC) is the same-day Velocity hotfix: dashboard 401 reload-loop fix on loopback mode; Docker now requires explicit
HERMES_DASHBOARD_INSECURE=1; MCP bare commands (npx,npm,node) resolve in Docker containers; Skills page source pills + category sidebar restored; Kanban workers respond to SIGTERM; Skills.sh catalog grew from 858 to 19,932 entries via sitemap. 28 commits, 21 merged PRs, 9 contributors. v0.15.2 (May 29, 2026, 13:37 UTC) is a packaging-only hotfix that bundlesplugin.yamlmanifests in wheel and sdist distributions so PyPI installs work without sideloading source. 4 contributors. ↩ -
Hermes Agent v0.15.0 release notes and the Hermes Agent releases page. “The Velocity release,” tag
v2026.5.28. Stats: 1,302 commits, 747 merged PRs, 321 community contributors. Refactorsrun_agent.py76% (16,083 → 3,821 lines across 14 modules). Adds the multi-agent Kanban platform (auto-decomposition, swarm topology, per-task model overrides, scheduled tasks, worktree management).session_searchredesigned 4,500× faster with the LLM dependency removed. Promptware defense against Brainworm-class prompt injection at three security chokepoints. Bitwarden Secrets Manager integration replaces multiple per-provider API keys with a single bootstrap token. Skill bundles let multiple skills load with one slash command. TUI session orchestrator for multi-session management in one terminal window. Krea 2 (Medium/Large) and FAL plugin support for image generation. xAI integration round adds a web-search plugin, OAuth upstream, retired-model detection, and natural TTS pauses in voice output. A patch release referenced on GitHub addresses dashboard 401 reload-loop, Docker--insecurerequiring explicitHERMES_DASHBOARD_INSECURE=1env var, MCP bare command resolution (npx,npm,node) in Docker, Skills page rendering, Kanban worker SIGTERM handling, full 19,932-entry Skills catalog via sitemap, and a small batch of.mddelivery, gateway probe safety, web URL redaction, kanban-worker vision capability, and hindsight observation defaults. ↩↩↩↩ -
Hermes Agent v0.11.0 Release Notes. April 23, 2026. “The Interface release” — full React/Ink rewrite of the interactive CLI with a Python JSON-RPC backend (
tui_gateway); pluggable transport architecture (agent/transports/); native AWS Bedrock via Converse API; five new inference paths (NVIDIA NIM, Arcee AI, Step Plan, Google Gemini CLI OAuth, Vercel ai-gateway); GPT-5.5 via Codex OAuth; QQBot as 17th messaging platform with QR-scan setup; expanded plugin surface (slash commands, tool dispatch, execution blocking, result transformation);/steer <prompt>for mid-run agent nudges that inject context after the next tool call without breaking prompt cache; shell hooks for lifecycle events without Python plugins; webhook direct-delivery mode that forwards payloads straight to a platform chat; smarter delegation with orchestrator roles + configurable spawn depth + file coordination; dashboard plugin system, live theme switching, i18n, mobile responsiveness. Stats since v0.9.0: 1,556 commits · 761 merged PRs · 1,314 files changed · 224,174 insertions · 29 community contributors. See also: Hermes Agent v0.11.0 GitHub release tag. ↩↩↩ -
Hermes Agent v0.10.0 Release Notes. April 16, 2026. “The Tool Gateway Release.” Nous Tool Gateway integration for paid Nous Portal subscribers — managed access to Firecrawl web search, FAL / FLUX 2 Pro image generation, OpenAI TTS, and Browser Use browser automation with no extra API keys. Per-tool opt-in via new
use_gatewayconfig field. Runtime prefers gateway over direct API keys when both are configured. Full integration withhermes toolsandhermes status. Replaces deprecatedHERMES_ENABLE_NOUS_MANAGED_TOOLSenv var. Implementation by @jquesnelle (emozilla). Hermes Agent CLI remains MIT-licensed and fully open source; the gateway is an integration with the existing Portal subscription product, not a paywall on the CLI. See also: Nous Portal for subscription pricing and signup. ↩↩↩