← ayumad.me

Hermes — Live Audit

Mac mini · read in-session 2026-09-10 · Hermes Agent v0.21.0 (2026.8.31)

A point-in-time audit of the agent that runs this site's automations. Every number below was read from the live install in a single session — the cron ledger, the session store, the model-usage tables and the memory database — not estimated.

Snapshot

v0.21.0
2026.8.31
6,640
commits behind
2
profiles
18
cron jobs / 17 live
147
skills
1
MCP server
91%
data volume used
1.4 GB
state.db
924
sessions
1,671
memory records

Broken — needing attention

Reminders ↔ Google Tasks Syncfailing

423 runs, 240 failed. The cause changed on the day of this audit: invalid_grant: Token has been expired or revoked, first seen 02:14, three in a row. At a 15-minute cadence that is roughly 96 failures a day until the token is re-authorised. The Google token file was last written Sep 8 09:30.

x-opencode-session HTTP 400intermittent

30 occurrences in 14 days, spaced exactly every 3 hours — the Marketplace Sweep cadence. Last one Sep 9 20:34. The sweep then ran clean at 02:36 and the gateway restarted at 02:41, and the affinity patch (Sep 8) is loaded, so this is intermittent rather than closed.

No real fallback providerfailing

config.yaml carries a fallback_providers value, but hermes fallback list answers “No fallback providers configured” — phantom protection. That is why the Sep 9 relay 503 took the Daily Brief down with no retry path. It also pointed at the same provider, which is useless during a relay outage, and at a legacy model id.

Album Elo Detectionfix staged, unverified

Failed four nights straight (Sep 6–9 at 23:30) on npm: command not found, exit 127 — a script-only cron does not inherit the shell PATH. The wrapper now exports $HOME/.hermes/node/bin. Reproduced in a minimal environment both ways: with the fix npm 10.9.8, without it command not found.

Ayumad.me Monitor reports into the voidsilent

This very site's staleness check is an agent job with deliver: local. Output is written, never delivered — if the deployed bundle goes stale, nobody is told.

Question Prompt is switched offdisabled

Disabled after 27 runs with 9 failures, last fire Sep 9 23:00. It was scheduled six times a day from the Questions Dashboard, so the reflective-prompt routine is currently not running.

Routing and cost

September usagevalue
API calls2,406
Input tokens19.3 M
Output tokens2.36 M
Cache reads219.4 M
Cron share of calls661 (27%)
Interactive share1,846 (73%)
Heaviest day (Sep 8)920 calls

Cache reads are 91% of all tokens — the ratio that makes a flat-rate subscription so hard to beat. Cost columns read $0.00: flat rate, as expected.

Modelcallscache
deepseek-v4-flash1,517171.9 M
mimo-v2.569935.9 M
deepseek-flash16018.0 M
qwen3.8-max941.5 M
gpt-5.6-luna10—
v4-flash-vision-exp10—

All on one relay provider except gpt-5.6-luna (included). The coder profile and subagent delegation both resolve to the cheap model; context compression runs on a long-context model. No wasted spend in the auxiliary paths.

Fixed and verified — during this audit

Vision model off a retired idfixed

The vision auxiliary pointed at a retired V4-vision model id that was only alive through temporary compatibility routing. Repointed to the current model and verified.

MCP handshake confirmedverified

The single MCP server connects over stdio in 812 ms with 2 tools discovered. Configuration-only checks were hiding nothing.

Health sentinel is right, not brokenverified

Its non-zero exit is the designed warning path, not a crash. The warnings it raised were true: the data volume is 92% full and four jobs are in a non-ok state.

Cost landmines still armed

ItemWhy it matters
Mixture-of-Agents enabled
aggregator on an abandoned provider
MoA is enabled: true with a premium aggregator and a reference model on a provider that was dropped for cost reasons. September shows only 10 reference-model calls, so it is not auto-firing — but it is a loaded gun aimed at a provider no longer being paid for.
max_tokens: 2048 Low for a coding daily driver on a 1M-context model — long diffs and reports get truncated mid-write. Raising it costs nothing; you only pay for tokens actually produced.
Weekly Kanban Dispatch Burns a full agent run every Monday on a board sitting at 6 done and 0 open. Either the workflow needs reviving or the job should become a script.
Install drift
6,640 commits · 4 modified + 1 untracked
A local patch that fixes the relay's session-affinity requirement exists only as an untracked file. The next upgrade can conflict with or clobber it.
Disk at 91%
19 GiB free
Reclaimable: a stale 516 MB emergency database backup, plus an offline session-store rebuild. The 656 MB of embedding models are load-bearing — those stay.

Worth building next

Local model on the homelab GPU
An RTX 3060 sits in the rack doing passthrough while ~660 cron calls a month go to a cloud relay. A 7–14B local model would take the routine load to zero marginal cost and survive relay outages.
Expose the subscription to other tools
A local OpenAI-compatible proxy backed by the existing subscription would let other CLI coding tools use it instead of paying twice for the same models.
Persistent specialist bots
Everything currently runs as an ephemeral session or a scheduled fire. Durable named specialists with their own memory boundaries would be the closest thing to a real fleet.
Finish the vault integration
The API server and webhook surfaces are disabled even though the credentials are already in place. The vault is the source of truth, so talking to the agent from inside it is the natural next surface.
Second profile for school and research
One profile carries everything. A separate one for coursework and research gets its own skills, schedule and memory with no cross-contamination.
Desktop panes and widgets
The desktop surface accepts UI plugins. A homelab and cron health pane reading the same stores the sentinel script uses is the obvious first one.

Deliberately unchanged

Compression and delegation routing — cheap, correct, matched to the flat-rate plan; auxiliary paths inherit the default model, so nothing is overspending there.

Session auto-prune with 90-day retention — already enabled, so the doctor's complaint about it is stale. The remaining work is the offline storage rebuild.

Webhooks, always-on voice, extra MCP servers — left off. No consumer, no event source, no security boundary yet. Fixing what is already broken comes before installing another integration.

Vault auto-sync, every 15 minutes — 426 of 426 runs clean. The only job with a perfect record.