Русская версия: README.ru.md
A curated list of what it actually takes to run a personal AI agent in production — not prompt tricks, but the boring layer: schedules, context budgets, secrets, sandboxes, delivery, and the failure modes nobody warns you about.
Everything here is written from an agent that has been running unattended since mid-2026 on one small server: cron jobs that fire at 03:10, a Telegram front end, health data, mail, and GitHub. The practice notes link to a series of 34 modules where each one is written out in full: Hermes-Agent-Ops.
- Start with the failure modes
- Scheduled work
- Context and cost
- Isolation and secrets
- Delivery and attention
- Memory that does not rot
- Watching the machine, not just the model
-
The job that ran and said nothing — a cron that fails silently is worse than one that crashes. Deliverability is a metric, not a hope. (delivery, cron)
-
The job that stopped existing — schedules drift: a one-shot that never re-armed, a job paused "for a day" in March. Audit the schedule itself, not the output. (cron, publishing)
-
The report nobody read — a digest delivered into the wrong channel is the same as not delivered. Route by kind, not by habit. (delivery, routing)
-
The context that quietly shrank — compaction is invisible until it drops the one detail you needed. Measure what compaction costs before you trust it. (context)
-
A cron for the crons — one job whose only purpose is to notice that another job did not run. (cron, observability)
-
Fresh-prompt linting — prompts rot faster than code: they reference jobs, files and names that no longer exist. (cron, evals)
-
From LLM job to script — the best cron is a collector script plus a small model call at the end, not an agent loop. (cron, cost)
-
Task evals and autopsy — when a job fails, reconstruct what it saw instead of re-running it and guessing. (evals)
-
Cost governance — a hard budget with a graceful downgrade beats a surprise invoice. (cost)
-
The cost dashboard — per-job, per-model spend, so "the agent got expensive" becomes a number with a cause. (cost, observability)
-
The thinking layer — pay for reasoning where the decision is, not where the text is. (cost, routing)
-
Compaction that never summarises — the alternative to summarising: delete what is stale, keep the rest verbatim. (context, cost)
-
Tool guardrails — an allow-list per tool, plus a hook that refuses the destructive shape before it runs. (isolation, secrets)
-
A sandbox without root — proot, one writable tree, and a home directory that is not mounted at all. (isolation)
-
Wallet guard — spend limits and a kill switch that does not depend on the agent's own judgement. (secrets, isolation)
-
Publishing an agent's work safely — the denylist that runs before
git push, because "I will remember to check" is not a control. (publishing, secrets)
-
Voice that lands as voice — an audio reply has to arrive as a voice note, at a pace a human can stand, or it may as well be text. (voice, delivery)
-
Delivery obligations — keep a record of what must be delivered and when, and check it; do not infer it from logs. (delivery)
-
Topic routing — one chat, many topics: status goes to the log, decisions go to the human, alerts go where they will be seen. (routing, delivery)
-
Quiet hours and cadence — a digest that arrives when nobody reads it trains the human to ignore digests. (delivery)
-
Memory in topics — a small always-loaded core plus one file per subject, read on demand. (memory)
-
Session housekeeping — sessions are a log, not a filing cabinet; name, archive, prune. (memory, context)
-
Skill library janitor — skills accumulate, contradict each other, and describe tools that were renamed last month. (memory, evals)
-
Blind-spot audit — ask what the agent is not looking at; the answer is usually more useful than a summary. (evals, observability)
-
Gateway supervisor — the process that talks to the chat platform is the single point of failure. (observability, isolation)
-
Harness probes — synthetic checks that prove the whole path works: model, tool, transport, permission. (observability, evals)
-
Ops as data — every job writes a row; the dashboard is a query, not a memory. (observability, datasets)
- Not prompt engineering. Different problem, different list.
- Not a framework. Most of these are a cron entry, a script and a JSON file.
- Not finished. Every item here was added because something broke.
Corrections and additions are welcome — especially failure modes, with the symptom, the cause, and the fix. One line per item, real link, no marketing.
MIT for the list and for the linked modules unless a module says otherwise.