A minimal, dependency-free agent harness in TypeScript — built to understand how agent runtimes like Claude Code and Codex actually work inside.
In 2026 every serious AI product runs on an agent harness: the runtime layer around an LLM that runs the loop, calls tools, manages context, enforces permissions, and carries state across turns. OpenAI open-sourced theirs; Anthropic's powers Claude Code; a dozen OSS projects compete on it. The core is famously small — "a while loop with tools" — but the engineering lives in everything that keeps that loop safe and cheap to run.
This repo is my way through that stack: I studied how the production
harnesses work (notes in docs/), then built a
faithful-but-tiny one. Zero runtime dependencies. Every line is meant to
be read.
const result = await runAgent("what does the note say?", [], {
provider: createOpenAIProvider({ apiKey, baseUrl, model }),
systemPrompt: "You are a concise assistant. Inspect before you conclude.",
tools, // ToolRegistry — read/write/edit/list_dir built in
permissions: defaultPolicy(), // allow reads, ask before writes, fail closed
toolContext: { workDir: process.cwd() },
}); ┌──────────────────────────────────────────────┐
│ runAgent() │
│ src/core/loop.ts │
│ │
│ ┌────────────────────────────────────┐ │
│ │ while budget && iterations left │ │
│ │ │ │
│ │ compact if over token budget │ │
│ │ │ │ │
│ │ ▼ │ │
│ │ provider.complete() ── text ──▶ onEvent stream
│ │ │ │ │ (REPL / UI / tests)
│ │ no tool calls? ── yes ──▶ done │ │
│ │ │ no │ │
│ │ permissions.resolve(each call) │ │
│ │ tools.execute(each allowed call) │ │
│ │ append tool_results ──▶ loop │ │
│ └────────────────────────────────────┘ │
└──────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
providers/ tools/ context/
openai · anthropic registry · tokens · compactor
(fetch + SSE) permissions · (keep tool_use/result
builtin fs tools pairs together!)
| Concern | Where | Choices worth reading |
|---|---|---|
| Agent loop | src/core/loop.ts |
model-decides-when-to-stop, iteration budget, abort support |
| Providers | src/providers/ |
OpenAI-compatible + Anthropic, plain fetch, hand-rolled SSE, streamed tool-arg reassembly |
| Tools | src/tools/ |
registry with uniform error capture + output caps; read/write/edit/list_dir with path-escape protection |
| Permissions | src/tools/permissions.ts |
allow / ask / deny, human-in-the-loop approver, fails closed |
| Context mgmt | src/context/ |
token budget → summarize old turns, keep recent tail, never split a tool_use from its results |
| Sessions | src/core/session.ts |
history persists per project, --resume to continue |
| CLI | src/cli.ts |
streaming REPL + one-shot -p mode |
git clone https://github.com/wjdjdakf17/mini-harness && cd mini-harness
npm install
# Works with any OpenAI-compatible endpoint (OpenAI, GLM, Ollama, vLLM, LiteLLM…)
export OPENAI_API_KEY=sk-...
export OPENAI_BASE_URL=https://api.openai.com/v1
export OPENAI_MODEL=gpt-4.1-mini
npm run cli # interactive; /save /sessions /exit
npm run cli -- -p "list the files and tell me what this project is"
# …or Anthropic:
# export MINI_HARNESS_PROVIDER=anthropic ANTHROPIC_API_KEY=sk-ant-... ANTHROPIC_MODEL=claude-sonnet-5Writes require interactive approval (allow write_file({"path":...})? [y/N]);
in one-shot mode there is no approver, so write-class tools fail closed —
that's the policy working, not a bug.
30 deterministic tests, no network, no keys — the provider boundary is faked with a scripted provider, so loop ordering, permission gates, compaction pair-safety, wire-format translation, and SSE chunk reassembly are all pinned:
npm testThe docs/ folder is the other half of the repo — what I learned building
this, with sources:
- What is an agent harness? — the term, why it took over, "the model sets the ceiling; the harness sets the price"
- The agent loop — anatomy, termination, events as the product surface
- Tools and permissions — tools as contracts, the allow/ask/deny gate, why no shell
- Context management — replay vs compaction vs lazy injection; the pair-preserving cut; the ARC-AGI-3 numbers
- The 2026 harness landscape — opencode · Pi · OpenHands · deepagents · Codex, and what actually separates them
- Design decisions — ADRs for this codebase, and what I'd build next
- References
No sandboxing (hence no shell tool), no MCP, no subagents, no real tokenizer (char-count estimation instead — it only decides when to compact). Each omission is deliberate and written down in the design decisions. This is a study instrument, not a product — if you need production, go use opencode or Codex; read this first so you know what they're doing for you.