Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

mini-harness

A minimal, dependency-free agent harness in TypeScript — built to understand how agent runtimes like Claude Code and Codex actually work inside.

In 2026 every serious AI product runs on an agent harness: the runtime layer around an LLM that runs the loop, calls tools, manages context, enforces permissions, and carries state across turns. OpenAI open-sourced theirs; Anthropic's powers Claude Code; a dozen OSS projects compete on it. The core is famously small — "a while loop with tools" — but the engineering lives in everything that keeps that loop safe and cheap to run.

This repo is my way through that stack: I studied how the production harnesses work (notes in docs/), then built a faithful-but-tiny one. Zero runtime dependencies. Every line is meant to be read.

const result = await runAgent("what does the note say?", [], {
  provider: createOpenAIProvider({ apiKey, baseUrl, model }),
  systemPrompt: "You are a concise assistant. Inspect before you conclude.",
  tools,              // ToolRegistry — read/write/edit/list_dir built in
  permissions: defaultPolicy(),   // allow reads, ask before writes, fail closed
  toolContext: { workDir: process.cwd() },
});

The whole harness in one diagram

            ┌──────────────────────────────────────────────┐
            │                 runAgent()                   │
            │               src/core/loop.ts               │
            │                                              │
            │   ┌────────────────────────────────────┐     │
            │   │  while budget && iterations left   │     │
            │   │                                    │     │
            │   │   compact if over token budget     │     │
            │   │            │                       │     │
            │   │            ▼                       │     │
            │   │   provider.complete() ── text ──▶ onEvent stream
            │   │            │                       │     │  (REPL / UI / tests)
            │   │   no tool calls? ── yes ──▶ done   │     │
            │   │            │ no                    │     │
            │   │   permissions.resolve(each call)   │     │
            │   │   tools.execute(each allowed call) │     │
            │   │   append tool_results ──▶ loop     │     │
            │   └────────────────────────────────────┘     │
            └──────────────────────────────────────────────┘
                    │                  │                │
                    ▼                  ▼                ▼
          providers/            tools/            context/
          openai · anthropic    registry ·        tokens · compactor
          (fetch + SSE)         permissions ·     (keep tool_use/result
                                 builtin fs tools   pairs together!)

What's implemented

Concern Where Choices worth reading
Agent loop src/core/loop.ts model-decides-when-to-stop, iteration budget, abort support
Providers src/providers/ OpenAI-compatible + Anthropic, plain fetch, hand-rolled SSE, streamed tool-arg reassembly
Tools src/tools/ registry with uniform error capture + output caps; read/write/edit/list_dir with path-escape protection
Permissions src/tools/permissions.ts allow / ask / deny, human-in-the-loop approver, fails closed
Context mgmt src/context/ token budget → summarize old turns, keep recent tail, never split a tool_use from its results
Sessions src/core/session.ts history persists per project, --resume to continue
CLI src/cli.ts streaming REPL + one-shot -p mode

Quickstart

git clone https://github.com/wjdjdakf17/mini-harness && cd mini-harness
npm install

# Works with any OpenAI-compatible endpoint (OpenAI, GLM, Ollama, vLLM, LiteLLM…)
export OPENAI_API_KEY=sk-...
export OPENAI_BASE_URL=https://api.openai.com/v1
export OPENAI_MODEL=gpt-4.1-mini

npm run cli          # interactive; /save /sessions /exit
npm run cli -- -p "list the files and tell me what this project is"

# …or Anthropic:
# export MINI_HARNESS_PROVIDER=anthropic ANTHROPIC_API_KEY=sk-ant-... ANTHROPIC_MODEL=claude-sonnet-5

Writes require interactive approval (allow write_file({"path":...})? [y/N]); in one-shot mode there is no approver, so write-class tools fail closed — that's the policy working, not a bug.

Tests

30 deterministic tests, no network, no keys — the provider boundary is faked with a scripted provider, so loop ordering, permission gates, compaction pair-safety, wire-format translation, and SSE chunk reassembly are all pinned:

npm test

Study notes

The docs/ folder is the other half of the repo — what I learned building this, with sources:

  1. What is an agent harness? — the term, why it took over, "the model sets the ceiling; the harness sets the price"
  2. The agent loop — anatomy, termination, events as the product surface
  3. Tools and permissions — tools as contracts, the allow/ask/deny gate, why no shell
  4. Context management — replay vs compaction vs lazy injection; the pair-preserving cut; the ARC-AGI-3 numbers
  5. The 2026 harness landscape — opencode · Pi · OpenHands · deepagents · Codex, and what actually separates them
  6. Design decisions — ADRs for this codebase, and what I'd build next
  7. References

What this is not

No sandboxing (hence no shell tool), no MCP, no subagents, no real tokenizer (char-count estimation instead — it only decides when to compact). Each omission is deliberate and written down in the design decisions. This is a study instrument, not a product — if you need production, go use opencode or Codex; read this first so you know what they're doing for you.

License

MIT

About

A minimal, dependency-free agent harness in TypeScript — built to understand how agent runtimes like Claude Code and Codex work inside.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages