Engineering workflows for Claude Code, Codex, and OpenCode: specialist agents, skills, task ledgers, contribution checks, orchestration loops, and safety hooks.
Forge is an installable engineering toolkit for Claude Code, Codex, and OpenCode. It combines role definitions, engineering methods, task ledgers, model-routing contracts, and local verification scripts. Prompts and source scripts are inspectable; host-specific hooks and execution permissions remain separate from portable skills. Local tests establish the documented contracts, not a comparative coding-performance score.
- Why Forge
- Install
- What's inside
- Release provenance
- How the pieces fit
- Repository layout
- Documentation
- Contributing
- License
LLM coding agents are only as good as the scaffolding around them. The same model produces dramatically different results depending on whether it has a sharp role, a proven method, scoped tools, and guardrails. Forge encodes that scaffolding:
- Specialists, not a generalist. Twenty agents each with a focused role, a concrete methodology, scoped tools, and a defined output format — a reviewer that thinks like a reviewer, a debugger that finds root causes, an auditor that traces taint to sinks.
- Orchestration for big work. Plan at Opus/Fable, decompose into a task ledger, route implementation to Sonnet, fan out mechanical work to Haiku, and iterate to verified done.
- Durable local execution. SQLite/WAL run history, idempotent lifecycle events, tamper-evident replay, digest-pinned workflow code/schema and worker definitions, fail-closed compatibility gates for replay, migration, restore, and effect retry, checkpointed recovery, offline lineage and signed provenance verification, and a strict separation between execution state, task planning, and privacy-safe receipts give long-running orchestration a recoverable foundation. The portable backend facade also models etcd-first distributed revisions, watch delivery, verified snapshots, and compaction recovery without treating provider metadata as canonical history.
- Stacked delivery, now native. Design and verify dependent PRs with a portable stack
manifest, default to GitHub's first-party
gh stack, or adapt the same safety protocol to vanilla GitHub, Graphite, Aviator, Sapling, and classic ghstack. - Authorization before effects. Declarative policy profiles bind exact actions to principals, resources, revisions, and one-use approvals, with staged previews and committed-effect receipts for GitHub mutations, releases, and production workflows.
- Discoverability without bloat.
/forge, CATALOG.md, bundles, and workflows route work to the smallest useful capability instead of dumping every skill into context. - Canonical capability contract. A body-aware v2 graph records component identity, instructions, tools, permissions, resources, eval links, and explicit Claude/Codex/ Agent Skills projections so host compatibility is reviewable and drift fails the gate. See Capability IR.
- Methodology on tap. Twenty-six skills inject engineering practices — TDD, root-cause debugging, threat modeling, safe migrations, orchestration, catalogs, task ledgers, and solve loops — exactly when the situation calls for them.
- One-keystroke workflows. Twenty-two slash commands wrap the everyday loop: forge, review, test, debug, plan, commit, PR, orchestrate, tasks, solve-loop, stack, stack-review.
- Safety by default. Lifecycle hooks block catastrophic commands and secret leaks, auto-format edits, inject repo context at session start, and notify you on completion — deterministically, without relying on the model to remember.
- Proven, not asserted. A real eval harness scores prompts and high-risk behavior
contracts (336/337 deterministic checks, one warning, plus cross-host scenarios and an opt-in
LLM judge); the test suite covers safety hooks, task sync, receipts, durable runtime replay, external effect
delivery, doctor, policy,
stacks, marketplace readiness, capability graph, rendering, semantic evidence, and conformance. Run
them yourself —
just check. - Auditable & self-validating. Read every prompt and script. CI validates structure, runs the tests, and scores the evals on every push. GitHub-backed task ledgers add stable issue identity, native task graphs, conflict stops, and resumable evidence.
- GitHub-native workflow bridge. A pinned
gh-awadapter projects Forge orchestration into deterministic workflow sources, read-only agent jobs, staged safe outputs, and source-to-lock drift checks. Its optional provider worker adds fenced leases, exact one-use approvals, account verification, idempotent recovery, and reference-only GitHub receipts without committing credentials. - Upstream contribution evidence. Review maintainer requirements, run explicit checks and verify digest-only receipts against a clean exact Git revision. The OSS contribution skill separates local verification from publication, legal attestations and merge authority.
Claude Code plugin (recommended — includes agents, commands, and hooks):
# In Claude Code:
/plugin marketplace add AlisinaDevelo/md-files
/plugin install forge@forgeCodex plugin (skills and orchestration methods):
codex plugin marketplace add AlisinaDevelo/md-files
codex plugin add forge@forgeOpenCode (Agent Skills and project instructions):
git clone https://github.com/AlisinaDevelo/md-files.git
cd md-files
./scripts/install-opencode.sh --copySee OpenCode support for copy, symlink, verification, and host-boundary details.
Repository marketplace installation is available for Claude Code and Codex, while OpenCode uses the Agent Skills installer. Forge has not been submitted to the public Claude or Codex directories; see Marketplace readiness for the dated publication state, publisher surfaces, and submission evidence.
As user-level symlinks (Claude agents, skills, commands):
git clone https://github.com/AlisinaDevelo/md-files.git && cd md-files
./scripts/install.sh # or --copy / --dry-runAs .agents skills (Codex, Zed, and OpenCode):
git clone https://github.com/AlisinaDevelo/md-files.git && cd md-files/zed
./install.sh # installs Forge skills into ~/.agents/skillsCherry-pick: everything is plain Markdown — copy any file into your own ~/.claude/
or project .claude/. See docs/getting-started.md for details.
Delegated, autonomous specialists with their own context and scoped tools.
| Agent | Role |
|---|---|
code-reviewer |
Severity-ranked review for correctness, security, and maintainability |
debugger |
Hypothesis-driven root-cause diagnosis |
security-auditor |
Defensive vulnerability review (OWASP/CWE, taint→sink) |
test-engineer |
Behavior-focused tests that reduce real risk |
architect |
Implementation plans and architectural trade-offs |
refactoring-specialist |
Behavior-preserving structural improvement |
performance-optimizer |
Measure-first bottleneck diagnosis and fixes |
database-expert |
Schema design, query tuning, safe migrations |
api-designer |
Consistent, evolvable API contracts |
frontend-specialist |
Component architecture, state, render performance |
accessibility-auditor |
WCAG audit and remediation |
dependency-auditor |
CVEs, license risk, safe upgrade planning |
devops-engineer |
CI/CD, containers, IaC, safe deploys |
docs-writer |
Accurate, example-driven documentation |
incident-responder |
Triage, mitigate, then root-cause + postmortem |
code-archaeologist |
Understand unfamiliar/legacy code before changing it |
migration-specialist |
Incremental, reversible framework/library/API migrations |
data-engineer |
Data pipelines, ETL/ELT, warehouse modeling, data quality |
sre |
SLOs, error budgets, capacity, toil reduction, reliability |
tech-lead |
Orchestrates large tasks across the specialists |
Methodologies and references injected into the current conversation when the situation
matches. Several use progressive disclosure — a lean SKILL.md plus deeper reference
files loaded only when needed.
| Skill | When it fires |
|---|---|
test-driven-development |
Implementing test-first (red-green-refactor) |
root-cause-debugging |
Diagnosing a bug or failure |
code-review-rubric |
Reviewing code (+ full checklist) |
refactoring-catalog |
Improving structure (+ smell→fix catalog) |
conventional-commits |
Writing commit messages |
pull-request-authoring |
Opening a reviewable PR |
api-design |
Designing or reviewing an API |
threat-modeling |
Security design review (STRIDE) |
safe-database-migrations |
Schema changes on live data |
performance-profiling |
Investigating performance |
observability |
Adding logs/metrics/traces |
technical-writing |
Writing developer docs |
git-workflow |
Branching, rebasing, conflicts, recovery, bisect |
error-handling |
Designing robust failure paths |
feature-flags |
Gating, progressive rollout, and flag cleanup |
caching-strategies |
Cache patterns, TTLs, invalidation, stampedes |
concurrency-and-parallelism |
Races, locks, async, idempotency |
prompt-engineering |
Authoring agents/skills/commands (+ patterns) |
forge-catalog |
Choose the right Forge command, agent, skill, bundle, or workflow |
orchestration |
Multi-model planning, delegation, integration, and verification |
task-ledger |
Jira/GitHub-issue-like local tasks with status, deps, agent, and model |
iterate-to-done |
Solve-loop discipline for draining a ledger until done or blocked |
stacked-changes |
GitHub-native and vendor-neutral stacked PR design, review, native reconciliation, restack, recovery, and landing |
doctor |
Read-only host, capability, repository-policy, and merge-readiness diagnostics |
policy |
Declarative authorization, staged previews, scoped approvals, and decision receipts |
User-triggered prompt templates with argument and shell injection.
| Command | Does |
|---|---|
/forge |
Choose the right Forge agent, skill, command, bundle, or workflow |
/review |
Review the current diff, severity-ranked |
/commit |
Draft a Conventional Commit for staged changes |
/test |
Write tests matching the repo's harness |
/debug |
Root-cause a bug before fixing |
/plan |
Step-by-step implementation plan |
/refactor |
Behavior-preserving cleanup |
/security-scan |
Defensive security review of the diff |
/pr |
Draft a PR description from the branch |
/optimize |
Measure-first performance fix |
/explain |
Explain a file, symbol, or system |
/docs |
Write docs grounded in the code |
/tidy |
Remove cruft from the diff, behavior-preserving |
/changelog |
Draft a changelog entry from commits since the last release |
/scaffold |
Scaffold a new module/component matching repo conventions |
/orchestrate |
Plan a big goal, create a task ledger, route work by agent/model, and drive it to done |
/tasks |
Create, list, update, or GitHub-sync the task ledger |
/solve-loop |
Drain ready ledger tasks with verify-before-done discipline |
/stack |
Plan, inspect, submit, reconcile, restack, repair, or land dependent pull requests |
/stack-review |
Review every stack layer bottom-up against its immediate parent |
/doctor |
Run the read-only Forge capability and merge-readiness preflight |
/policy |
Evaluate policy, stage effects, issue approvals, authorize, and record outcomes |
Deterministic guardrails the harness runs on lifecycle events — no model memory required.
| Hook | Event | Effect |
|---|---|---|
session-context |
SessionStart | Injects current branch, ahead/behind, dirty count, and recent commits as context |
guard-bash |
PreToolUse(Bash) | Blocks catastrophic commands (rm -rf /, force-push to main, fork bombs) |
scan-secrets |
PreToolUse(Write/Edit) | Blocks writing credentials into files |
format-file |
PostToolUse(Write/Edit) | Auto-formats edited files with the installed formatter |
notify |
Stop | Desktop notification when a turn finishes |
output-styles/— selectable system-prompt modes: Concise Engineer (answer-first, no preamble) and Mentor (teaches the why as it works). Ship with the plugin; pick one via/config.statusline/— a status line showing model · dir · git · context% · cost.settings/— examplesettings.json(permission allowlist, deny rules for secrets, status line, output style) to pair with the plugin.
instructions/— aCLAUDE.mdtemplate library, stack-agnostic engineering principles, and language snippets (TypeScript, Python, Go).mcp/— example Model Context Protocol server configs with least-privilege guidance.
evals/— deterministic prompt-quality and behavior-contract checks, shared cross-host scenarios, and an opt-in LLM-judge eval that scores agents against real tasks.tests/— pytest cases covering safety hooks, task sync, receipts, durable runtime replay, external effect delivery, doctor, stacks, and conformance.just checkruns it all.
Tagged releases publish deterministic Claude, Codex, .agents, and OpenAI skills-only
bundles with SHA-256 manifests, SPDX SBOMs, and an offline verifier. Hosted releases can
also supply GitHub artifact attestations; local-only releases do not imply hosted provenance. See
release provenance for consumer verification and the
threat model.
flowchart LR
goal["big goal\n/orchestrate"] --> ledger["task ledger\n/tasks"]
ledger --> route["route by agent + model\nOpus/Fable · Sonnet · Haiku"]
route --> loop["solve loop\n/solve-loop"]
loop --> done["verified done"]
ledger --> topology{"one PR or stack?"}
topology --> stack["stack graph\n/stack"]
stack --> stackreview["incremental review\n/stack-review"]
stackreview --> done
plan["plan\narchitect / /plan"] --> impl[implement]
impl --> review["review\ncode-reviewer / /review"]
review --> test["test\ntest-engineer / /test"]
test --> debug["debug\ndebugger / /debug"]
debug --> ship["ship\n/commit · /pr"]
guard(["guardrails — always on\nguard-bash · scan-secrets · format-file · notify"])
guard -. wraps .-> impl
guard -. wraps .-> review
guard -. wraps .-> debug
guard -. wraps .-> ship
Agents go deep on focused jobs; skills supply the method; commands trigger the loop; hooks keep it safe. For big tasks, the main conversation acts as the conductor so it can spawn specialists in parallel; the ledger keeps the run honest. See docs/usage-patterns.md.
.claude-plugin/ Claude Code marketplace manifest
.agents/plugins/ Codex marketplace manifest
data/ generated catalog, capability graph, bundles, and workflow metadata
plugins/forge/ the Forge plugin
.claude-plugin/ plugin manifest
.codex-plugin/ Codex plugin manifest
agents/ 20 specialist subagents
skills/ 26 progressive-disclosure skills
commands/ 22 slash commands
hooks/ 5 lifecycle hooks (session-context, guard, secrets, format, notify)
output-styles/ selectable system-prompt modes
instructions/ CLAUDE.md templates, principles, language guides
mcp/ example MCP server configs
statusline/ status line script
settings/ example settings.json
evals/ prompt eval harness, shared scenarios, and result evidence
tests/ runnable hook, task-ledger, receipt, doctor, stack, and conformance tests
docs/ getting started, usage, architecture, rationale, CI
scripts/ validation, installation, release, and marketplace checks
.github/ CI, issue/PR templates, CODEOWNERS, dependabot
- Getting started — install options and first steps
- OSS release contract — contribution checks, research basis, native compatibility evidence, and release acceptance
- OpenHands SDK — optional native AgentSkills compatibility verification
- Usage patterns — how the components combine in real workflows
- Bundles & workflows — focused capability sets and ordered playbooks
- Quality bar — validation and safety standards for Forge components
- Competitive audit — what Forge borrows from larger skill libraries
- Frontier roadmap — research-backed interoperability, identity, evaluation, release evidence, and connected-execution priorities
- Stacked changes — GitHub-native stacks, provider adapters, safety model, review flow, CI, and recovery
- GitHub native stacks — remote inspect/import, SHA-guarded reconciliation, divergence classes, mutation authority, and preview fallback
- Policy plane — action envelopes, profiles, approvals, staged previews, adapter integration, and privacy-safe decision evidence
- Cross-host conformance — shared scenarios, host adapters, live evidence, result schemas, and release gates
- Release provenance — deterministic bundles, SBOMs, attestations, offline verification, and threat model
- Marketplace readiness — honest directory status, publisher surfaces, asset policy, and submission smoke-test matrix
- OpenAI Agent Plugins compatibility — current universal plugin contract, Forge audit, submission boundary, and staged compatibility plan
- OpenAI submission packet — reproducible candidate archive evidence, five positive/three negative cases, and the external publication boundary
- Capability IR — body-aware graph, deterministic host renderer, adapter contract, migration workflow, and current compiler boundary
- Durable runtime — local SQLite/WAL history, deterministic replay, transactional outbox/inbox effects, generation-fenced heartbeats, lease evidence, checkpointed recovery, human-input waits, signals, MCP Tasks projection, reviewed migrations, offline lineage verification, signed trace/provenance evidence, idempotency, hash-chain verification, and explicit at-least-once boundaries
- Runtime provenance — signed trace correlation, privacy defaults, offline trust verification, key rotation, retention, and incident response
- GitHub Agentic Workflows — pinned gh-aw compilation, read-only agents, staged safe outputs, policy evidence, native lock verification, and durable episode correlation
- Architecture — how the repo is organized and why
- Design rationale — the decisions and trade-offs behind Forge
- CI & headless usage — run Forge in pipelines and automated review
- Evals — the evidence layer · Tests — hook test suite
- Contributing — add an agent, skill, command, or hook
- Changelog
Contributions are welcome — new agents, skills, commands, and hooks, or improvements to
existing ones. Run ./scripts/validate.sh before opening a PR. See
CONTRIBUTING.md for the conventions and the quality bar.
MIT © Alisina Karimi