Skip to content

Latest commit

 

History

211 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🔨 Forge

Engineering workflows for Claude Code, Codex, and OpenCode: specialist agents, skills, task ledgers, contribution checks, orchestration loops, and safety hooks.

License: MIT Claude Code Agents Skills Commands Tests Prompt evals


Forge is an installable engineering toolkit for Claude Code, Codex, and OpenCode. It combines role definitions, engineering methods, task ledgers, model-routing contracts, and local verification scripts. Prompts and source scripts are inspectable; host-specific hooks and execution permissions remain separate from portable skills. Local tests establish the documented contracts, not a comparative coding-performance score.

Table of contents

Why Forge

LLM coding agents are only as good as the scaffolding around them. The same model produces dramatically different results depending on whether it has a sharp role, a proven method, scoped tools, and guardrails. Forge encodes that scaffolding:

  • Specialists, not a generalist. Twenty agents each with a focused role, a concrete methodology, scoped tools, and a defined output format — a reviewer that thinks like a reviewer, a debugger that finds root causes, an auditor that traces taint to sinks.
  • Orchestration for big work. Plan at Opus/Fable, decompose into a task ledger, route implementation to Sonnet, fan out mechanical work to Haiku, and iterate to verified done.
  • Durable local execution. SQLite/WAL run history, idempotent lifecycle events, tamper-evident replay, digest-pinned workflow code/schema and worker definitions, fail-closed compatibility gates for replay, migration, restore, and effect retry, checkpointed recovery, offline lineage and signed provenance verification, and a strict separation between execution state, task planning, and privacy-safe receipts give long-running orchestration a recoverable foundation. The portable backend facade also models etcd-first distributed revisions, watch delivery, verified snapshots, and compaction recovery without treating provider metadata as canonical history.
  • Stacked delivery, now native. Design and verify dependent PRs with a portable stack manifest, default to GitHub's first-party gh stack, or adapt the same safety protocol to vanilla GitHub, Graphite, Aviator, Sapling, and classic ghstack.
  • Authorization before effects. Declarative policy profiles bind exact actions to principals, resources, revisions, and one-use approvals, with staged previews and committed-effect receipts for GitHub mutations, releases, and production workflows.
  • Discoverability without bloat. /forge, CATALOG.md, bundles, and workflows route work to the smallest useful capability instead of dumping every skill into context.
  • Canonical capability contract. A body-aware v2 graph records component identity, instructions, tools, permissions, resources, eval links, and explicit Claude/Codex/ Agent Skills projections so host compatibility is reviewable and drift fails the gate. See Capability IR.
  • Methodology on tap. Twenty-six skills inject engineering practices — TDD, root-cause debugging, threat modeling, safe migrations, orchestration, catalogs, task ledgers, and solve loops — exactly when the situation calls for them.
  • One-keystroke workflows. Twenty-two slash commands wrap the everyday loop: forge, review, test, debug, plan, commit, PR, orchestrate, tasks, solve-loop, stack, stack-review.
  • Safety by default. Lifecycle hooks block catastrophic commands and secret leaks, auto-format edits, inject repo context at session start, and notify you on completion — deterministically, without relying on the model to remember.
  • Proven, not asserted. A real eval harness scores prompts and high-risk behavior contracts (336/337 deterministic checks, one warning, plus cross-host scenarios and an opt-in LLM judge); the test suite covers safety hooks, task sync, receipts, durable runtime replay, external effect delivery, doctor, policy, stacks, marketplace readiness, capability graph, rendering, semantic evidence, and conformance. Run them yourself — just check.
  • Auditable & self-validating. Read every prompt and script. CI validates structure, runs the tests, and scores the evals on every push. GitHub-backed task ledgers add stable issue identity, native task graphs, conflict stops, and resumable evidence.
  • GitHub-native workflow bridge. A pinned gh-aw adapter projects Forge orchestration into deterministic workflow sources, read-only agent jobs, staged safe outputs, and source-to-lock drift checks. Its optional provider worker adds fenced leases, exact one-use approvals, account verification, idempotent recovery, and reference-only GitHub receipts without committing credentials.
  • Upstream contribution evidence. Review maintainer requirements, run explicit checks and verify digest-only receipts against a clean exact Git revision. The OSS contribution skill separates local verification from publication, legal attestations and merge authority.

Install

Claude Code plugin (recommended — includes agents, commands, and hooks):

# In Claude Code:
/plugin marketplace add AlisinaDevelo/md-files
/plugin install forge@forge

Codex plugin (skills and orchestration methods):

codex plugin marketplace add AlisinaDevelo/md-files
codex plugin add forge@forge

OpenCode (Agent Skills and project instructions):

git clone https://github.com/AlisinaDevelo/md-files.git
cd md-files
./scripts/install-opencode.sh --copy

See OpenCode support for copy, symlink, verification, and host-boundary details.

Repository marketplace installation is available for Claude Code and Codex, while OpenCode uses the Agent Skills installer. Forge has not been submitted to the public Claude or Codex directories; see Marketplace readiness for the dated publication state, publisher surfaces, and submission evidence.

As user-level symlinks (Claude agents, skills, commands):

git clone https://github.com/AlisinaDevelo/md-files.git && cd md-files
./scripts/install.sh        # or --copy / --dry-run

As .agents skills (Codex, Zed, and OpenCode):

git clone https://github.com/AlisinaDevelo/md-files.git && cd md-files/zed
./install.sh                # installs Forge skills into ~/.agents/skills

Cherry-pick: everything is plain Markdown — copy any file into your own ~/.claude/ or project .claude/. See docs/getting-started.md for details.

What's inside

Agents

Delegated, autonomous specialists with their own context and scoped tools.

Agent Role
code-reviewer Severity-ranked review for correctness, security, and maintainability
debugger Hypothesis-driven root-cause diagnosis
security-auditor Defensive vulnerability review (OWASP/CWE, taint→sink)
test-engineer Behavior-focused tests that reduce real risk
architect Implementation plans and architectural trade-offs
refactoring-specialist Behavior-preserving structural improvement
performance-optimizer Measure-first bottleneck diagnosis and fixes
database-expert Schema design, query tuning, safe migrations
api-designer Consistent, evolvable API contracts
frontend-specialist Component architecture, state, render performance
accessibility-auditor WCAG audit and remediation
dependency-auditor CVEs, license risk, safe upgrade planning
devops-engineer CI/CD, containers, IaC, safe deploys
docs-writer Accurate, example-driven documentation
incident-responder Triage, mitigate, then root-cause + postmortem
code-archaeologist Understand unfamiliar/legacy code before changing it
migration-specialist Incremental, reversible framework/library/API migrations
data-engineer Data pipelines, ETL/ELT, warehouse modeling, data quality
sre SLOs, error budgets, capacity, toil reduction, reliability
tech-lead Orchestrates large tasks across the specialists

Skills

Methodologies and references injected into the current conversation when the situation matches. Several use progressive disclosure — a lean SKILL.md plus deeper reference files loaded only when needed.

Skill When it fires
test-driven-development Implementing test-first (red-green-refactor)
root-cause-debugging Diagnosing a bug or failure
code-review-rubric Reviewing code (+ full checklist)
refactoring-catalog Improving structure (+ smell→fix catalog)
conventional-commits Writing commit messages
pull-request-authoring Opening a reviewable PR
api-design Designing or reviewing an API
threat-modeling Security design review (STRIDE)
safe-database-migrations Schema changes on live data
performance-profiling Investigating performance
observability Adding logs/metrics/traces
technical-writing Writing developer docs
git-workflow Branching, rebasing, conflicts, recovery, bisect
error-handling Designing robust failure paths
feature-flags Gating, progressive rollout, and flag cleanup
caching-strategies Cache patterns, TTLs, invalidation, stampedes
concurrency-and-parallelism Races, locks, async, idempotency
prompt-engineering Authoring agents/skills/commands (+ patterns)
forge-catalog Choose the right Forge command, agent, skill, bundle, or workflow
orchestration Multi-model planning, delegation, integration, and verification
task-ledger Jira/GitHub-issue-like local tasks with status, deps, agent, and model
iterate-to-done Solve-loop discipline for draining a ledger until done or blocked
stacked-changes GitHub-native and vendor-neutral stacked PR design, review, native reconciliation, restack, recovery, and landing
doctor Read-only host, capability, repository-policy, and merge-readiness diagnostics
policy Declarative authorization, staged previews, scoped approvals, and decision receipts

Commands

User-triggered prompt templates with argument and shell injection.

Command Does
/forge Choose the right Forge agent, skill, command, bundle, or workflow
/review Review the current diff, severity-ranked
/commit Draft a Conventional Commit for staged changes
/test Write tests matching the repo's harness
/debug Root-cause a bug before fixing
/plan Step-by-step implementation plan
/refactor Behavior-preserving cleanup
/security-scan Defensive security review of the diff
/pr Draft a PR description from the branch
/optimize Measure-first performance fix
/explain Explain a file, symbol, or system
/docs Write docs grounded in the code
/tidy Remove cruft from the diff, behavior-preserving
/changelog Draft a changelog entry from commits since the last release
/scaffold Scaffold a new module/component matching repo conventions
/orchestrate Plan a big goal, create a task ledger, route work by agent/model, and drive it to done
/tasks Create, list, update, or GitHub-sync the task ledger
/solve-loop Drain ready ledger tasks with verify-before-done discipline
/stack Plan, inspect, submit, reconcile, restack, repair, or land dependent pull requests
/stack-review Review every stack layer bottom-up against its immediate parent
/doctor Run the read-only Forge capability and merge-readiness preflight
/policy Evaluate policy, stage effects, issue approvals, authorize, and record outcomes

Hooks

Deterministic guardrails the harness runs on lifecycle events — no model memory required.

Hook Event Effect
session-context SessionStart Injects current branch, ahead/behind, dirty count, and recent commits as context
guard-bash PreToolUse(Bash) Blocks catastrophic commands (rm -rf /, force-push to main, fork bombs)
scan-secrets PreToolUse(Write/Edit) Blocks writing credentials into files
format-file PostToolUse(Write/Edit) Auto-formats edited files with the installed formatter
notify Stop Desktop notification when a turn finishes

Output styles, status line & settings

  • output-styles/ — selectable system-prompt modes: Concise Engineer (answer-first, no preamble) and Mentor (teaches the why as it works). Ship with the plugin; pick one via /config.
  • statusline/ — a status line showing model · dir · git · context% · cost.
  • settings/ — example settings.json (permission allowlist, deny rules for secrets, status line, output style) to pair with the plugin.

Instructions & MCP

  • instructions/ — a CLAUDE.md template library, stack-agnostic engineering principles, and language snippets (TypeScript, Python, Go).
  • mcp/ — example Model Context Protocol server configs with least-privilege guidance.

Evidence — evals & tests

  • evals/ — deterministic prompt-quality and behavior-contract checks, shared cross-host scenarios, and an opt-in LLM-judge eval that scores agents against real tasks.
  • tests/ — pytest cases covering safety hooks, task sync, receipts, durable runtime replay, external effect delivery, doctor, stacks, and conformance. just check runs it all.

Release provenance

Tagged releases publish deterministic Claude, Codex, .agents, and OpenAI skills-only bundles with SHA-256 manifests, SPDX SBOMs, and an offline verifier. Hosted releases can also supply GitHub artifact attestations; local-only releases do not imply hosted provenance. See release provenance for consumer verification and the threat model.

How the pieces fit

flowchart LR
  goal["big goal\n/orchestrate"] --> ledger["task ledger\n/tasks"]
  ledger --> route["route by agent + model\nOpus/Fable · Sonnet · Haiku"]
  route --> loop["solve loop\n/solve-loop"]
  loop --> done["verified done"]
  ledger --> topology{"one PR or stack?"}
  topology --> stack["stack graph\n/stack"]
  stack --> stackreview["incremental review\n/stack-review"]
  stackreview --> done

  plan["plan\narchitect / /plan"] --> impl[implement]
  impl --> review["review\ncode-reviewer / /review"]
  review --> test["test\ntest-engineer / /test"]
  test --> debug["debug\ndebugger / /debug"]
  debug --> ship["ship\n/commit · /pr"]

  guard(["guardrails — always on\nguard-bash · scan-secrets · format-file · notify"])
  guard -. wraps .-> impl
  guard -. wraps .-> review
  guard -. wraps .-> debug
  guard -. wraps .-> ship
Loading

Agents go deep on focused jobs; skills supply the method; commands trigger the loop; hooks keep it safe. For big tasks, the main conversation acts as the conductor so it can spawn specialists in parallel; the ledger keeps the run honest. See docs/usage-patterns.md.

Repository layout

.claude-plugin/        Claude Code marketplace manifest
.agents/plugins/       Codex marketplace manifest
data/                  generated catalog, capability graph, bundles, and workflow metadata
plugins/forge/         the Forge plugin
  .claude-plugin/        plugin manifest
  .codex-plugin/         Codex plugin manifest
  agents/                20 specialist subagents
  skills/                26 progressive-disclosure skills
  commands/              22 slash commands
  hooks/                 5 lifecycle hooks (session-context, guard, secrets, format, notify)
  output-styles/         selectable system-prompt modes
instructions/          CLAUDE.md templates, principles, language guides
mcp/                   example MCP server configs
statusline/            status line script
settings/              example settings.json
evals/                 prompt eval harness, shared scenarios, and result evidence
tests/                 runnable hook, task-ledger, receipt, doctor, stack, and conformance tests
docs/                  getting started, usage, architecture, rationale, CI
scripts/               validation, installation, release, and marketplace checks
.github/               CI, issue/PR templates, CODEOWNERS, dependabot

Documentation

  • Getting started — install options and first steps
  • OSS release contract — contribution checks, research basis, native compatibility evidence, and release acceptance
  • OpenHands SDK — optional native AgentSkills compatibility verification
  • Usage patterns — how the components combine in real workflows
  • Bundles & workflows — focused capability sets and ordered playbooks
  • Quality bar — validation and safety standards for Forge components
  • Competitive audit — what Forge borrows from larger skill libraries
  • Frontier roadmap — research-backed interoperability, identity, evaluation, release evidence, and connected-execution priorities
  • Stacked changes — GitHub-native stacks, provider adapters, safety model, review flow, CI, and recovery
  • GitHub native stacks — remote inspect/import, SHA-guarded reconciliation, divergence classes, mutation authority, and preview fallback
  • Policy plane — action envelopes, profiles, approvals, staged previews, adapter integration, and privacy-safe decision evidence
  • Cross-host conformance — shared scenarios, host adapters, live evidence, result schemas, and release gates
  • Release provenance — deterministic bundles, SBOMs, attestations, offline verification, and threat model
  • Marketplace readiness — honest directory status, publisher surfaces, asset policy, and submission smoke-test matrix
  • OpenAI Agent Plugins compatibility — current universal plugin contract, Forge audit, submission boundary, and staged compatibility plan
  • OpenAI submission packet — reproducible candidate archive evidence, five positive/three negative cases, and the external publication boundary
  • Capability IR — body-aware graph, deterministic host renderer, adapter contract, migration workflow, and current compiler boundary
  • Durable runtime — local SQLite/WAL history, deterministic replay, transactional outbox/inbox effects, generation-fenced heartbeats, lease evidence, checkpointed recovery, human-input waits, signals, MCP Tasks projection, reviewed migrations, offline lineage verification, signed trace/provenance evidence, idempotency, hash-chain verification, and explicit at-least-once boundaries
  • Runtime provenance — signed trace correlation, privacy defaults, offline trust verification, key rotation, retention, and incident response
  • GitHub Agentic Workflows — pinned gh-aw compilation, read-only agents, staged safe outputs, policy evidence, native lock verification, and durable episode correlation
  • Architecture — how the repo is organized and why
  • Design rationale — the decisions and trade-offs behind Forge
  • CI & headless usage — run Forge in pipelines and automated review
  • Evals — the evidence layer · Tests — hook test suite
  • Contributing — add an agent, skill, command, or hook
  • Changelog

Contributing

Contributions are welcome — new agents, skills, commands, and hooks, or improvements to existing ones. Run ./scripts/validate.sh before opening a PR. See CONTRIBUTING.md for the conventions and the quality bar.

License

MIT © Alisina Karimi

About

my own instructions, skills, hooks and agents I have made for my own use in AI assisted development, now made into the plugin "Forge"

Topics

Resources

Contributing

Security policy

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages