RepoPilot is a repository-aware code understanding and automatic repair system for real software repositories.
It focuses on the part that decides whether a coding agent is actually useful in practice: issue localization, structure-aware retrieval, patch generation, sandbox verification, adversarial review, repair memory, and benchmark-driven iteration.
- repository-aware retrieval with lexical, dense, symbol, call, import, and graph signals
- Tree-sitter + code-graph localization for cross-file bug analysis
- multi-candidate patch generation with AST-aware and semantic rewrite strategies
- sandbox-first patch validation with
git apply --check, test execution, and repair replay - adversarial patch review, policy-aware ranking, and badcase analysis
- SWE-bench-style evaluation pipeline with public-dataset parsing and harness-compatible prediction export
RepoPilot is designed as a strong large-model application project rather than a toy prompt demo:
- Problem: repository bug fixing is not just code generation; it also requires localization, impact analysis, validation, and failure recovery.
- Method: combine graph-aware retrieval, structured planning, patch portfolio ranking, tool verification, and memory replay into one repair loop.
- Outcome: build a reusable workflow for repository issue -> candidate patch -> validation -> repair -> benchmark export.
RepoPilot turns a repository issue into an end-to-end repair workflow:
- retrieve relevant code with lexical, embedding, symbol, call, import, and graph signals
- localize likely files, functions, and propagation paths
- generate multiple patch candidates
- validate unified diffs with
git apply --check - run candidate patches in isolated sandboxes
- score and rank candidates with graph, policy, and adversarial signals
- replay repair context across rounds
- persist repair memory, user preferences, and thread-level context
- expose compressed context and a decision tree for operator inspection
Most repo agents stop at "LLM + retrieval + diff text."
RepoPilot pushes deeper on the parts that decide whether a coding agent is actually useful:
- semantic AST rewrite planning for Python targets
- graph-aware file ranking and patch selection
- parallel patch tournament execution in sandbox copies
- adversarial review that challenges weak or superficial patches
- counterexample-driven regeneration when the first patch is not robust
- persistent repair memory plus repo-level preference learning
- thread-aware memory hooks for multi-round workflows
- compressed context packets for large-repo and long-run stability
- decision-tree output for explainability
These are project-internal validation signals used to guide iteration, not external leaderboard claims:
- local test suite covers workflow, public SWE-bench parsing, graph-aware helpers, and patch selection logic
- official/public evaluation pathway supports local
json,jsonl, andparquetexports plus harness-compatibleall_preds.jsonl - benchmark pipeline supports 12 official repository mirrors and 500-case batch execution flow
- public-eval summaries track
strict_pass@1,env_adjusted_pass@1,average_overall, path recall, repair rounds, and reproducibility tags
- Tree-sitter code graph for Python, JavaScript, TypeScript, and TSX
- dense + sparse hybrid retrieval with rerank
- symbol / call / import / graph propagation signals
- impacted-file prediction from local code graph structure
- structured root-cause hypothesis
- implementation blueprint
- execution-unit decomposition
- acceptance-bundle generation
- rule-based and LLM-assisted patch generation
- AST-aware function-scoped patch candidates
- semantic AST rewrite-plan candidates
- execution-unit-driven multi-candidate patch generation
- patch portfolio ranking
- graph-priority bonus
- policy priors learned from past repairs
- adversarial penalty for weak patch archetypes
- dynamic candidate budget for sandbox evaluation
- failure parsing from
pytest,git apply, and CI feedback - repair-context replay across rounds
- counterexample-driven patch regeneration
- sandbox-first validation before worktree mutation
- long-term repair memory in
.repopilot/memory.sqlite3 - repo-level preference persistence
- thread-level profile persistence
- user-style inference from issue text and successful repair history
- compressed context packet
- decision tree with Mermaid-ready structure
- persistent trace and checkpoint artifacts
- multi-agent repo diagnosis and repair workflow
- real OpenAI-compatible LLM integration
- graph-based retrieval and rerank
- semantic AST rewrite candidate generation
- parallel patch sandbox execution
- adversarial patch review
- counterexample-driven regeneration
- worktree patch apply with guardrails
- GitHub PR / CI / comment integration
- interactive CLI mode with guided prompts
- FastAPI dashboard for operator workflows
- SWE-bench-style evaluation runner
- public-eval JSON / Markdown export
Recent verified local state for this repository:
22 passedtest suite- interactive CLI available through
run_repo_pilot.py --interactive - compressed context and decision tree emitted in real runs
- repo-level preference persistence loaded in real smoke runs
- semantic AST rewrite candidates integrated into the patch pipeline
- graph-aware and policy-aware patch selection active
Representative real smoke signals observed during recent runs:
- real repo smoke overall around
0.838 - semantic-AST-oriented smoke overall around
0.863 - thread memory storage verified
- repo preference loading verified
These are project-internal validation signals, not external leaderboard claims.
RepoPilot supports two benchmark layers:
Used to validate repeatable end-to-end agent behavior on controlled tasks:
- retrieval
- graph localization
- repair context replay
- memory
- dashboard / operator path
- benchmark reporting
RepoPilot can load official/public instances from local json, jsonl, or parquet, run the agent, and export harness-compatible all_preds.jsonl.
Important: local approximate runs are clearly labeled as approximate when the repository snapshot is not the exact official base commit.
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e .[dev]If you want local parquet-based public-dataset parsing, install the benchmark extras as well:
.\.venv\Scripts\python.exe -m pip install -e .[dev,benchmark]Copy .env.example to .env.
DashScope / Qwen example:
DASHSCOPE_API_KEY=your-key
QWEN_MODEL=qwen-plus
QWEN_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1Embedding example:
EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
EMBEDDING_API_KEY=your-key
EMBEDDING_MODEL=text-embedding-v4Generic OpenAI-compatible example:
LLM_BASE_URL=https://api.example.com/v1
LLM_API_KEY=your-key
LLM_MODEL=gpt-4o-miniGitHub integration:
GITHUB_TOKEN=github_pat_or_classic_tokenFor guided usage without memorizing flags:
.\.venv\Scripts\python.exe run_repo_pilot.py --interactiveMain production path:
.\.venv\Scripts\python.exe run_repo_pilot.py --repo . --issue "API returns an unstable JSON schema; locate the endpoint and stabilize the contract." --run-tests --apply-sandbox --save-run --use-llm --graphApply a validated patch back to the repository:
.\.venv\Scripts\python.exe run_repo_pilot.py --repo . --issue "Fix a failing package import at startup." --run-tests --apply-sandbox --apply-worktree --save-run --use-llm --graphCreate a PR and inspect CI:
.\.venv\Scripts\python.exe run_repo_pilot.py --repo . --issue "GitHub integration smoke" --run-tests --apply-sandbox --apply-worktree --create-pr --poll-ci --comment-body "RepoPilot validation passed." --save-run --use-llm --graphStart the local dashboard:
.\.venv\Scripts\python.exe -m uvicorn app.api.server:app --reload --port 8000Open:
http://127.0.0.1:8000/repo-pilot/ui
The dashboard exposes:
- run configuration
- selected patch
- failure signals
- compressed context
- decision tree
- full JSON payload
RepoPilot persists:
- repair memories
- selected patch evidence
- repair journals
- repo-level preference profiles
- thread-level profiles
Primary local stores:
.repopilot/memory.sqlite3
.repopilot/traces.sqlite3
.repopilot/approvals.sqlite3
Local benchmark:
.\.venv\Scripts\python.exe run_benchmark.py --use-llm --run-tests --apply-sandbox --save-runSWE-style evaluation:
.\.venv\Scripts\python.exe run_swe_bench_style.py --cases benchmarks\swe_style_cases.json --work-dir .repopilot\swe_runs --max-cases 11 --write-json .repopilot\reports\public_eval_latest.json --write-markdown .repopilot\reports\public_eval_latest.mdOfficial/public instance pathway:
.\.venv\Scripts\python.exe run_swe_bench_style.py --dataset-path path\to\official_swe_bench.jsonl --dataset-name princeton-nlp/SWE-bench_Verified --dataset-split test --instance-id django__django-16527 --write-preds .repopilot\reports\all_preds.jsonl --write-json .repopilot\reports\official_eval_summary.jsonOfficial dataset download/export example:
.\.venv\Scripts\python.exe -c "from datasets import load_dataset; import json, pathlib; ds=load_dataset('princeton-nlp/SWE-bench_Verified', split='test'); target=pathlib.Path('.repopilot/official_datasets/swe_bench_verified_test_500.jsonl'); target.parent.mkdir(parents=True, exist_ok=True); f=target.open('w', encoding='utf-8', newline=''); [f.write(json.dumps({k: ds[i][k] for k in ds.column_names}, ensure_ascii=False) + chr(10)) for i in range(len(ds))]; f.close(); print(target)"Network note:
- downloading
SWE-bench_Verifiedfrom Hugging Face requires outbound access to Hugging Face - running official/public cases against remote repositories also requires outbound access to GitHub for
git clone - if GitHub is blocked, RepoPilot can still export predictions and parse official dataset rows locally, but end-to-end public-case execution will stop at repository preparation
RepoPilot is already a strong engineering agent prototype, but it is still short of the best frontier coding agents in a few places:
- semantic AST rewrite is still candidate-oriented, not yet a full executable rewrite operator
- thread memory exists and persists, but long multi-session planning can still go deeper
- patch synthesis quality is stronger than before, but still not at the level of the best closed commercial agents
- official SWE-bench claims still depend on exact base-commit reproduction
These are active design targets, not hidden weaknesses.
RepoPilot is being developed as:
- a serious GitHub-visible coding-agent project
- a strong interview project for AI agent / coding-agent roles
- a testbed for graph retrieval, semantic patching, and continual repair learning