Compress repositories before sending to GPT-4, Claude, Gemini, or Aider.
pip install contextforgeReal-world repositories are TOO BIG for AI context windows:
Repository | Tokens | Cost/Call (GPT-4)
Python requests |154,061 |$2.31
Your monorepo? |500K+ |$7.50+
50%+ of those tokens are noise: test files, documentation, changelogs, binaries.
You're paying AI models to read test_utils.py and HISTORY.md. πΈ
ContextForge is a CLI tool that intelligently compresses repositories into optimized context packages.
contextforge forge ./my-repo --budget 50000 --model gpt4 -o optimized.md
1.Scans your repository (file types, sizes, categories) 2.Analyzes Git history to find important files (like yek) 3.Prioritizes using 6 weighted signals (commit frequency, recency, etc.) 4.Filters noise (auto-excludes tests, binaries, node_modules) 5.Truncates smartly (keeps function signatures, cuts implementations) 6.Exports optimized packages tailored for specific models
Real Results π Tested on the Python requests library (154K tokens):
Metric |Before |After |Savings| Tokens |154,061 |50,406 |67% β| Cost/Call|$2.31 | $0.77 | $1.54 saved| Files |112(all)| 54 (signal only)| 58 removed| Speed | - |0.9 seconds|β‘|
class Response: """Response object.""" def raise_for_status(self): if self.status_code >= 400: raise HTTPError(...) @property def content(self) -> bytes: if self._content is False: raise ... def iter_content(self, chunk_size=1, decode_unicode=False): # ... 200 lines of implementation
class Response: """Response object.""" def raise_for_status(self): ... # [Implementation omitted: ~723 tokens] @property def content(self) -> bytes: ... # [Implementation omitted: ~136 tokens] def iter_content(self, chunk_size=1, decode_unicode=False): ... # [~532 tokens]
β Multi-model support: GPT-4 Markdown, Claude XML, Gemini, Aider formats β Git-intelligent ranking: Ranks files by commit frequency, recency, branch relevance β Visual analytics: Pie charts showing token distribution (matplotlib) β Configurable: YAML-based config with include/exclude patterns β Dry-run mode: Preview what would be included before committing β Caching: Token counts cached for speed on repeated runs β Smart truncation: AST-aware signature extraction (tree-sitter) β 48 tests passing, clean architecture, fully typed Python
From PyPI (coming soon)pip install contextforge# Or from sourcegit clone https://github.com/YOUR_USERNAME/contextforge.gitcd contextforgepip install -e .
cd your-project contextforge init
contextforge analyze . --visualize
contextforge forge . --budget 100000 --model gpt4 -o context.md
cat context.md | claude
contextforge analyze ./my-repo [OPTIONS]Options: --top-n N Show top N files (default: 20) --by-category Group by file type --git-ranking Show Git-based importance --visualize Generate pie chart (PNG/SVG) --save-report Export analysis as JSON --open Open chart in browser
contextforge forge ./my-repo [OPTIONS]
Options: --budget, -b N Token budget (default: 100000) --model, -m MODEL Target: gpt4, claude, gemini, aider (default: gpt4) --format, -f FMT Output: markdown, xml, json, aider, gemini --output, -o FILE Output file path --include-tests Include test files --include-docs Include full documentation --strategy STRAT Truncation: signature, headline, aggressive, none --dry-run Preview without generating output --verbose, -v Detailed output
contextforge init [PATH] Creates .contextforge/config.yaml with sensible defaults.
contextforge cache --stats # Show cache info contextforge cache --clear # Clear all caches
Create .contextforge/config.yaml in your project root: defaults: token_budget: 100000 target_model: gpt4 output_format: markdown
scoring: git_commit_frequency: 0.30 git_recency: 0.20 git_author_match: 0.10 git_branch_relevance: 0.10 file_type_heuristics: 0.15 dependency_centrality: 0.15
filters: exclude_patterns: - "/node_modules/" - "/pycache/" - "/*.test.js" - "/test_*.py"
include_patterns: - "README.md" - "package.json" - "pyproject.toml"
truncation: strategy: signature # signature, headline, aggressive, none
Smart truncation (signature extraction) supported for:
Python β | JavaScript/TypeScript β | Go β | Java β Rust (experimental) | Ruby (experimental) Plus plain text/markdown compression.
ContextForge uses multi-signal weighted scoring: Priority = 0.30 Γ commit_frequency
- 0.20 Γ recency_score
- 0.10 Γ author_match
- 0.10 Γ branch_relevance
- 0.15 Γ file_type_score
- 0.15 Γ dependency_centrality
All signals are normalized 0-1 and time-decayed (half-life: 90 days).
| Strategy | Description | Best For |
|---|---|---|
signature |
Keep function/class signatures, omit bodies | Source code |
headline |
Keep headings + first paragraph | Documentation |
aggressive |
Headers only, maximum compression | Large docs |
none |
Include everything | Small files |
- Language: Python 3.10+
- Token Counting:
tiktoken(OpenAI's tokenizer) - Git Analysis:
gitpython - CLI Framework:
Typer - Terminal UI:
Rich - Visualization: Matplotlib + Plotly
- AST Parsing: Tree-sitter
Tier Price Features Core (CLI) |Free | All basic features Pro |$12/month| IDE plugins (VS Code, JetBrains), team sync Enterprise | Custom | SSO, audit logs, self-hosted, SLA
β Core CLI functionality β Multi-model export formats β Git-based prioritization β Visualization VS Code extension JetBrains plugin GitHub App (PR integration) Web dashboard ML-based importance scoring Team sharing/collaboration
Contributions welcome!
1.Fork the repo 2.Create feature branch (git checkout -b feature/amazing) 3.Commit changes (git commit -m 'Add amazing feature') 4.Push to branch (git push origin feature/amazing) 5.Open Pull Request
MIT
+Inspired by yek for Git-based file importance +Built with OpenAI's tiktoken +Powered by tree-sitter for AST parsing
If ContextForge saves you money, β Star the repo!