Skip to content

About

Smart AI Context Compression - Save 67% on token costs by intelligently compressing repositories for GPT-4, Claude, Gemini

Resources

Stars

2 stars

Watchers

1 watching

Forks

Latest commit

Β 

History

1 Commit

Folders and files

Repository files navigation

ContextForge - Smart AI Context Compression

Save 67% on AI Token Costs

Compress repositories before sending to GPT-4, Claude, Gemini, or Aider.

Installation

pip install contextforge

The Problem 😫

Real-world repositories are TOO BIG for AI context windows:

Repository | Tokens | Cost/Call (GPT-4) Python requests |154,061 |$2.31
Your monorepo? |500K+ |$7.50+ 50%+ of those tokens are noise: test files, documentation, changelogs, binaries.

You're paying AI models to read test_utils.py and HISTORY.md. πŸ’Έ

The Solution ✨

ContextForge is a CLI tool that intelligently compresses repositories into optimized context packages.

One Command 🎯

contextforge forge ./my-repo --budget 50000 --model gpt4 -o optimized.md

What It Does πŸ› οΈ

1.Scans your repository (file types, sizes, categories) 2.Analyzes Git history to find important files (like yek) 3.Prioritizes using 6 weighted signals (commit frequency, recency, etc.) 4.Filters noise (auto-excludes tests, binaries, node_modules) 5.Truncates smartly (keeps function signatures, cuts implementations) 6.Exports optimized packages tailored for specific models

Real Results πŸ“Š Tested on the Python requests library (154K tokens):

Metric |Before |After |Savings| Tokens |154,061 |50,406 |67% ↓| Cost/Call|$2.31 | $0.77 | $1.54 saved| Files |112(all)| 54 (signal only)| 58 removed| Speed | - |0.9 seconds|⚑|

Before vs After Example

Instead of this (7,472 tokens):

class Response: """Response object.""" def raise_for_status(self): if self.status_code >= 400: raise HTTPError(...) @property def content(self) -> bytes: if self._content is False: raise ... def iter_content(self, chunk_size=1, decode_unicode=False): # ... 200 lines of implementation

You get this (~2,500 tokens):

class Response: """Response object.""" def raise_for_status(self): ... # [Implementation omitted: ~723 tokens] @property def content(self) -> bytes: ... # [Implementation omitted: ~136 tokens] def iter_content(self, chunk_size=1, decode_unicode=False): ... # [~532 tokens]

=== Signature mode: ~4,974 tokens omitted ===

Same understanding. 67% fewer tokens. πŸŽ‰

Features ⚑

βœ… Multi-model support: GPT-4 Markdown, Claude XML, Gemini, Aider formats βœ… Git-intelligent ranking: Ranks files by commit frequency, recency, branch relevance βœ… Visual analytics: Pie charts showing token distribution (matplotlib) βœ… Configurable: YAML-based config with include/exclude patterns βœ… Dry-run mode: Preview what would be included before committing βœ… Caching: Token counts cached for speed on repeated runs βœ… Smart truncation: AST-aware signature extraction (tree-sitter) βœ… 48 tests passing, clean architecture, fully typed Python

Installation πŸ’Ώ

From PyPI (coming soon)pip install contextforge# Or from sourcegit clone https://github.com/YOUR_USERNAME/contextforge.gitcd contextforgepip install -e .

Quick Start πŸƒ

1. Initialize configuration in your project

cd your-project contextforge init

2. Analyze token distribution

contextforge analyze . --visualize

3. Generate optimized context package

contextforge forge . --budget 100000 --model gpt4 -o context.md

4. Use with your favorite AI tool!

cat context.md | claude

CLI Reference πŸ“–

analyze - Explore repository

contextforge analyze ./my-repo [OPTIONS]Options: --top-n N Show top N files (default: 20) --by-category Group by file type --git-ranking Show Git-based importance --visualize Generate pie chart (PNG/SVG) --save-report Export analysis as JSON --open Open chart in browser

forge - Create optimized package

contextforge forge ./my-repo [OPTIONS]

Options: --budget, -b N Token budget (default: 100000) --model, -m MODEL Target: gpt4, claude, gemini, aider (default: gpt4) --format, -f FMT Output: markdown, xml, json, aider, gemini --output, -o FILE Output file path --include-tests Include test files --include-docs Include full documentation --strategy STRAT Truncation: signature, headline, aggressive, none --dry-run Preview without generating output --verbose, -v Detailed output

init - Create config

contextforge init [PATH] Creates .contextforge/config.yaml with sensible defaults.

cache - Manage cache

contextforge cache --stats # Show cache info contextforge cache --clear # Clear all caches

Configuration βš™οΈ

Create .contextforge/config.yaml in your project root: defaults: token_budget: 100000 target_model: gpt4 output_format: markdown

scoring: git_commit_frequency: 0.30 git_recency: 0.20 git_author_match: 0.10 git_branch_relevance: 0.10 file_type_heuristics: 0.15 dependency_centrality: 0.15

filters: exclude_patterns: - "/node_modules/" - "/pycache/" - "/*.test.js" - "/test_*.py"

include_patterns: - "README.md" - "package.json" - "pyproject.toml"

truncation: strategy: signature # signature, headline, aggressive, none

Supported Languages 🌐

Smart truncation (signature extraction) supported for:

Python βœ… | JavaScript/TypeScript βœ… | Go βœ… | Java βœ… Rust (experimental) | Ruby (experimental) Plus plain text/markdown compression.

How It Works πŸ”¬

Prioritization Algorithm

ContextForge uses multi-signal weighted scoring: Priority = 0.30 Γ— commit_frequency

  • 0.20 Γ— recency_score
  • 0.10 Γ— author_match
  • 0.10 Γ— branch_relevance
  • 0.15 Γ— file_type_score
  • 0.15 Γ— dependency_centrality

All signals are normalized 0-1 and time-decayed (half-life: 90 days).

Truncation Strategies

Strategy Description Best For
signature Keep function/class signatures, omit bodies Source code
headline Keep headings + first paragraph Documentation
aggressive Headers only, maximum compression Large docs
none Include everything Small files

Tech Stack πŸ› οΈ

  • Language: Python 3.10+
  • Token Counting: tiktoken (OpenAI's tokenizer)
  • Git Analysis: gitpython
  • CLI Framework: Typer
  • Terminal UI: Rich
  • Visualization: Matplotlib + Plotly
  • AST Parsing: Tree-sitter

Monetization πŸ’°

Tier Price Features Core (CLI) |Free | All basic features Pro |$12/month| IDE plugins (VS Code, JetBrains), team sync Enterprise | Custom | SSO, audit logs, self-hosted, SLA

Roadmap πŸ—ΊοΈ

βœ…Core CLI functionality βœ…Multi-model export formats βœ…Git-based prioritization βœ…Visualization VS Code extension JetBrains plugin GitHub App (PR integration) Web dashboard ML-based importance scoring Team sharing/collaboration

Contributing 🀝

Contributions welcome!

1.Fork the repo 2.Create feature branch (git checkout -b feature/amazing) 3.Commit changes (git commit -m 'Add amazing feature') 4.Push to branch (git push origin feature/amazing) 5.Open Pull Request

Testing βœ…

Run all testspytest tests/ -v# Current status: 48 tests passing βœ…

License πŸ“„

MIT

Acknowledgments πŸ™

+Inspired by yek for Git-based file importance +Built with OpenAI's tiktoken +Powered by tree-sitter for AST parsing

Built with frustration and hope. πŸš€

     If ContextForge saves you money, ⭐ Star the repo!

About

Smart AI Context Compression - Save 67% on token costs by intelligently compressing repositories for GPT-4, Claude, Gemini

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages