Skip to content

Repository files navigation

Repo Context Engine

Give coding agents the smallest useful slice of a codebase instead of dumping the whole repository into a prompt.

CI Python 3.11+ License: MIT

Repo Context Engine scans a repository, extracts symbols and imports, ranks files against a natural-language task, and emits a compact Markdown context pack for Claude Code, Codex, Cursor, Copilot, or any file-aware coding agent.

No embeddings. No external service. No source code leaves your machine.

The problem

Coding agents fail when they see too little context—but also when they see too much irrelevant context. Large prompt dumps increase cost, hide important constraints, and make edits less precise. This project treats context selection as a deterministic engineering step.

Features

  • Multi-language repository scan for Python, JS/TS, Go, Rust, Java, Kotlin, Ruby, PHP, and common config formats
  • AST-based Python symbol and import extraction
  • Lightweight structural extraction for other languages
  • Task-aware ranking across paths, symbols, content, and dependency centrality
  • Character-budgeted context packs with file excerpts
  • Default exclusion of dependencies, virtual environments, builds, and VCS data
  • JSON repository maps for downstream automation
  • Zero runtime dependencies and no network access

Quick start

git clone https://github.com/zchstime/repo-context-engine.git
cd repo-context-engine
python -m pip install -e .

Find the files most relevant to a task:

repo-context --repo /path/to/project query "fix session validation and add tests"

Build a context pack that fits a fixed character budget:

repo-context --repo /path/to/project pack \
  "fix session validation and add tests" \
  --budget 18000 \
  --output context.md

Then attach context.md to your coding agent.

How ranking works

natural-language task
        ↓ tokenize
path matches + symbol matches + content matches + dependency centrality
        ↓ score
ranked files
        ↓ budget
context pack with excerpts and verification instructions

The score is intentionally inspectable. Every result explains whether it matched the file path, declared symbols, previewed content, or dependency centrality. You can understand why a file was selected before trusting the context pack.

Example output

  9.40  src/auth/session.py  path: session; symbols: validation; dependency centrality: 6
  4.25  tests/test_session.py content: session, validation
  1.50  src/api/middleware.py dependency centrality: 10

Use as a library

from repo_context_engine import ContextEngine

engine = ContextEngine("/path/to/repo")
for item in engine.rank("trace payment webhook retries"):
    print(item.score, item.record.path, item.reasons)

prompt_context = engine.context_pack("trace payment webhook retries")

Privacy

The engine is offline and deterministic. Generated context packs contain source excerpts, so review them before sharing outside the trust boundary of the repository.

Test

PYTHONPATH=src python -m unittest discover -s tests -v

Roadmap

  • Git diff-aware ranking
  • Language Server Protocol symbol adapter
  • Import graph visualization
  • Tokenizer-aware budgets for major model families
  • MCP server mode

If you believe better context beats bigger prompts, Star the repository and open an issue with the language or ranking signal you need next.

License

MIT

About

Build compact, ranked repository context packs for Codex, Claude Code, Cursor, and other coding agents.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages