AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Code Context Engine

mcp-elara-labs-code-context-engine · by elara-labs

Save 94% on AI coding tokens. Index your codebase, agents search instead of reading files. Works with Claude Code, Codex, Copilot, Cursor, Gemini CLI. Local MCP server, free, open source.

No reviews yet
0 installs
5 views
0.0% view→install

Install

$ agentstack add mcp-elara-labs-code-context-engine

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-elara-labs-code-context-engine)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Code Context Engine? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Code Context Engine

Index your codebase. AI searches instead of re-reading files.94% token savings, reproducibly benchmarked.

Website · Docs · Why CCE? · Benchmark · GitHub

Python 3.11+ · macOS · Linux · Windows

           

One command. Auto-detects your editor. Zero cloud, zero config.


Use cases

| | Use case | How CCE helps | |---|---|---| | 💰 | Reduce Claude Code costs | 94% fewer input tokens per session | | 🔒 | Keep code private | Everything local, no cloud indexing | | 🔄 | Multi-editor teams | One index across Claude Code, Cursor, VS Code, Gemini CLI | | 🧠 | Cross-session memory | Decisions and context survive restarts | | | Faster responses | Less context = faster Claude replies | | 📊 | Track actual savings | Dollar amounts, not estimates |


Quick start

One command. 30 seconds.

uvx --from "code-context-engine[local]" cce init    # install + index + configure, one shot

Or if you prefer a persistent install:

uv tool install "code-context-engine[local]"    # or: pipx install "code-context-engine[local]"
cd /path/to/your/project
cce init

Restart your editor. Done. Every question now hits the index instead of re-reading files.

> Already have Ollama? Skip [local] and use uv tool install code-context-engine instead. CCE auto-detects Ollama at localhost:11434 and uses nomic-embed-text.

System requirements

Python 3.11+ and a C compiler (for tree-sitter grammars).

| Platform | Setup | |----------|-------| | macOS | xcode-select --install | | Ubuntu/Debian | sudo apt install build-essential cmake | | Fedora/RHEL | sudo dnf install gcc gcc-c++ cmake | | Windows | Visual Studio Build Tools (C++ workload) + CMake |

Tested on macOS, Linux, Windows with Python 3.11/3.12/3.13.

cce init auto-detects your editor and writes the right config. To target a specific agent, use --agent claude, --agent codex, --agent copilot, or --agent all.

| Editor | Config written | Instructions | |--------|---------------|--------------| | Claude Code | .mcp.json | CLAUDE.md | | VS Code / Copilot | .vscode/mcp.json | .github/copilot-instructions.md | | Cursor | .cursor/mcp.json | .cursorrules | | Gemini CLI | .gemini/settings.json | GEMINI.md | | OpenAI Codex | ~/.codex/config.toml (user-global, per-project section) | AGENTS.md | | OpenCode | opencode.json | | | Tabnine | .tabnine/agent/settings.json | TABNINE.md |

Multiple editors in the same project? All get configured in one command.

Codex note: Codex CLI reads MCP servers from ~/.codex/config.toml only — it has no per-project config. cce init adds one [mcp_servers.cce--] section per project so multiple projects coexist; cce uninstall removes only the section for the current project.

  my-project · 38 queries · last query 5m ago

  ⛁ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶  88% tokens saved

  Input savings   1.9M  tokens   $27.78
  Output savings  4.8k  tokens   $0.36
  ──────────────────────────────────────────
  Total saved   1.9M  tokens   $28.15

  Breakdown:
    retrieval              84%  ▰▰▰▰▰▰▰▰▰▰    1.8M   $26.76 · 12 calls
    chunk compression       3%  ▰▱▱▱▱▱▱▱▱▱   68.5k    $1.03 · 12 calls
    output compression*    
Content-Hash Embedding Cache

SHA-256 fingerprint per chunk, salted with model name. Re-index skips unchanged code. Binary float32 storage (10x smaller than JSON). Typical re-index: 96% cache hit, under 1 second.

sqlite-vec: 2 MB instead of 217 MB

Replaced LanceDB with sqlite-vec. Same cosine-distance quality, 99% smaller install. WAL mode + PRAGMA NORMAL for 80% write speedup. Vectors, FTS5, code graph, and compression cache all in three SQLite files.

Deterministic Grammar Compression

Memory entries compressed without LLM calls. Drops articles, fillers, pronouns. Three levels (lite/full/ultra, 20-60% savings). Code, paths, URLs preserved byte-for-byte. Same input always yields same output.

Fail-Closed Hook Design

5 Claude Code lifecycle hooks capture session context. Every hook runs `curl ... || true`, so a crashed server never blocks the user. SessionStart injects bootstrap context; others capture silently.

Multi-Provider Pricing

Dollar estimates in `cce savings` support 15+ models across Anthropic, OpenAI, and Google. Static pricing ships with CCE, live Anthropic pricing is fetched and cached 7 days. Configure `pricing.model` (e.g. `gpt-4o`, `gemini-2.5-pro`, `sonnet`) or override with `pricing.input` / `pricing.output` for custom rates.

Append-Only Savings Ledger

7 buckets track every token saved: retrieval, chunk compression, output compression, memory recall, grammar, turn summarization, progressive disclosure. Survives restarts. Powers CLI and dashboard analytics.

---

## CLI at a glance

```bash
cce init                    # Index + install hooks + register MCP
cce                         # Status banner
cce savings                 # Token savings with dollar estimates
cce savings --all           # All projects
cce dashboard               # Web dashboard with live charts
cce search "auth flow"      # Test a query
cce status                  # Index health + config
cce services                # Ollama + dashboard + MCP status
cce commands add-rule '...' # Project rules for Claude
cce uninstall               # Clean removal of all CCE artifacts

Run cce list for the full command reference.


Configuration

Zero-config by default. Override what you need in ~/.cce/config.yaml or .context-engine.yaml:

compression:
  level: standard          # minimal | standard | full
  output: standard         # off | lite | standard | max
  ollama_url: http://localhost:11434   # point at a remote Ollama if desired

retrieval:
  top_k: 20
  confidence_threshold: 0.5

pricing:
  model: opus              # opus | sonnet | haiku | gpt-4o | gemini-2.5-pro | ...
  # input: 15.0            # override $/1M input tokens
  # output: 75.0           # override $/1M output tokens

Remote Ollama: If you run Ollama on another machine in your network, set compression.ollama_url (e.g. http://nas.local:11434) or export CCE_OLLAMA_URL — the env var wins. CCE probes the endpoint and falls back to truncation-only compression when it's unreachable, so a flaky link won't break indexing.


Output Compression

CCE also compresses Claude's responses (same concept as Caveman):

| Level | Style | Savings | |-------|-------|---------| | off | Full output | 0% | | lite | No filler or hedging | ~30% | | standard | Fragments, drop articles | ~65% | | max | Telegraphic | ~75% |

Tell Claude: "switch to max compression" or "turn off compression". Code blocks and commands are never compressed.


Disk Footprint

| Component | Size | |-----------|------| | Core install (Ollama backend) | ~17 MB | | With [local] extra (fastembed + ONNX) | ~189 MB | | Embedding model (one-time download) | ~60 MB (fastembed) or managed by Ollama | | Index per project (small/medium/large) | 5-60 MB |

No GPU required. With Ollama, embeddings are handled by the Ollama server. With the [local] extra, the embedding model runs on CPU via ONNX Runtime.


Supported Languages

AST-aware chunking (tree-sitter parsed, 10 extensions):

| Language | Extensions | |----------|-----------| | Python | .py | | JavaScript | .js, .jsx | | TypeScript | .ts, .tsx | | PHP | .php | | Go | .go | | Rust | .rs | | Java | .java |

Language-aware fallback chunking (40+ extensions):

| Category | Languages | |----------|-----------| | Web | HTML, CSS, SCSS, LESS, Vue, Svelte | | Systems | C, C++, C#, Zig, Nim | | Mobile | Swift, Kotlin, Dart | | Functional | Haskell, Scala, Clojure, Elixir, Erlang, F# | | Scripting | Ruby, Perl, Lua, R, Bash/Zsh | | Data/Config | JSON, YAML, TOML, XML, SQL, GraphQL, Protobuf | | DevOps | Terraform, HCL, Dockerfile | | Docs | Markdown |

All other text files are chunked by line range. Binary files are skipped.


Documentation

| Page | Content | |------|---------| | How Much Are You Spending on AI Coding Tokens? | The math on input vs output tokens | | What is CCE? (Complete Guide) | Setup, tools, how it works, FAQ | | How to Save Claude Code Tokens | Cost breakdown and savings guide | | Benchmark Deep Dive | Full FastAPI benchmark methodology | | Comparison with Alternatives | CCE vs Cursor, Aider, Continue, Greptile | | Examples | Real conversations with Claude | | How It Works | Full 9-stage pipeline | | CLI Reference | Every command with output | | Configuration | All config options |


FAQ

Does CCE affect response quality?

No. Quality stays the same or slightly improves.

CCE replaces "dump the entire file" with "search for the relevant function." The model still gets the code it needs (0.90 Recall@10 in benchmarks). Less irrelevant context means less noise competing for attention, which can improve the model's focus on your actual question.

How does output token savings work?

CCE writes output compression rules directly into your agent's instruction files (CLAUDE.md, AGENTS.md, .cursorrules, etc.) during cce init. These rules apply to the entire session, not just CCE tool responses, so every reply from the agent follows them.

Set the level in ~/.cce/config.yaml or .context-engine.yaml:

compression:
  output: max       # off | lite | standard | max

Then re-run cce init to update instruction files. Or change at runtime:

set_output_level output_level=max

| Level | Savings | What it does | |-------|---------|--------------| | off | 0% | No compression | | lite | ~25% | Removes filler/hedging/pleasantries + diff-only for code changes | | standard | ~70% | Drops articles, fragments, short synonyms + diff-only for code | | max | ~80% | Telegraphic style + diff-only for code |

Default is standard. All levels include code output rules that tell the model to show only changed lines (not full file rewrites), which is where most output tokens go in coding sessions. The max level produces very terse prose (similar to "caveman mode"). Code blocks, paths, and commands are never compressed regardless of level.

Where do the savings come from?

Most savings are input tokens (what goes into the model):

| Layer | Type | Typical savings | |-------|------|-----------------| | Retrieval | Input | 94% (full files → relevant chunks) | | Chunk compression | Input | 89% (chunks → signatures) | | Grammar compression | Input | 13% (article/filler removal) | | Turn summarization | Input | varies (session history) | | Progressive disclosure | Input | varies (tool payloads) | | Output compression | Output | 25-80% (depends on level) |

Output tokens cost 5x more per token (e.g. Opus: $15/1M input vs $75/1M output), so even a small output reduction has outsized cost impact.


Roadmap

  • [x] Multi-repo benchmarks (FastAPI, chi, fiber)
  • [ ] More benchmarks (Django, Express)
  • [ ] Tree-sitter support for C, C++, Ruby, Swift, Kotlin
  • [ ] Docker support for remote mode

See [CHANGELOG.md](CHANGELOG.md) for shipped features.


Contributing

Contributions welcome. See https://github.com/elara-labs/code-context-engine/blob/main/CONTRIBUTING.md for setup.


License

MIT. See [LICENSE](LICENSE).

Authors

Acknowledgments

Claude Code · MCP · sqlite-vec · Tree-sitter · fastembed · Ollama


If CCE saves you tokens, give it a star.

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.