AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Bgts Context Engine

mcp-bgts-ai-org-bgts-context-engine · by bgts-ai-org

Deterministic code-graph context engine for AI coding agents - PostgreSQL + Apache AGE + pgvector, MCP & REST, no LLM in the loop.

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add mcp-bgts-ai-org-bgts-context-engine

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-bgts-ai-org-bgts-context-engine)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
yesterday

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Bgts Context Engine? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Deterministic code-graph context for AI coding agents.

Ask "why does the login timeout fire on the meeting webhook?" and get the eight symbols that actually answer it — ranked, budgeted, and reproducible.

[](https://pypi.org/project/bgts-context-engine/) [](https://pypi.org/project/bgts-context-engine/) [](https://github.com/bgts-ai-org/bgts-context-engine/actions/workflows/ci.yml) [](LICENSE) [](docs/mcp.md) [](https://github.com/bgts-ai-org/bgts-context-engine/stargazers)

[Quick start](#quick-start) · [Use it from your agent](#use-it-from-your-agent) · [How it works](#how-it-works) · [Supported models](#supported-models) · [Documentation](#documentation) · [Türkçe](README.tr.md)


Why this exists

An agent working on an unfamiliar repository has to decide what to read before it can decide what to change. The usual answer is embedding search over chunked files. It is cheap to build and wrong in a specific way: it returns text that reads like the question rather than code that participates in the behaviour. Ask about a login timeout and you get the five files that mention timeouts, not the one function that sets it and the three callers that break when you change it.

That information is structural, and it has an exact answer. handleLogin calls refreshSession, which reads SESSION_TTL, which is written in exactly one place. That is a graph walk.

BGTS Context Engine indexes your repositories into that graph — symbols, calls, references, type hierarchies, HTTP routes, cross-language bridges — and answers questions by walking it. Embeddings are used in one place only: finding entry points when the task text names nothing recognisable. They never affect ranking.

The same task text, against the same commit, returns the same context pack. No model in the retrieval path, no clock, no randomness. When an agent makes a bad change you can replay exactly what it was told, find the stage that surfaced the wrong symbol, and fix that stage.

The engine is published so people can run it. Organisations that want the same thing inside their own perimeter — help with indexing, deployment, scoring tuned to their repositories, or the agent stack around it — can engage BGTS for consulting. Write to opensource-ai@bgts.com.

UI walkthrough of the BGTS Context Engine web interface.

Quick start

# 1. PostgreSQL 16 with Apache AGE + pgvector, in one database
docker compose -f deploy/docker-compose.yml up -d

# 2. Install and migrate
pip install bgts-context-engine
cp .env.example .env
bce migrate

# 3. Index something
bce index --repo /path/to/your/repo --name my-service

# 4. Ask
bce context --task "fix the login timeout in the meeting webhook"

Then serve it:

bce serve        # REST at :8000/docs, web UI at :8000/ui/
bce serve-mcp    # MCP over stdio, for agents (install the [mcp] extra; see below)

Use it from your agent

The MCP surface is behind the mcp extra (Python MCP SDK 1.x: mcp>=1.0,<2). Install it so bce is on the user PATH, not only inside a project .venv. Cursor and VS Code spawn bce serve-mcp themselves and do not activate the venv:

pip install "bgts-context-engine[mcp]"
bce --version   # must work in a new terminal, with no venv activated

Every MCP client uses the same stdio command plus the database in the environment. Do not leave bce serve-mcp running in a terminal for the editor: stdout is the protocol, so the process stays silent, and the IDE starts its own copy.

One command per editor. Run it in the project the agent works on (the repository you indexed), pointing at the engine's .env:

bce --env-file /path/to/engine/.env cursor-init --repo-id my-service   # Cursor
bce --env-file /path/to/engine/.env claude-init --repo-id my-service   # Claude Code

cursor-init writes .cursor/mcp.json (merged into an existing one) and the rule .cursor/rules/bgts-context-engine.mdc, which tells the agent to call get_context_for_task first, treat its file:line entries as verified locations instead of grepping for them, and not to edit a file just because it was listed. claude-init writes .mcp.json, a marked section in CLAUDE.md, and a UserPromptSubmit hook (bce precontext) that runs the graph once per prompt and hands the agent the answer before its first turn — the flow the agent benchmark measured (−20 % tokens, half the search output, same or better checks; 14 tasks on a React/TypeScript codebase, same model and machine). Cursor's prompt hook cannot add context, so there the rule does that job. --no-hook, --repo-id (repeatable) and --bce-command adjust the files; both commands are safe to rerun.

After you add or change the MCP config, restart Cursor or VS Code (or Command Palette → “Developer: Reload Window”). The server should then show as enabled with eight tools (ten with indexing enabled). Setup detail: [docs/mcp.md](docs/mcp.md). The manual equivalents:

Cursor — user config ~/.cursor/mcp.json (applies to every project), or a project .cursor/mcp.json that stays local (the directory is gitignored):

{
  "mcpServers": {
    "bgts-context-engine": {
      "command": "bce",
      "args": ["serve-mcp"],
      "env": { "BCE_DB_HOST": "localhost", "BCE_DB_NAME": "bce" }
    }
  }
}

VS Code — user MCP settings, or a project .vscode/mcp.json (also gitignored):

{
  "servers": {
    "bgts-context-engine": {
      "type": "stdio",
      "command": "bce",
      "args": ["serve-mcp"],
      "env": { "BCE_DB_HOST": "localhost", "BCE_DB_NAME": "bce" }
    }
  }
}

Claude Code — one command:

claude mcp add bgts-context-engine --env BCE_DB_HOST=localhost -- bce serve-mcp

Claude Desktop — same block as Cursor, in claude_desktop_config.json.

If command: "bce" stays disconnected, the editor cannot see bce on PATH. Install as above, or skip a permanent install with uvx:

"command": "uvx",
"args": ["--from", "bgts-context-engine[mcp]", "bce", "serve-mcp"]

Then ask your agent something that needs the repository rather than the file you have open: "what breaks if I change the session TTL?" The agent calls get_context_for_task, and the other tools in [docs/mcp.md](docs/mcp.md) let it drill from there — exact callers, the blast radius of a change, the symbol behind a name — without guessing at file names.

What comes back

Not a list of file paths. A ranked pack, with the reasoning attached:

{
  "anchors": {
    "python::api::webhooks::handle_meeting_webhook#a3f1": ["explicit", "lexical"],
    "python::auth::session::refresh_session#88c2":        ["lexical", "semantic"]
  },
  "context": {
    "items": [
      { "symbol_id": "...refresh_session#88c2", "name": "refresh_session", "kind": "function",
        "file_id": "my-service:src/auth/session.py", "line": 41, "detail_level": "full",
        "graph_distance": 0, "score": 11.42, "tokens": 214, "content": "def refresh_session(...)" },
      { "symbol_id": "...SESSION_TTL#4b0d", "name": "SESSION_TTL", "kind": "constant",
        "file_id": "my-service:src/auth/config.py", "line": 12, "detail_level": "signature",
        "graph_distance": 2, "score": 6.10,  "tokens": 31,  "content": "SESSION_TTL: int" }
    ],
    "used_tokens": 1388, "budget": 1500, "included": 20, "skipped": 0
  },
  "coverage": {
    "anchor_source_count": 3, "connected_component_ratio": 0.875,
    "top_candidate_margin": 1.84, "orphan_ratio": 0.0,
    "touches_god_node": false, "commit_mismatch": false,
    "confidence": "high"
  }
}

Three things here that a vector store cannot give you:

anchors says why the engine looked where it did, and which independent sources agreed. Three sources agreeing is usually right; one is a guess.

coverage is a trust report. confidence: "low" means the engine found something but could not corroborate it — the moment for an agent to ask a follow-up question instead of editing. commit_mismatch means the index is behind your working tree.

detail_level falls off with graph distance: the symbol you are changing arrives in full, its neighbours as signatures, the outer ring as name @ file:line. That is how twenty genuinely relevant symbols fit in 1500 tokens — short enough for an agent to carry on every turn. Every item also names its file_id and line, so the agent opens the file instead of searching for the symbol.

How it works

task text
   │
   ├─ anchors       four independent sources nominate entry points:
   │                explicit names, task history, full-text, vector
   ├─ expansion     fixed-shape graph walk: callers 2 hops, callees 1,
   │                references, type hierarchy, same-file siblings
   ├─ scoring       weighted sum over reference kind, task signal, centrality,
   │                distance, leaf penalty, edge provenance
   ├─ scope         drop repositories this caller may not see
   ├─ narrowing     keep the top N
   ├─ assembly      fit the token budget, cheaper detail further out
   └─ coverage      report how much of this is trustworthy

Callers reach two hops and callees only one, on purpose: when you change a function, what breaks is upstream of it. Reference kind carries the heaviest weight, because a place that writes a value is where the bug lives while a place that reads it is usually just downstream. Centrality saturates at degree 20, because a logger touches everything and explains nothing.

The full formula, every weight, and the confidence thresholds are in [docs/retrieval.md](docs/retrieval.md).

Features

  • Code graph, not chunks. Symbols, CALLS, REFERENCES, INHERITS, IMPLEMENTS,

IMPORTS, HTTP ROUTES_TO handlers, and WHY: comments bound to what they explain.

  • Deterministic by construction. Sorted traversal, stable tiebreaks, versioned scoring

weights. bce bench verifies it by running each case repeatedly and comparing output.

  • Six languages. Python, JavaScript and TypeScript built in; Java, C# and Go behind the

langs extra. [Adding one](docs/languages.md#adding-a-language) touches two files.

  • Cross-language call edges. React Native and Expo bridges connect

NativeModules.Foo.bar() in TypeScript to bar in Objective-C, Swift or Kotlin — a hole no single parser can see.

  • Edge provenance you can audit. scip from a real compiler index, treesitter from

syntax, heuristic from a pattern match. Scored differently, reported per response.

  • Incremental re-indexing. git diff decides what to re-parse. Symbol ids survive file

moves and reformatting, so history and embeddings stay valid.

  • One database. Apache AGE and pgvector in the same PostgreSQL, so one query joins a

graph traversal, a vector search and a SQL filter — and one pg_dump backs up the index.

  • MCP and REST from one implementation. A focused tool set over stdio, the same functions

over HTTP. Nothing to drift.

  • A UI that explains itself. /ui ships in the wheel and replays a real retrieval call

stage by stage: anchors lighting up, expansion spreading, candidates scored and cut.

  • Runs offline. The default embedding provider is deterministic arithmetic over token

digests. No API key, no network, repeatable benchmarks. openai talks to any OpenAI-compatible /v1/embeddings server (vLLM, TEI, Ollama), so a model such as jina-code-embeddings-1.5b can run inside the perimeter.

Supported models

Embeddings only find entry points when the task text names nothing the graph already knows. They never rank the answer. Out of the box that seed is hashing: deterministic arithmetic, no API key, no network. For a real code model, set BCE_EMBEDDING_PROVIDER and BCE_EMBEDDING_MODEL to one of these:

| Model | Provider | Dimension | | --- | --- | --- | | voyage-code-3 | Voyage AI (voyage) | 1024 | | voyage-code-4 | Voyage AI (voyage) | 1024 | | jina-code-embeddings-1.5b | OpenAI-compatible (openai) | 1536 |

Voyage is a hosted API — pip install "bgts-context-engine[embed]" and BCE_VOYAGE_API_KEY. Jina is the on-prem path: any server that speaks /v1/embeddings (vLLM, TEI, Ollama). Switching the model or the dimension is a re-index (bce migrate --reset-embeddings). The knobs are in [docs/deployment.md](docs/deployment.md).

Coming next — same openai socket, not yet a fitted retrieval profile:

— the smaller sibling of 1.5b, for hosts that cannot hold 1.5B parameters.

code retriever.

Where it fits

| | Embedding RAG | Language server | BGTS Context Engine | | --- | --- | --- | --- | | Retrieval basis | text similarity | compiler index | code graph + anchors | | Cross-file, cross-repo | weak | per project | yes | | Cross-language edges | no | no | yes, heuristic | | Same query, same answer | no | yes | yes | | Ranked for a task | by similarity | not ranked | yes, with coverage | | Token budget aware | chunk count | no | yes, detail by distance | | Explains its own answer | no | no | anchors + provenance + confidence |

A language server is exact but scoped to what you have open. Embedding search is broad but unaccountable. This sits between them: repository-wide and cross-language like the former, exact and reproducible like the latter.

Measuring it

Retrieval quality claims are worthless without the task set they were measured on, so the harness ships instead of a leaderboard. You give it your own tasks and the symbols you believe answer them:

bce bench --cases my-tasks.json --out report.json

Each case is a task text plus its ground-truth symbol_ids. The report gives recall, precision, precision@1 and MRR per case, median and p95 latency, and two pass/fail checks that matter more than the scores: every case is run repeatedly and must return a byte-identical ordering, and any case with a scoped principal must not surface a repository that principal cannot read.

Building the case file is the real work — it means deciding, by hand, what the right answer is. It is also the only honest way to know whether a change to the scoring weights helped. The format and a worked example are in [docs/deployment.md](docs/deployment.md#benchmarking).

Roadmap

Ordered by how often it comes up, not by difficulty:

  • Scope enforcement on every layer. Layer 3 applies the per-user repository filter;

Layers 1 and 2 do not. Until that closes, the API belongs behind a proxy — see [SECURITY.md](SECURITY.md).

  • Streamable HTTP transport for MCP. Today the MCP surface is stdio only, so the server

runs next to the agent. Remote transport makes one index serve a team.

  • More languages. Rust, Kotlin and PHP are the most requested. The provider interface is

the contribution path with the least friction — see [docs/languages.md](docs/languages.md#adding-a-language).

  • Wider SCIP ingestion. Compiler-grade edges beat syntax-derived ones and are scored as

such; more toolchains means more of the graph carries scip provenance.

  • A published benchmark corpus. An open task set over public repositories, so results

are comparable between projects rather than only between your own runs.

Requests and disagreements belong in issues — what people actually ask for reorders this list.

Documentation

| | | | --- | --- | | [Architecture](docs/architecture.md) | the deterministic line, the three layers, indexing | | [Retrieval](docs/retrieval.md) | anchors, expansion, every

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.