Install
$ agentstack add mcp-lorenzo-cambiaghi-lynxmcp ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Lynx
A 100% local MCP server for semantic code search — AST-aware chunking, hybrid BM25 + dense retrieval, and an optional code knowledge graph. Works with any MCP client (Claude Code, Cursor, Windsurf, Antigravity, ...).
[](https://github.com/lorenzo-cambiaghi/LynxMCP/actions/workflows/test.yml) [](LICENSE)
[](https://glama.ai/mcp/servers/lorenzo-cambiaghi/LynxMCP)
[](https://glama.ai/mcp/servers/lorenzo-cambiaghi/LynxMCP)
Your AI assistant greps file names and guesses. Lynx gives it real retrieval over your code, your library docs, and your PDFs — without a single byte leaving your machine.
💸 What it saves you — every wrong file your AI opens is billed tokens
Agentic coding burns tokens re-reading files the assistant grepped into the wrong place. Lynx hands it the right code in one tool call, with file:line and symbol — measured on real codebases:
| Tokens to get the answer into context | Agentic grep | Lynx | | |---|---:|---:|:--| | Django 5.2 — Python, 158k lines | 4,150 | 1,725 | −58% | | Json.NET — C#, 69k lines | 6,590 | 1,540 | −77% | | Guava — Java, 181k lines | 5,892 | 807 | −86% |
Plus: outline triage is 2.4× fewer tokens, and the code arrives in 1 tool call instead of 2+ (chunks included, with symbol + file:line + score). The token cut holds across languages — even where grep ranks results just as well, because Lynx returns the whole function in one call instead of match-lines plus a follow-up read.
That's real money at today's frontier API prices. For 25 engineers (≈31,500 retrievals/month), the yearly API bill Lynx removes:
| Flagship model (input $/1M) | Django (Python) | Json.NET (C#) | Guava (Java) | |---|---:|---:|---:| | Claude Fable 5 — Anthropic flagship ($10) | ≈ $85,000 | ≈ $95,000 | ≈ $95,000 | | GPT‑5.5 — OpenAI flagship ($5) | ≈ $42,000 | ≈ $47,000 | ≈ $47,000 | | Claude Opus 4.8 — top coding model ($5) | ≈ $42,000 | ≈ $47,000 | ≈ $47,000 |
Token deltas are measured ([Django](benchmarks/RESULTS.md) · [Json.NET](benchmarks/RESULTScsharp.md) · [Guava](benchmarks/RESULTSjava.md)). The yearly figures add one eliminated grep round‑trip re‑billing a 20k‑token context; the conservative floor (tool output only, zero assumptions) is $0.4k–1.6k/mo depending on model and codebase. Run it for your own team, prices and codebase: CLI python benchmarks/savings_calculator.py --devs N, or the interactive [savings calculator](benchmarks/savings_calculator.html) — pick the codebase and model from drop‑downs and edit the $/1M price live (presets in [benchmarks/pricing.json](benchmarks/pricing.json) + [measured.json](benchmarks/measured.json), yours to change).
- AST-aware indexing — tree-sitter parses 18+ languages (19 grammars, counting TSX) and indexes whole functions/classes, not arbitrary text windows.
- Hybrid retrieval — dense embeddings + code-tokenized BM25, fused with RRF; optional cross-encoder reranker.
- Token-efficient triage —
view=outlinereturns signatures instead of bodies, so an agent scans the candidates for ~2.4× fewer tokens and reads only the code it picks ([measured](docs/OUTLINE.md)). - Code knowledge graph (opt-in) — who-calls-what, inheritance, imports: ask "what breaks if I change this?" and get the actual blast radius — or export it as a single, shareable, offline graph view (
lynx graph export). - Joinable as SQL — search and the graph are also served as rows over a local HTTP API, so you can correlate your code with tickets, PRs, or logs in [DuckDB](docs/DUCKDB.md) or [Coral](docs/CORAL.md) — no data leaves your machine.
- Multi-source — index codebases, public docs sites (fetched once, on demand; JS-rendered SPAs supported via optional headless Chromium), and PDFs side by side.
- Live index — a file watcher re-indexes saves in ~2s. No manual rebuild ritual.
- [Web manager UI](docs/GUIDE.md#lynxmanager--guided-setup-web-ui-diagnostics-new-in-v09) —
lynx manager uigives you guided setup, a query playground, diagnostics, and client config snippets.
LynxManager — guided setup, query playground & diagnostics, all in the browser. Full walkthrough →
Shareable graph views — lynx graph export --symbol GetVoxel writes one self-contained, offline file (no server, no CDN): the symbol's blast radius — who calls it (above) and what it calls (below). Attach it to a PR or archive it for an audit.
Quickstart
# 1. Install the CLI (isolated, no venv ritual)
pipx install lynx-mcp
# or: uv tool install lynx-mcp
# 2. Create a config pointing at your project
lynx manager init
# 3. Build the index (downloads the ~130MB embedding model on first run)
lynx build
Then register Lynx in your MCP client (Claude Code shown; see the [full guide](docs/GUIDE.md) for Cursor, Antigravity, and generic stdio clients — or let lynx manager ui generate the snippet for you):
{
"mcpServers": {
"lynx": {
"command": "lynx",
"args": ["serve", "--config", "/absolute/path/to/config.json"]
}
}
}
Prefer zero terminal? There are double-click installers for macOS and Windows.
The tools your AI gets
The tool set is fixed — it does not grow with the number of sources, so your client's tool list (and context window) stays small. Tools take a source argument where relevant.
| Tool | What it answers | |------|-----------------| | search(query, source?, outline?) | Primary hybrid search. Omit source to search every source at once (RRF-fused). outline=true returns signatures-only for cheap triage (see below). | | deep_search(queries, source?) | Escalation: tries multiple query phrasings until one passes a quality threshold. | | graph_query(operation, symbol?) | callers, callees, subclasses, superclasses, imports, neighbors, shortest_path, overview, surprising_connections, status. | | find_definition(symbol) | Where is X defined? (AST-precise when the graph is on, BM25 fallback otherwise.) | | find_usages(symbol) | Every use of X — calls and non-call references (generics, decorators, docs). | | find_tests_for(symbol) | Are there tests for X? | | find_similar(snippet) | Does code like this already exist? | | describe_symbol(symbol) | One-shot context for X: definition + who calls it + what it calls + its tests, in a single call. | | impact(symbol) | Blast radius: everything that reaches X transitively through the call graph (with hop distance) + the tests to re-run. | | module_summary(file) | A file as a unit: the symbols it defines, what it imports, and which files depend on it. (graph) | | repo_overview() | "What is this and where do I start": detected languages, frameworks, entry points, and build/test/run commands. | | export_graph(target, mode?) | Render a shareable, offline graph view — a symbol's blast radius or a file hub — as a single self-contained file. (graph) | | search_diff(query, base?) | Search only the files changed vs a base branch — built for code review. | | feedback(trying_to_do, tried, stuck) | The agent files a report when the index couldn't answer — stored 100% locally, your signal for tuning sources. | | list_sources / get_rag_status / update_source_index | Introspection and maintenance. |
Retrieval tools carry MCP readOnlyHint annotations (clients can auto-approve them); the only write is export_graph, which saves a graph view file. The server ships its usage playbook in the MCP handshake (instructions + a lynx://guide resource) — your agent knows how to query well without any rules-file setup.
(graph) tools need the optional code knowledge graph enabled for the source. The tool set is per-capability, never per-source.
How it works
flowchart LR
A["Your code + docs + PDFs"] --> B["Tree-sitterAST chunker"]
B --> C["bge-smalldense embeddings"]
B --> D["code-tokenizedBM25"]
B --> G["Code knowledge graph(opt-in)"]
C --> R{{"RRF fusion"}}
D --> R
Q(["Your query"]) --> R
R --> RR["Optionalreranker"]
RR --> RES["Ranked codefile : line : symbol"]
G --> GT["Graph toolscallers · subclasses · usages"]
classDef store fill:#fff3e6,stroke:#e8742c,color:#24292f;
classDef out fill:#e8742c,stroke:#e8742c,color:#fff;
class C,D,G store;
class RES,GT out;
Everything runs locally: HuggingFace models are downloaded once, then Lynx switches to offline mode. No telemetry, no cloud index, no code upload. The only network access is the model download and the explicit webdoc fetch step you trigger yourself.
Restricted networks / air-gapped machines
The embedding model is a public HuggingFace model (BAAI/bge-small-en-v1.5, ~130MB) — no account or token is required. If you hit We couldn't connect to 'https://huggingface.co', the machine simply can't reach the Hub (firewall, proxy, DNS, or an offline box).
You usually don't need to do anything. When the HuggingFace download fails, Lynx automatically falls back to a copy of the model hosted on this repo's GitHub Releases and installs it from there — including on the installer's first run. You only need the steps below if GitHub is also unreachable, or if you want to use a mirror / a shared cache / your own host.
- Point the fallback elsewhere — if you can't reach github.com either but you host the archive somewhere reachable (an internal server, an artifact store), set the base URL and the automatic fallback uses it:
``bash export LYNX_MODEL_ARCHIVE_BASE_URL=https:///lynx-models # expects /BAAI--bge-small-en-v1.5.zip (produced by --export-archive) ``
- Mirror — point Lynx at a reachable HuggingFace mirror and (optionally) a shared cache, then download normally:
``bash export HF_ENDPOINT=https:// # e.g. an internal proxy or hf-mirror.com export HF_HOME=/shared/hf-cache # optional: shared/persistent cache lynx manager install --model ``
- Transfer an archive — on a machine with access, export the model, copy the file to the offline machine (USB,
scp, an internal share…), then import it:
```bash # online machine lynx manager install --model lynx manager install --export-archive bge-small.zip
# offline machine — local path or a direct download URL both work lynx manager install --from-archive /path/to/bge-small.zip lynx manager install --from-archive "https:///bge-small.zip" `` A URL only works if it serves the file **directly, with no authentication and no interstitial page**. A GitHub Release asset on a public repo is the easiest option — the bundled Publish model archive workflow can create one for you. **Google Drive does *not* work as a --from-archive` URL for this model: for files larger than ~100MB Drive returns a "can't scan for viruses" HTML page instead of the file, so the import would get HTML, not a zip (Lynx detects this and tells you). Use Drive only to hand the file to a person, who downloads it in a browser and passes the local path**.
- Check what's configured —
lynx manager doctorreports the active cache dir, whether a mirror is set, and whether the model is present.
Why not just let the agent grep?
Grep is great when you know the identifier. It fails when you (or the agent) know the behavior: "where do we clamp the camera zoom?" matches nothing literal. Agentic grep also burns tokens — every wrong file the agent opens is context spent. Lynx answers behavioral queries in one tool call with file + line + symbol citations, and the graph layer answers structural questions (callers, inheritance) that grep fundamentally cannot — polymorphic dispatch leaves no textual trace.
Honest counterpoint: on a small repo that fits in the agent's context, built-in tools are fine. Lynx pays off on large codebases, on framework docs your model's training data has gone stale on, and on repeated sessions where re-exploring from scratch is waste.
Benchmarks (reproducible)
On the django/ package of Django 5.2 (883 files, ~158k lines), 20 behavioral questions with known ground-truth files — full methodology, per-task results, and an intentionally strong grep baseline in benchmarks/RESULTS.md:
| | Agentic grep | Lynx | |---|---|---| | median tokens to answer (tool output + required follow-up read) | 4,150 | 1,725 | | tool round-trips before the code is in context | 2+ | 1 (chunks included, with symbol + file:line + score) | | hit@1 / MRR | 45% / 0.64 | 55% / 0.67 | | "what inherits from Field?" — full descendant tree (100 classes) | 101 grep rounds | 4 graph calls, same recall, file:line per edge |
The ranking quality is comparable (Django's docstring-rich code is grep's best case — we say so in the report). The structural difference is not: every tool round-trip is a full model inference over the growing context, and class-relation questions force grep into one round per discovered class while graph_query reads resolved inheritance edges.
Second language, sparser docs — the gap widens. The same test on Json.NET (C#) — Src/Newtonsoft.Json/, 240 files, 69k lines, 15 behavioral questions (RESULTS_csharp.md). With C#'s PascalCase identifiers and fewer narrative comments, Lynx wins every metric, ranking included:
| | Agentic grep | Lynx | |---|---|---| | median tokens to answer | 6,590 | 1,540 (−77%) | | hit@1 / MRR | 33% / 0.47 | 47% / 0.58 |
This is the counter-example the Django report predicts: move off grep's best case and the lexical baseline drops, while semantic retrieval holds.
**Third language, grep's best case — the token gap holds anyway. Guava (Java)** — com/google/common/, 606 files, 181k lines, 15 questions (RESULTS_java.md). Guava's self-documenting class names (BloomFilter, RateLimiter, Splitter) are ideal for lexical search — so here grep actually out-ranks Lynx. The metric you pay for still collapses:
| | Agentic grep | Lynx | |---|---|---| | median tokens to answer | 5,892 | 807 (−86%) | | hit@1 / MRR | 73% / 0.81 | 60% / 0.70 |
The honest takeaway across all three. Ranking parity swings with how self-documenting the code is — Lynx ahead on C#, level on Python, behind on Guava. But the token cost — the line on your invoice — drops 58–86% every time, because Lynx hands back the whole function in one call instead of match-lines plus a follow-up read. That's the number that scales to a team's monthly bill.
# reproduce — Python (Django)
git clone --depth 1 --branch 5.2 https://github.com/django/django.git benchmarks/_target/django
python benchmarks/run_benchmark.py && python benchmarks/structural_demo.py
# reproduce — C# (Json.NET)
git clone --depth 1 https://github.com/JamesNK/Newtonsoft.Json.git benchmarks/_target/jsonnet
python benchmarks/run_benchmark.py --tasks benchmarks/tasks_jsonnet.json \
--target-dir benchmarks/_target/jsonnet --storage-dir benchmarks/_storage_csharp \
--results-json benchmarks/results_csharp.json --results-md benchmarks/RESULTS_csharp.md
# reproduce — Java (Guava)
git clone --depth 1 https://github.com/google/guava.git benchmarks/_target/guava
python benchmarks/run_benchmark.py --tasks benchmarks/tasks_guava.json \
--target-dir benchmarks/_target/guava --storage-dir benchmarks/_storage_java \
--results-json benchmarks/results_java.json --results-md benchmarks/RESULTS_java.md
Two ways to read a result: full vs outline
Every Lynx search ranks the same way (hybrid dense + BM25 over whole functions). What differs is **how much of
…
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: lorenzo-cambiaghi
- Source: lorenzo-cambiaghi/LynxMCP
- License: Apache-2.0
- Homepage: https://github.com/lorenzo-cambiaghi/LynxMCP
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.