# SaveContext

> Git LFS for your LLM context — 20x Context Savings — an MCP server that keeps big documents out of the model and returns byte-exact quotes for a fraction of the tokens.

- **Type:** MCP server
- **Install:** `agentstack add mcp-alexanderboger-savecontext`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [AlexanderBoger](https://agentstack.voostack.com/s/alexanderboger)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** Apache-2.0
- **Upstream author:** [AlexanderBoger](https://github.com/AlexanderBoger)
- **Source:** https://github.com/AlexanderBoger/SaveContext
- **Website:** https://pypi.org/project/savecontext/

## Install

```sh
agentstack add mcp-alexanderboger-savecontext
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# 🗜️ SaveContext

**Git LFS for your LLM context window.**

Keep big documents *outside* the model — hand it a small searchable handle instead.
The model reads a cheap summary to find what it needs, then pulls back the **exact
original words** to quote, using a fraction of the tokens.

An **MCP server** for Claude Code, Claude Desktop, and any Model Context Protocol
client — with a standalone CLI too.

[](https://github.com/AlexanderBoger/SaveContext/blob/master/LICENSE)
[](https://www.python.org)
[](https://modelcontextprotocol.io)
[](#reference)
[](#reference)

---

```
  ingest 150k-token contract ─▶ ctx://acme@v1
                                 ├─ ~700-token semantic brief
                                 ├─ block map  (b0000 … b0042)
                                 ├─ protected atoms  (money · dates · obligations …)
                                 └─ raw source stored compressed, never auto-sent

  ask 6 questions ─▶ lookup(...)  ─▶  one call · top blocks · exact atoms
                                       · verbatim quotes · confidence flag
```

## Install

With [uv](https://docs.astral.sh/uv/), add SaveContext to Claude Code in one line — no
clone, no `pip`, no virtualenv:

```bash
claude mcp add savecontext -- uvx savecontext
```

Then open Claude Code, run `/mcp` to confirm the tools loaded, and **point the model at
a big document and just ask** — it ingests, looks up, and quotes on its own. To pin
where the database lives, add `-e SAVECONTEXT_DB=/abs/path.db` before the `--`.

Using a different MCP client (Claude Desktop, etc.)? Point its server config at the same
command: `uvx savecontext`. Prefer no MCP at all? The [CLI](#reference) does everything
standalone.

## Why it's different

Most context tools are lossy summarizers — helpful, but they blur the exact numbers and
can't tell you when they're guessing. SaveContext does the two things they can't:

- **🎯 Exact, with receipts.** Facts come back as the *original characters* with their
  source location, never a paraphrase — every span reconstructs byte-for-byte. A summary
  is for *finding*; the quote is for *trusting*.
- **🙅 Says "not found" instead of inventing.** When the answer isn't in the document,
  retrieval is flagged weak and the agent abstains — rather than hallucinating a
  plausible one.
- **📉 Flat cost.** The payload to answer questions stays ~7,300 tokens whether the
  document is 15k or 150k. Pasting scales linearly; this doesn't.
- **🔁 One call, many questions.** `lookup` answers a batch of queries in a single round
  trip, not one retrieve-cycle per question.

## By the numbers

20.5×less contextat 150k tokens
~7.3kflat token cost,any document size
54/54live answers correct,incl. not&nbsp;found
100%byte-exact quotes,0 hallucinations

  
  
  

Measured with real agents on a contract set plus **held-out** documents (an RFC and a
synthetic incident log the retriever was never tuned on), answering fact questions that
have distractor values planted next to the true ones, graded against ground truth.
Crossover is ≈ 5k tokens — below that, just paste; above it, the gap widens with size.

> Honest note: live runs are single samples — treat small differences as noise. The
> durable findings are the flat-vs-scaling cost curve, byte-exact retrieval, and zero
> hallucination on absent facts. Dense technical prose is the current retrieval frontier.
> Reproduce it yourself: [`examples/benchmark_suite/`](https://github.com/AlexanderBoger/SaveContext/blob/master/examples/benchmark_suite).

---

## Reference

The 11 tools

All tools return clean JSON. Handles are stable and readable: `ctx://label@vN` for
contexts, `out://label@vN` for outputs.

| Tool | Purpose |
|------|---------|
| `ingest` | Store long text; return handle, brief, block map, protected-atom summary, safety notes. **Does not return the raw source.** Optional `agent_brief` lets the model supply its own summary. |
| `lookup` | **Batched retrieval** — answer several queries in one call. Per query: the top block(s), the exact atoms inside them, a verbatim best-matching sentence, and a `confidence` flag (`strong`/`weak`) so the agent can abstain instead of guessing. |
| `brief` | Task-specific brief: relevant blocks + atoms for a given task, with uncertainty notes. |
| `expand` | Lazily expand a block / atom / section / search match at `fidelity = summary \| facts \| quotes \| full`. |
| `quote` | Exact verbatim quote (by atom id or search query) with surrounding context and source ref. |
| `map` | Full (or paginated) block map + atom summary, fetched lazily. |
| `set_brief` | Replace a context's brief with an agent-authored one. Raw source and atoms unchanged. |
| `diff` | Diff new source against a stored context; report added/removed/changed **atoms** and meaning impact. Creates a new version. |
| `zip_output` | Store a large generated **output**; return section map + preview. |
| `expand_output` | Expand part of a stored output at `fidelity = preview \| section \| full`. |
| `audit` | Compression ratio, atom counts, per-category preservation estimate, `safe_for` / `unsafe_for`, warnings. |

**`expand` fidelity levels:** `summary` (extractive lead sentences, cheapest) · `facts`
(summary **plus** exact atoms, `[money] $500,000`) · `quotes` (verbatim spans you can
cite) · `full` (raw block text — the only path that returns source prose).

What a turn looks like

```text
1. User pastes a 150,000-token contract.
2. Claude → ingest(source_text=…, label="acme")
        ← ctx://acme@v1, ~700-token brief, block map, atoms.
3. User: "Liability cap? Notice period? Governing law?"
4. Claude → lookup(ctx://acme@v1, ["liability cap", "termination notice", "governing law"])
        ← per query: the governing block, its atoms, a verbatim sentence, confidence.
5. Claude → quote(ctx://acme@v1, search_query="shall not exceed")   # verify exact wording
        ← "…total liability shall not exceed $750,000…" + source offsets.
6. Claude answers using only the relevant, verbatim, cited context.
```

A runnable version is in [`examples/demo.py`](https://github.com/AlexanderBoger/SaveContext/blob/master/examples/demo.py) (`python examples/demo.py`).

Command line (savectx)

The package ships a CLI that shares the same store as the MCP server — ingest on one,
query from the other:

```bash
savectx ingest report.md --label q4 --type auto      # ingest a file
savectx ingest - --label pasted                       # or read stdin
savectx list                                          # what's in the store
savectx lookup ctx://q4@v1 --query "liability cap" --query "fees"   # batched
savectx brief  ctx://q4@v1 --task "liability risks"
savectx expand ctx://q4@v1 --selector liability --fidelity facts
savectx quote  ctx://q4@v1 --search "shall not exceed"
savectx audit  ctx://q4@v1
savectx diff   ctx://q4@v1 report_v2.md               # version + change report
savectx serve                                         # run the MCP server
```

Bare `savecontext` (and `python -m savecontext`) runs the MCP server, so MCP configs
keep working unchanged.

Configuration

Set via environment (or the MCP `env` block):

| Variable | Default | Purpose |
|---|---|---|
| `SAVECONTEXT_DB` | `~/.savecontext/savecontext.db` | Store location. `:memory:` for ephemeral. Shared by CLI and server. |
| `SAVECONTEXT_TRANSPORT` | `stdio` | `stdio` \| `http` \| `sse` |
| `SAVECONTEXT_HOST` / `SAVECONTEXT_PORT` | FastMCP defaults | For `http`/`sse` transport |
| `SAVECONTEXT_EMBEDDINGS` | off | `1` enables a local sentence-transformers retrieval backend (default is dependency-free BM25) |
| `SAVECONTEXT_LLM_SUMMARY` | off | `1` has a local Ollama model write an abstractive brief at ingest (fails closed to extractive) |

The `ctx://` handle scheme is stable. `SAVECONTEXT_*` vars take precedence; the
pre-rename `CONTEXTSAVER_*` / `CONTEXTVAULT_*` names are still honored as fallbacks.

How it works (architecture + the loss-aware guarantee)

| Layer | Module | Behaviour |
|-------|--------|-----------|
| Tokenizer | `tokenizer.py` | tiktoken `cl100k_base`, else `chars/4` |
| Chunking | `chunking.py` | deterministic: headings + paragraph packing to a token budget, exact offsets |
| Extraction | `extraction.py` | regex atoms: dates, money, %, numbers, URLs, emails, entities, negations, obligations, code ids, file paths |
| Retrieval | `retrieval.py` | **BM25** + concept-group expansion + answer-type intents + saturation scoring; optional embedding backend behind the same `rank` interface |
| Summary | `summarize.py` | extractive: ranked blocks + lead sentences + high-value atom lines; optional local-LLM |
| Diff | `diffing.py` | atom-set diff + coarse block text diff + risk impact |
| Storage | `storage.py` | SQLite metadata + **zstd-compressed** raw source; blocks/atoms as offsets |
| Service / Server | `service.py`, `server.py` | the eleven tools; FastMCP over stdio / HTTP / SSE |

**Loss-aware guarantee.** Blocks and atoms are stored as *offsets into the raw source*,
so any span reconstructs byte-for-byte. The brief is explicitly flagged lossy; exact
wording always comes from `quote()` or `expand(fidelity="quotes")`. The full raw source
is only ever returned via `expand(fidelity="full")`.

Who writes the brief?

Three summarizer options, best first — all keep the protected atoms exact; only the
prose brief changes:

1. **The calling agent (recommended).** Claude already has the document in context at
   `ingest` time, so it writes a faithful, dense brief for free — inline via
   `ingest(agent_brief=…)` or afterwards via `set_brief`. (`brief_mode='agent'`)
2. **Local LLM (optional).** `SAVECONTEXT_LLM_SUMMARY=1` → a local Ollama model writes an
   abstractive brief; fails closed to extractive. (`brief_mode='abstractive(local-llm)'`)
3. **Extractive (default).** Rule-based real sentences from the source — free, offline,
   deterministic, always available. (`brief_mode='extractive'`)

Development & tests

```bash
git clone https://github.com/AlexanderBoger/SaveContext && cd SaveContext
python -m pip install -e ".[tokenizer,dev]"
python -m pytest            # 80 tests
```

Covers tokenizer / handles / chunking / extraction / BM25 retrieval units, an end-to-end
pass over all eleven tools (offset exactness, verbatim quotes, versioning, money-change
diff, audit preservation buckets, error handling), the batched-`lookup` retrieval +
confidence logic, and the CLI flow.

Roadmap

- ✅ BM25 + hybrid semantic ranking (concept groups, answer-type intents) — no model required
- ✅ Batched `lookup` with per-query confidence / abstention
- ✅ Verbatim-integrity + extraction-recall reporting in `audit`
- ✅ PDF / DOCX / HTML loaders and recursive folder ingest
- ✅ HTTP / SSE transport in addition to stdio
- ⬜ Embedding backend benchmarked against the held-out suite (dense-prose retrieval)
- ⬜ Vector-store persistence so embeddings aren't recomputed per query
- ⬜ LLM-assisted atom extraction for fuzzy / multi-span facts

## License

[Apache License 2.0](https://github.com/AlexanderBoger/SaveContext/blob/master/LICENSE) — see also [NOTICE](https://github.com/AlexanderBoger/SaveContext/blob/master/NOTICE). If you redistribute
SaveContext or build on it, retain the copyright and the `NOTICE` file, per the license
(that's how attribution is required).

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [AlexanderBoger](https://github.com/AlexanderBoger)
- **Source:** [AlexanderBoger/SaveContext](https://github.com/AlexanderBoger/SaveContext)
- **License:** Apache-2.0
- **Homepage:** https://pypi.org/project/savecontext/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-alexanderboger-savecontext
- Seller: https://agentstack.voostack.com/s/alexanderboger
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
