# Whittle

> Carves your agent's tool outputs down to what matters. Never cuts what doesn't come back.

- **Type:** MCP server
- **Install:** `agentstack add mcp-firstops-dev-whittle`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [firstops-dev](https://agentstack.voostack.com/s/firstops-dev)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** Apache-2.0
- **Upstream author:** [firstops-dev](https://github.com/firstops-dev)
- **Source:** https://github.com/firstops-dev/whittle
- **Website:** https://github.com/firstops-dev/whittle

## Install

```sh
agentstack add mcp-firstops-dev-whittle
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# 🪓 whittle

**Cut your AI agent's token bill. Lose nothing that matters.**

[](https://github.com/firstops-dev/whittle/actions/workflows/ci.yml)
[](https://github.com/firstops-dev/whittle/releases)
[](https://www.npmjs.com/package/@firstops/whittle)
[](https://pkg.go.dev/github.com/firstops-dev/whittle)
[](LICENSE)

  Install ·
  How it works ·
  Compression ·
  Model routing ·
  Benchmarks ·
  Why whittle?

Long agent sessions drown in tokens, and most compressors buy their ratio by silently destroying what agents need: array rows vanish, file reads get gutted, identifiers come back mangled. Whittle holds one line: **lossless or clearly marked, code never touched, every anomaly fails open to the original bytes.**

Built for engineers running long Claude Code sessions who refuse to trade fidelity for token savings.

## Highlights

- 🪓 **Write-time compression**: a Claude Code PostToolUse hook whittles each tool output *before* it enters context, so the savings repeat on every later turn. History is never mutated, so prompt caches stay intact.
- 🔒 **Lossless or marked**: JSON reshapes byte-exact (rows are never dropped), logs keep every error plus an honest `[N lines omitted]`, source code passes through untouched.
- 🎯 **Token-honest numbers**: savings measured against calibrated tokenizer counts, not byte counts that overstate by up to 4×.
- 🧭 **Opt-in model router**: hard reasoning stays on your strongest model, trivia drops to the cheapest, per one auditable policy file driven by trained-classifier signals ([how it works](docs/ROUTER.md)).
- 🛟 **Fail-open, local, yours**: if whittle is down your agent runs on originals, a rejected rewrite retries your original request, zero credentials leave your box, and the Go binary has zero external dependencies.

## Install

```sh
npm install -g @firstops/whittle   # or: go install github.com/firstops-dev/whittle/cmd/whittle@latest
whittle setup                      # hook + local daemon + optional ML sidecar, one command
```

Tool outputs are whittled from now on; `whittle stats` shows what you're saving. Try it with zero commitment: `npx -y @firstops/whittle compress build.log`. (Homebrew: `brew install firstops-dev/tap/whittle`. Linux runs the daemon under systemd, see [notes](docs/compression.md).)

**Optional: turn on model routing.**

```sh
whittle policy init                              # calibrated policy, your model ids auto-detected
whittle route -install                           # background service (or `whittle route` in a terminal)
export ANTHROPIC_BASE_URL=http://127.0.0.1:45873
```

## See it

One live feed, both surfaces: `whittle watch` streaming real routing verdicts (trivia dropped to haiku with the classifier signals inline, hard work kept on opus) and real compression carves. Captured from a live session; every line is genuine output.

```
$ whittle compress -stats demo/build.log
... [59 lines omitted]
2026-07-04T10:00:58Z INFO  compiling module pkg/058 ok (406ms)
2026-07-04T10:00:41Z ERROR migrate failed: relation "users" does not exist
... [79 lines omitted]
2026-07-04T10:02:19Z ERROR shutdown: connection reset by peer
2026-07-04T10:02:20Z INFO  144 ok, 2 errors

whittle: action=compressed detected=log strategy=ansi_strip+log_compressor tokens=4728->163
```

Both errors, their context, and the summary survive; 138 lines of noise collapse into two counted markers. Across **5,000 real agent sessions**: **22% tool-output reduction at zero measured information loss**: mechanically lossless on 15,846/15,846 items, and a blinded 4-judge panel found 0/120 material loss on the lossy prose path. Full receipts, including an honest side-by-side against headroom: [bench/](bench/README.md).

## Compression: what happens to each content type

JSON is reshaped **losslessly** (byte-exact reconstruction). Logs and terminal streams are cut lossily but **marked and exactly accounted**. Markdown file reads keep code fences, tables, and headings byte-exact while prose is compressed by an optional local model with fidelity guards. Source code is **never touched**. Every path is guarded: the worst case is always *not compressed*, never *corrupted*.

The full per-type contract, ML prose path, architecture, and performance tables: [docs/compression.md](docs/compression.md) · what each guarantee is pinned by: [GUARANTEES.md](GUARANTEES.md).

## Model routing (opt-in)

Whittle's second surface: a local proxy on `ANTHROPIC_BASE_URL` that sends each request to the cheapest model tier that can still handle it, per a policy you can read in one screen.

- **Calibrated out of the box**: `whittle policy init` writes a conservative default (hard reasoning → strongest tier, confident chit-chat → cheapest, *everything else untouched*) with your account's real model ids auto-detected. [What the default does](router/policies/default.md) · [how to write your own policy](docs/POLICY.md).
- **Multi-signal, not keyword-matching**: a trained 14-subject classifier (probability-mass thresholded, so an *uncertain* classification never escalates), a contrastive difficulty score, and your own keywords. Every log line shows each signal's value against its gate.
- **Rewrites the model, never your history**: prompt-cache prefixes survive; capabilities the cheaper model rejects are stripped automatically; credentials pass through untouched.
- **Never blocks you**: bad policy, dead classifier, or a rejected rewrite all fall back to your original request. Unset the env var and you're direct again.
- **Savings you can measure**: every request logs requested model, served model, and real token usage. `whittle watch` streams both feeds live (routing verdicts + compression carves) in one view.

The router itself (engine, policy design, signal composition) is whittle's own; two pretrained models power its ML signals (the 14-subject `domain` classifier and the text embedder, both from [vLLM Semantic Router](https://github.com/vllm-project/semantic-router)). Architecture, signal math, and precise credits: [docs/ROUTER.md](docs/ROUTER.md).

## Why write-time?

Most context compressors are read-time proxies: they rewrite your conversation history on every LLM call, which breaks prompt caches, terminates your API traffic, and makes lossy compression the default. Whittle compresses each output **once, at the moment it's born**, before it enters history: savings compound across every later turn, and nothing sits in your request path.

| | whittle | read-time compressors |
|---|---|---|
| compresses | once, at write time | re-runs on every LLM call |
| prompt cache | intact, history never rewritten | invalidated by history rewrites |
| your request path | nothing resident in it | a proxy that must stay up |
| loss model | lossless or marked | lossy by default, recover on demand |
| cost of failure | original bytes | a broken or blocked call |

The full argument, and why routing is a different kind of proxy: [docs/why-whittle.md](docs/why-whittle.md).

## FAQ

**Will it break my agent?** No, and that is the core design constraint. Every path fails open: if whittle is down, declines, or errors, your agent sees original bytes. The router likewise: worst case is your request untouched.

**Does it need Python?** No. The deterministic compressors (JSON, logs, terminal, markdown structure) are pure Go. Python powers the *optional* prose model and router classifiers; `whittle setup` installs it if `python3` exists, and everything else works without it.

**Are token savings dollar savings?** Not 1:1. Under prompt caching, cheap cache-reads dominate the bill, so a 22% token cut is roughly 3–5% of session cost. We publish both numbers rather than pretending otherwise: [bench/](bench/README.md).

**How does it compare to headroom?** On identical bytes, headroom's defaults compress ~5 points more, by dropping rows whittle refuses to drop. On conversation-shaped content whittle leads *while staying lossless*. Which trade you want is the whole point: [bench/SIDEBYSIDE.md](bench/SIDEBYSIDE.md).

**Where does my data go?** Nowhere. Hook, daemon, models, and router all run on your machine. The router forwards your own credentials to Anthropic and logs token *counts*, never prompt text.

**Which agents?** Claude Code today (hook + router). Cursor, Codex, and OpenCode adapters are on the roadmap; the compression engine is also a plain Go library and HTTP service.

## Verify it yourself

Every claim here is checkable from a clone; that is the point.

```sh
make test                          # guarantees as executable tests (GUARANTEES.md)
go run ./bench                     # corpus reductions + fidelity, SHA-pinned, CI-gated
python bench/calibrate_tokens.py   # reproduces the token-estimator MAE
```

## Contributing

The bar: guarantees are executable, see [CONTRIBUTING.md](CONTRIBUTING.md). Fidelity bugs (whittle changing an output's meaning) are treated as urgent.

Near-term roadmap: a tagged release that ships the router, agent adapters (Cursor, Codex, OpenCode), Linux packaging. Good first issues: adapters, packaging, detection corpus cases.

## Acknowledgments

Whittle's log-selection strategy, several detection heuristics, and the tabular parser were adapted from [Headroom](https://github.com/headroomlabs-ai/headroom) (Apache-2.0); their compaction work is excellent; we wanted the write-time position and the stricter fidelity contract it demands. The router's two pretrained models (the `domain` classifier and the text embedder behind the similarity signals) come from [vLLM Semantic Router](https://github.com/vllm-project/semantic-router) ([whitepaper](https://vllm-semantic-router.com/white-paper)); the routing engine and policy design are whittle's own. See [NOTICE](NOTICE).

If whittle's fidelity contract is the trade you want, a ⭐ helps other agent users find it.

## License

Apache-2.0.

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [firstops-dev](https://github.com/firstops-dev)
- **Source:** [firstops-dev/whittle](https://github.com/firstops-dev/whittle)
- **License:** Apache-2.0
- **Homepage:** https://github.com/firstops-dev/whittle

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-firstops-dev-whittle
- Seller: https://agentstack.voostack.com/s/firstops-dev
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
