# Audrey

> Persistent memory and continuity engine for Claude Code and AI agents.

- **Type:** MCP server
- **Install:** `agentstack add mcp-evilander-audrey`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [Evilander](https://agentstack.voostack.com/s/evilander)
- **Installs:** 0
- **Category:** [Databases](https://agentstack.voostack.com/c/databases)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [Evilander](https://github.com/Evilander)
- **Source:** https://github.com/Evilander/Audrey
- **Website:** https://www.npmjs.com/package/audrey

## Install

```sh
agentstack add mcp-evilander-audrey
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

The local-first memory firewall for AI agents.

  
    Give Codex, Claude Code, Claude Desktop, Cursor, Windsurf, VS Code, JetBrains, Ollama-backed agents,
    and custom agent services one durable memory layer they can check before they touch tools.
  

  
    
    
    
  

## In Plain English

AI coding assistants are brilliant but forgetful. They'll happily rerun the same broken command they ran yesterday, forget the rules your team agreed on last week, and treat every new session like it's day one.

Audrey is the memory they're missing. It quietly keeps track of what worked, what failed, and what you told it — then checks that memory **before** the agent does something, so it can say "hold on, this exact command failed last time, and here's what fixed it" instead of repeating the mistake. Everything lives in one local file on your machine: no cloud, no account, and nothing about your code ever leaves your computer.

That's the whole idea. The rest of this README is the detail.

## Why Audrey Exists

Agents forget the exact mistakes they made yesterday. They repeat broken commands, lose project-specific rules, miss contradictions, and treat every new session like a cold start.

Audrey Guard is the headline loop: record what happened, remember what mattered, check before action, return `allow`, `warn`, or `block` with evidence, then validate whether the memory helped.

Audrey turns those hard-won lessons into a local memory runtime:

- `audrey guard --tool Bash "npm run deploy"` runs memory-before-action from the terminal.
- `memory_recall` finds durable context by semantic similarity.
- `memory_preflight` checks prior failures, risks, rules, and relevant procedures before an action.
- `memory_reflexes` converts remembered evidence into trigger-response guidance agents can follow.
- `memory_validate` closes the loop after the action: `helpful`, `used`, or `wrong` outcomes feed salience and can bind back to the exact preflight event, evidence ids, and Guard action fingerprint.
- `memory_dream` consolidates episodes into principles and applies decay.
- `audrey impact` and `audrey doctor` tell a human or CI system whether the runtime is doing real work and is actually ready.

It is not a hosted vector database, a notes app, or a Claude-only plugin. Audrey is a SQLite-backed continuity layer that can sit under any local or sidecar agent loop.

  

## Quick Start

Requires Node.js 20+.

```bash
npx audrey doctor
npx audrey demo --scenario repeated-failure
npx audrey guard --tool Bash "npm run deploy"
```

`doctor` verifies Node, the MCP entrypoint, provider selection, memory-store health, and host config generation. The repeated-failure demo is no-key, no-host, and no-network: it creates a temporary store, records a failed deploy, teaches Audrey the fix, then shows Audrey Guard blocking the repeat attempt with evidence.

Expected first-run shape:

```text
Audrey Doctor v1.0.2
Store health: not initialized
Verdict: ready
```

After the first real memory write, `doctor` should report the store as healthy.

## Install Into Agent Hosts

Preview host setup without editing config files:

```bash
npx audrey install --host codex --dry-run
npx audrey install --host claude-code --dry-run
npx audrey install --host generic --dry-run
```

Generate raw config blocks:

```bash
npx audrey mcp-config codex
npx audrey mcp-config generic
npx audrey mcp-config vscode
npx audrey hook-config claude-code
```

Claude Code can be registered directly:

```bash
npx audrey install
claude mcp list
```

For memory-before-action hooks, preview with `npx audrey hook-config
claude-code`, then apply with `npx audrey hook-config claude-code --apply
--scope project` for `.claude/settings.local.json` or `--scope user` for
`~/.claude/settings.json`. Audrey merges the hook block into existing settings
and writes a timestamped backup before changing a non-empty file. The generated
`PreToolUse` hook runs `audrey guard --hook --fail-on-warn`; the `PostToolUse`
and `PostToolUseFailure` hooks record redacted tool traces. Verify the active
hook set inside Claude Code with `/hooks`.

All local MCP paths default to local embeddings and one shared SQLite-backed memory directory. **Set a distinct `AUDREY_DATA_DIR` per tenant, agent identity, or concurrent host.** SQLite uses WAL mode without an advisory lock, so two processes sharing a directory will contend on writes. Isolation is a hard requirement for multi-agent setups, not a recommendation.

Installer-generated host config does not include provider API keys by default. Prefer setting `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GOOGLE_API_KEY`, or `GEMINI_API_KEY` in the host runtime environment; use `npx audrey install --include-secrets` only if you explicitly accept argv/config exposure.

## Use With Ollama And Local Agents

Ollama runs models; Audrey supplies memory. Start Audrey as a local REST sidecar and expose its routes as tools in your agent loop:

```bash
AUDREY_AGENT=ollama-local-agent npx audrey serve
curl http://localhost:7437/health
curl http://localhost:7437/v1/status
```

Runnable example:

```bash
AUDREY_AGENT=ollama-local-agent npx audrey serve
OLLAMA_MODEL=qwen3 node examples/ollama-memory-agent.js "What should you remember about Audrey?"
```

Core sidecar tools:

| Agent Need | REST Route |
|---|---|
| Check memory before acting | `POST /v1/preflight` |
| Get reflex rules for an action | `POST /v1/reflexes` |
| Store a useful observation | `POST /v1/encode` |
| Recall relevant context | `POST /v1/recall` |
| Get a turn-sized memory packet | `POST /v1/capsule` |
| Check health | `GET /v1/status` |

## What Ships

| Surface | Status |
|---|---|
| MCP stdio server | 20 tools plus status/recent/principles resources and briefing/recall/reflection prompts |
| CLI | `doctor`, `demo`, `guard`, `install`, `mcp-config`, `hook-config`, `status`, `dream`, `reembed`, `observe-tool`, `promote`, `impact` |
| REST API | Hono server with `/health` and `/v1/*` routes |
| JavaScript SDK | Direct TypeScript/Node import from `audrey` |
| Python client | `pip install audrey-memory`, calls the REST sidecar |
| Storage | Local SQLite plus `sqlite-vec`, no hosted database required |
| Deployment | npm package, Docker, Compose, host-specific MCP config generation |
| Safety loop | preflight warnings, reflexes, redacted tool traces, contradiction handling |

## Memory Model

Audrey is built around the parts of memory that matter for agents:

- Episodic memory: specific observations, tool results, preferences, and session facts.
- Semantic memory: consolidated principles extracted from repeated evidence.
- Procedural memory: remembered ways to act, avoid, retry, or verify.
- Affect and salience: emotional weight and importance influence recall.
- Interference and decay: stale, conflicting, or low-confidence memories lose authority over time.
- Contradiction handling: competing claims are tracked instead of silently overwritten.
- Tool-trace learning: failed commands and risky actions become future preflight warnings.

The product bet is simple: the next generation of useful agents will not just retrieve facts. They will remember what happened, decide whether a memory is still trustworthy, and use that memory before touching tools.

## Use Audrey From Code

### JavaScript

```js
import { Audrey } from 'audrey';

const brain = new Audrey({
  dataDir: './audrey-data',
  agent: 'support-agent',
  embedding: { provider: 'local', dimensions: 384 },
});

await brain.encode({
  content: 'Stripe returns HTTP 429 above 100 req/s',
  source: 'direct-observation',
  tags: ['stripe', 'rate-limit'],
});

const memories = await brain.recall('stripe rate limit');

await brain.waitForIdle();
brain.close();
```

### Python

```bash
pip install audrey-memory
```

```python
from audrey_memory import Audrey

brain = Audrey(base_url="http://127.0.0.1:7437", agent="support-agent")
memory_id = brain.encode("Stripe returns HTTP 429 above 100 req/s", source="direct-observation")
results = brain.recall("stripe rate limit", limit=5)
brain.close()
```

## Production Readiness

Audrey is close to a 1.0-ready local memory runtime, but production depends on how it is embedded. Treat it like stateful infrastructure.

Release gates used for this package:

```bash
npm run release:gate
npm run python:release:check
npm run bench:guard:card
npm run bench:guard:validate
npx audrey doctor
npx audrey demo
```

Recommended runtime checks:

```bash
npx audrey doctor --json
npx audrey status --json --fail-on-unhealthy
npx audrey install --host codex --dry-run
```

Production controls you still own:

- Set one `AUDREY_DATA_DIR` per tenant, environment, or isolation boundary.
- Pin `AUDREY_EMBEDDING_PROVIDER` and `AUDREY_LLM_PROVIDER` explicitly.
- Back up the SQLite data directory before provider or dimension changes.
- Keep API keys and raw credentials out of encoded memory content.
- Use `AUDREY_API_KEY` if the REST sidecar is reachable beyond the local process boundary.
- Run `npx audrey dream` on a schedule so consolidation and decay stay current.
- Add application-level encryption, retention, access control, and audit logging for regulated environments.

## Environment Variables

| Variable | Default | Purpose |
|---|---|---|
| `AUDREY_DATA_DIR` | `~/.audrey/data` | SQLite memory store path. Use one per tenant or agent identity for isolation. |
| `AUDREY_AGENT` | `local-agent` | Logical agent identity stamped on writes. |
| `AUDREY_EMBEDDING_PROVIDER` | `local` | `local`, `gemini`, `openai`, or `mock`. Cloud providers require explicit opt-in. |
| `AUDREY_LLM_PROVIDER` | auto | `anthropic`, `openai`, or `mock`. |
| `AUDREY_DEVICE` | `gpu` | Local embedding device (`gpu` or `cpu`). Falls back to CPU if GPU init fails. |
| `AUDREY_PORT` | `7437` | REST sidecar port. |
| `AUDREY_HOST` | `127.0.0.1` | REST sidecar bind address. Set to `0.0.0.0` only with `AUDREY_API_KEY`. |
| `AUDREY_API_KEY` | unset | Bearer token required for non-loopback REST traffic. |
| `AUDREY_ALLOW_NO_AUTH` | `0` | Set to `1` to allow non-loopback bind without an API key. Don't. |
| `AUDREY_ENABLE_ADMIN_TOOLS` | `0` | Set to `1` to enable export, import, and forget routes/tools. Disabled by default. |
| `AUDREY_PROMOTE_ROOTS` | unset | Colon/semicolon-separated extra roots for `audrey promote --yes` writes. By default writes are restricted to `process.cwd()`. |
| `AUDREY_DEBUG` | `0` | Set to `1` to print MCP info logs (server started, warmup completed). Errors always log. |
| `AUDREY_PROFILE` | `0` | Set to `1` to emit per-stage timings via MCP `_meta.diagnostics`. |
| `AUDREY_DISABLE_WARMUP` | `0` | Set to `1` to skip background embedding warmup at MCP boot. |
| `AUDREY_ONNX_VERBOSE` | `0` | Set to `1` to restore ONNX runtime EP-assignment warnings (suppressed by default). |
| `AUDREY_PRAGMA_DEFAULTS` | `1` | Set to `0` to revert SQLite PRAGMA tuning to better-sqlite3 defaults. |
| `AUDREY_CONTEXT_BUDGET_CHARS` | `4000` | Default Memory Capsule character budget. |

## Benchmarks

Audrey ships three benchmark families.

### Performance snapshot

`npm run bench:perf-snapshot` measures encode and hybrid recall latency at multiple corpus sizes against the in-process mock provider. It reports p50/p95/p99 plus machine provenance so the numbers are reproducible and honest about what they cover.

```bash
npm run build
npm run bench:perf-snapshot                                 # default sizes 100, 1000, 5000
node benchmarks/perf-snapshot.js --sizes 1000,10000 --json  # custom shape
```

Sample output from `benchmarks/snapshots/perf-0.22.2.json` (24-core Ryzen 9 7900X3D, Node 25.5.0, mock 64-dim embedding, hybrid recall, limit 5):

| Corpus size | Encode p50 (ms) | Encode p95 (ms) | Recall p50 (ms) | Recall p95 (ms) | Recall p99 (ms) |
|---|---|---|---|---|---|
| 100 | 0.33 | 0.59 | 0.54 | 1.82 | 2.71 |
| 1,000 | 0.31 | 2.15 | 1.57 | 2.36 | 21.18 |
| 5,000 | 0.31 | 1.84 | 2.09 | 3.42 | 16.58 |

These numbers cover Audrey's own pipeline (SQLite + sqlite-vec + hybrid ranking) and exclude embedding-provider cost. Real-world recall p95 with a local 384-dim provider is typically 5-15x higher; with a hosted provider it is dominated by the API round-trip. Run on your own hardware before quoting numbers anywhere.

### Behavioral regression suite

`npm run bench:memory:check` is a release gate. It runs a small set of retrieval and lifecycle scenarios (information extraction, knowledge updates, multi-session reasoning, conflict resolution, privacy boundary, overwrite, delete-and-abstain, semantic/procedural merge) against Audrey and three weak baselines (vector-only, keyword+recency, recent-window) and asserts Audrey doesn't regress. The baseline comparisons exist to catch correctness regressions in retrieval logic, not to make marketing claims.

```bash
npm run bench:memory          # full regression suite (writes JSON + report)
npm run bench:memory:check    # release gate, exits non-zero on regression
```

### GuardBench comparative suite

`npm run bench:guard:check` runs Audrey's local GuardBench comparative suite:
ten pre-action scenarios across Audrey Guard, no-memory, recent-window,
vector-only, and FTS-only adapters. The scenarios cover exact repeated
failures, required procedures, changed file scopes, changed commands,
recovered failures, recall degradation, redaction safety, conflicting
instructions, and noisy stores. It writes
`benchmarks/output/guardbench-summary.json`,
`benchmarks/output/guardbench-manifest.json`, and
`benchmarks/output/guardbench-raw.json`. The emitted manifest, summary, and raw
output shapes are validated by JSON schemas under `benchmarks/schemas/`.

Latest local result in this checkout: 10/10 scenarios passed, 100% prevention
rate, 0% false-block rate, 0 raw secret leaks, 0 published artifact leaks in
the raw-secret sweep, and 3.529ms / 27.78ms
p50/p95 guard latency under the mock-provider methodology.

**Methodology caveats, on purpose.** All numbers above are produced against
the in-process mock 64-dim embedding provider documented in the run's
`provenance` block. They characterize Audrey's controller and SQLite path,
not real-provider end-to-end latency or production false-positive rates. The
100% prevention rate is over the 5 GuardBench scenarios that expect a
`block` decision (the suite is 10 scenarios total, mixed across allow / warn
/ block). Local baseline decision accuracy was: no-memory 10%, recent-window
60%, vector-only 40%, and FTS-only 10%; none of the local baselines passed
the GuardBench decision-plus-evidence contract, which since v1.0.1 requires
the correct decision plus at least one returned evidence id for `block` /
`warn` scenarios (no longer Audrey-specific lineage phrasing — see
`CHANGELOG.md#101---2026-05-15`). External-system numbers for Mem0 and Zep
are explicitly out of scope for this Stage-A artifact; live credentialed
runs land in a v2 paper after raw evidence bundles publish.

```bash
npm run bench:guard
npm run bench:guard:check
npm run bench:guard:manifest
npm run bench:guard:validate
npm run bench:guard:card
npm run bench:guard:bundle
npm run bench:guard:bundle:verify
npm run bench:guard:leaderboard
npm run bench:guard:adapter-registry:validate
npm run bench:guard:adapter-module:validate
npm run bench:guard:adapter-self-test
npm run bench:guard:adapter-self-test:validate
npm run bench:guard:publication:verify
npm run bench:guard:adapter-smoke
npm run bench:guard:adapter-conformance
npm run bench:guard:external:dry-run
npm run bench:guard:mem0 -- --dry-run
npm run bench:guard:zep -- --dry-run
node benchmarks/adapter-self-test.mjs --adapter ./path/to/adapter.mjs
node benchmarks/guardbench.js --adapter ./path/to/adapter.mjs --check
```

External GuardBench adapters are ESM modules that export either `default`,
`adapter`, or `createGuardBenchAdapter()`. The adapter receives scenario seed
data and the proposed action, but the harness withholds `expectedDecision` and
`requiredEvidence` until scoring. Start from
`benchmarks/adapt

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [Evilander](https://github.com/Evilander)
- **Source:** [Evilander/Audrey](https://github.com/Evilander/Audrey)
- **License:** MIT
- **Homepage:** https://www.npmjs.com/package/audrey

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-evilander-audrey
- Seller: https://agentstack.voostack.com/s/evilander
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
