Install
$ agentstack add mcp-udjin-labs-mnemostack ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ● Filesystem access Used
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
mnemostack
[](https://pypi.org/project/mnemostack/) [](https://pypi.org/project/mnemostack/) [](https://github.com/udjin-labs/mnemostack/actions/workflows/ci.yml) [](https://opensource.org/licenses/Apache-2.0)
> Self-hosted hybrid memory & retrieval for AI apps.
mnemostack is a durable retrieval layer over your own Qdrant (and optional Memgraph): semantic, keyword (BM25), temporal, and graph recall, fused with Reciprocal Rank Fusion and refined by an 8-stage ranking pipeline — with payload filters for multi-tenant isolation, optional LLM answer synthesis (confidence + citations), and an ingest path that enriches and projects structured fields. One recall(query) call, usable as a Python library, an HTTP service, or an MCP server.
Flagship use case — durable memory for AI agents. Long-running agents hit the same wall: context gets compacted, sessions restart, useful decisions disappear, and the next run pays the re-orientation tax again. mnemostack gives them a persistent memory layer to query when the context window is not enough — durable, searchable, scoped, and explainable, not just embedded and hoped for.
The same engine backs other retrieval-heavy work: RAG over mixed corpora, multi-tenant or per-user knowledge stores, and time-aware search backends — anywhere pure vector similarity falls short on its own.
Status: Actively developed — public API is stable; new functionality lands additively in minor releases. Breaking changes are rare and called out in [CHANGELOG.md](CHANGELOG.md).
Quickstart: agent memory over MCP
The fastest on-ramp is the MCP server — it gives Claude Desktop, Claude Code, Cursor, ChatGPT, or another MCP-capable agent durable memory in a few commands. Building an app instead of wiring up an agent? Use the [HTTP API or the Python library](#three-ways-to-use-mnemostack) over the same collection.
1. Install
pip install 'mnemostack[mcp]'
Run a local Qdrant for the vector store:
docker run -p 6333:6333 qdrant/qdrant:latest
Optional: run Memgraph for graph-backed memory:
docker run -p 7687:7687 memgraph/memgraph:latest
2. Start the MCP server
export GEMINI_API_KEY=your-key-here
mnemostack mcp-serve --provider gemini --collection my-memory
Claude Desktop config example:
{
"mcpServers": {
"mnemostack": {
"command": "mnemostack",
"args": ["mcp-serve", "--provider", "gemini", "--collection", "my-memory"],
"env": {
"GEMINI_API_KEY": "your-key-here"
}
}
}
}
Claude will then be able to call mnemostack_search, mnemostack_answer, and graph tools.
3. Store memory
Index a folder of notes, docs, transcripts, or project context:
mnemostack index ./my-notes/ --provider gemini --collection my-memory --recreate
--recreate drops the existing collection, so it asks for confirmation first; pass --yes to skip the prompt (required in scripts/CI — non-interactive runs without it exit with code 2).
For a running app or assistant, use the streaming Ingestor API shown below to store messages as they arrive.
4. Recall memory
From an agent, ask a memory-style question and let the MCP tools retrieve the right facts.
From the shell, test the same collection directly:
mnemostack search "what did we decide about auth" --provider gemini --collection my-memory
mnemostack answer "what did we decide about auth" --provider gemini --collection my-memory
Why hybrid memory?
Vector search answers "what sounds similar?" Real retrieval over a growing corpus needs to answer "what actually matters for this query?" — which takes exact matches, semantic similarity, relationship tracking, recency, user/project scope, and feedback from past recalls. Mnemostack uses hybrid retrieval so recall is reliable instead of embedding roulette. (Agent memory is the most demanding version of this problem, which is why it's the flagship use case.)
Use cases
Agent and chatbot memory (flagship):
- Long-running coding agents that need to survive compaction and session restarts.
- Chat assistants and conversational bots that remember a user's earlier messages, preferences, and decisions across sessions — scope each user's memory with payload
filtersso one user never sees another's history. - Personal assistant memory for preferences, recurring tasks, and long-horizon context.
- Multi-agent context sharing through one durable memory backend.
- Session compaction recovery when the useful details no longer fit in the prompt.
Beyond agents — the same engine as a retrieval backend:
- A searchable knowledge base in memory: ingest your docs/notes/FAQ once, then serve hybrid search or grounded
answer()(with confidence and source citations) over it from the CLI, HTTP, or library. - RAG over mixed corpora (code, docs, transcripts) where exact-token and temporal recall beat pure vector similarity.
- Multi-tenant or per-user knowledge stores — payload
filtersisolate each tenant's data inside every retriever, or enable service-key auth (serve --auth) for a hard, key-resolved tenant boundary with optional per-tenant quotas (see the [HTTP API](#http-server-optional)). - Time-aware search and knowledge bases — "what changed last week", point-in-time graph facts, freshness-weighted ranking.
- Team or project knowledge recall across docs, notes, tickets, and chat history.
Three ways to use Mnemostack
- MCP server — for agent users. Start
mnemostack mcp-serve, connect your agent, and use memory tools from the chat/runtime you already use. - HTTP API — for app developers. Run
mnemostack serveand call/recall,/answer,/feedback,/health,/metrics, or/docsfrom any language. - Python SDK — for library users. Compose retrievers, stores, rerankers, graph tools, and the streaming ingest API inside your own Python application.
Architecture
Mental model
Think of it as a storage hierarchy for agent memory:
- Context window = RAM. Fast, limited (typically 100K–200K tokens for many agent models; larger windows exist, but usable working context is often much smaller after tools, instructions, MCP output, and context rot — ~45K usable tokens on a 200K window is a realistic working number in long-running agent sessions. Clears on session restart.
- mnemostack corpus = Disk. Persistent, searchable, grows forever — every fact the agent has ever seen, queryable on demand.
recall(query)= page fault handler. When the agent needs something that isn't in the current context, it pulls the exact fact from storage with a single hybrid query — not a grep, not a reload of the whole corpus.
The practical effect: you stop re-explaining your project to the agent after every /compact. You stop losing momentum to the re-orientation tax that shows up in any agent with session compaction. mnemostack solves it at the library level, not tied to any single agent runtime.
How it works, in one paragraph
On each recall(query): the configured retrievers (Vector and Temporal by default, with BM25 and Memgraph when configured) run in parallel and return ranked lists. Reciprocal Rank Fusion merges them. The optional 8-stage pipeline can reweight results using query classification, exact-token rescue, gravity/hub dampening, freshness, inhibition-of-return, curiosity boosts, Q-learning weights supplied through its state store, and graph resurrection. An optional LLM reranker does a final ordering pass. You get a list of RecallResult with source, score, and provenance — ready to hand to a model.
Where mnemostack fits
Most memory tools in the agent ecosystem pick one axis and optimize for it: simple vector similarity for RAG, framework-bound memory tied to a specific agent library, platform-level runtimes with audit and compliance features, or CLI wrappers over a single vendor's session store. Each makes sense for its scope.
mnemostack takes a different slice: it is a recall quality layer, offered as a plain Python package. Four retrievers (Vector + BM25 + Memgraph + Temporal), RRF fusion, an 8-stage pipeline, and an optional LLM reranker — composed to handle mixed workloads on the same corpus: exact-token lookups, semantic queries, temporal questions, and multi-hop reasoning, without forcing you to choose one mode over another.
We are not a replacement for your agent framework and not a full platform runtime. We are the piece that actually finds the right fact in a growing corpus. Drop mnemostack into your own Python agent or application, or let a higher-level service call recall() over a plain function boundary. The retrievers, pipeline, and reranker are individually composable — take only the parts you need.
Design
See [ARCHITECTURE.md](ARCHITECTURE.md) for detailed design: pipeline stages, Qdrant schema, Memgraph temporal model, consolidation runtime, MCP tools.
Storage, index, and retrieval layers
- Storage/index layer: Qdrant stores vector points and payloads; BM25 indexes exact-token corpora; Memgraph stores temporal graph facts; the Temporal retriever handles time-aware vector recall.
- Fusion layer: Reciprocal Rank Fusion merges ranked lists from Vector, BM25, Memgraph, and Temporal retrievers, with optional static or adaptive weights.
- Recall pipeline: the 8-stage pipeline can classify the query, rescue exact tokens, dampen gravity/hubs, blend freshness, apply inhibition-of-return, add curiosity boosts, use Q-learning state, and resurrect graph-linked memories.
- Feedback loop: HTTP and MCP recall can apply existing state; explicit
/feedbackormnemostack feedbackupdates usefulness signals without silently training on every response. - Inference layer: optional LLM reranking and answer generation sit on top of recall, so retrieval still works when the LLM is unavailable.
Pipeline state
The 8-stage pipeline can use a small state store between calls (Q-learning weights, inhibition-of-return history, per-document gravity/hub counters). FileStateStore(path) persists it to a JSON file. HTTP recall applies existing state and can record inhibition-of-return exposure with --auto-record-ior; Q-learning updates only through explicit /feedback calls. CLI/MCP recall still apply existing state but do not collect feedback automatically. For deterministic benchmarks, call build_full_pipeline(enable_stateful_stages=False) so IoR/Q-learning/curiosity state cannot affect scores. For multi-process servers, implement your own StateStore (three methods: get(), set(), update()) backed by Redis or your database.
Graceful degradation
Any retriever can fail (Memgraph down, Qdrant unreachable, BM25 corpus empty). Recaller logs and continues with the remaining sources. The LLM reranker is wrapped in try/except by convention — if the LLM is rate-limited, the pre-rerank order is returned. This is deliberate: a memory stack that goes dark because one component hiccuped is worse than a slightly degraded one.
One exception: query expansion runs before retrieval, so a misconfigured expansion step (query_expansion=True without an expansion_llm, or a provider error inside it) surfaces as an error instead of degrading silently — see [ARCHITECTURE.md](ARCHITECTURE.md) for the full fail-open contract. Degradations themselves are visible, not silent: every HTTP/MCP response carries degraded tags, and the full per-retriever trace is available opt-in via include_trace.
Comparison and benchmarks
On LoCoMo, Mnemostack reaches 82.9% strict accuracy in our evaluation setup. The table below includes our baseline runs and externally reported numbers for context. Results depend on dataset version, configuration, judge model, scoring rules, and query type. Treat externally reported numbers as directional unless they were run with the same harness and settings.
Benchmarks
Full LoCoMo runs use the official SNAP-Research dataset (10 samples / 1986 QA) from a clean state. Across the tables below: Strict = exact match, Combined = strict + partial. Counts in cells are correct / total.
Some LoCoMo cat_5 questions have empty ground-truth answers. Under the current scorer, these are counted as correct because there is no expected answer to match. To avoid overstating recall quality, we also report signal-only scores with those questions removed. Signal-only scores are computed on the 1,540 questions with non-empty ground-truth answers.
LoCoMo, current judge (gemini-3-flash-preview)
| Run | Strict (full) | Combined (full) | Strict (signal-only) | Combined (signal-only) | | --- | --- | --- | --- | --- | | Baseline v0.3.0 (Vector + BM25 + 8-stage pipeline) | 76.7% (1524 / 1986) | 88.1% (1750 / 1986) | 70.0% (1078 / 1540) | 84.7% (1304 / 1540) | | Retrieval improvements (window_size=3, query expansion, top-K 25) | 82.5% (1639 / 1986) | 92.2% (1832 / 1986) | 77.5% (1193 / 1540) | 90.0% (1386 / 1540) | | v0.4.5 + photo captions (same config as above) | 82.9% (1647 / 1986) | 92.7% (1842 / 1986) | 78.0% (1201 / 1540) | 90.6% (1396 / 1540) |
> Honest numbers disclaimer. (full) is the headline aggregate across all 1986 questions, the format vendors typically report — some publish only their strongest sub-category, we publish the full aggregate because it's what actually predicts behavior on mixed workloads. (signal-only) strips the cat_5 auto-pass artifact described above, so what you read there is the real recall quality on questions that have a ground-truth answer.
Per-category breakdown (v0.4.5 + photo captions run):
| Category | Strict | Combined | | --- | --- | --- | | cat_1 single-hop lists | 51.4% | 88.3% | | cat_2 temporal | 79.8% | 85.7% | | cat_3 open-domain reasoning | 62.5% | 79.2% | | cat_4 multi-hop reasoning | 88.0% | 94.6% | | cat_5 adversarial open-domain | 100.0% | 100.0% |
Notes:
- Judge model matters:
gemini-3-flash-previewis more accurate than the previous Gemini Flash judge on synonyms, partial matches, and empty ground truth. cat_5questions have empty ground truth in this new run and are auto-scored as correct by the benchmark harness. That makes the newcat_5strict score (446 / 446, 100.0%) useful for aggregate harness accounting, but not directly comparable to the historicalcat_5strict score (89.7%) from the older adversarial-question evaluation.- Pipeline: Vector retrieval with Gemini embeddings + BM25 + RRF + 8-stage reranking pipeline. The LLM reranker is not part of the benchmark loop (it is a runtime/server feature), so
rerank_modedoes not affect these numbers. - The v0.4.5 run additionally ingests the photo captions (
blip_caption) that LoCoMo attaches to image-sharing turns — 697 of the 1540 signal questions cite image turns as evidence, and earlier runs silently dropped that content. Answer prompts also show the time of day of each memory since v0.4.5.
Historical LoCoMo results (gemini-2.5-flash judge)
| Metric | First full run | mnemostack 0.2.1 | | --- | --- | --- | | Strict | 66.4% (1319 / 1986) | 67.8% (1346 / 1986) | | Partial | 12.8% (254 / 1986) | 12.6% (250 / 1986) | | Wrong | 20.8% (413 / 1986) | 19.6% (390 / 1986) | | Combined | 79.2% (1573 / 1986) | 80.4% (1596 / 1986) |
By question category (combined not tracked for the first full run):
| Category | First run Strict | 0.2.1 Strict | 0.2.1 Combined | Δ Strict | | --- | --- | --- | --- | --- | | cat_1 single-hop lists | 34.8% | 34.4% | 74.1% | −0.4pp | | cat_2 temporal | 64.5% | 69.8% | 77.9% | +5.3pp | | cat_3 open-domain reasoning | 31.2% | 41.7% | 49.0% | +10.5pp | | cat_4 multi-hop reasoning | 69.2% | 69.6% | 82.0% | +0.4pp | | cat_5 adversarial open-domain | 90.1% | 89.7% | 89.7% | −0.4pp |
Last historical run: 2026-04-27, mnemostack 0.2.1, same dataset, judged by gemini-2.5-flash.
#
…
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: udjin-labs
- Source: udjin-labs/mnemostack
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.