AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified Apache-2.0 Self-run

SodaMem

mcp-sodamem-sodamem · by SodaMem

Agentic memory infra for AI agents

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add mcp-sodamem-sodamem

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-sodamem-sodamem)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3d ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of SodaMem? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Evidence-grounded temporal memory for AI agents.

Every memory can name the turn it came from, and knows when it stopped being true.

[](https://github.com/SodaMem/SodaMem/blob/main/LICENSE) [](https://github.com/SodaMem/SodaMem/blob/main/pyproject.toml) [](https://github.com/SodaMem/SodaMem/tree/main/benchmarking/artifacts/) [](https://github.com/SodaMem/SodaMem/blob/main/benchmarking/README.md#locomo-cat-1-4) [](https://github.com/SodaMem/SodaMem/discussions)

English · 简体中文 · 日本語 · 한국어 · Français · Español · Deutsch · Português

Accuracy against estimated API cost per question. The quadrant that matters is up and to the left.


Benchmark

92.8% (464/500) on LongMemEval.

| | | |---|---| | reader / planner / judge | deepseek-v4-flash | | judge prompts | the LongMemEval benchmark's own evaluate_qa.py templates, byte-identical | | store | longmemeval_s_500_Hobs_entitysubj, 500 users / 235,840 facts |

Every answer and every retrieved memory is published in benchmarking/artifacts/ — 500 answers verbatim, 8,427 evidence rows. Re-grade them with any judge, or hand the retrieved context to your own reader and see what the number does. Neither needs access to anything of ours.

86.88% (1338/1540) on LoCoMo. End-to-end QA accuracy, LLM-as-judge.

| | | |---|---| | reader / planner / judge | deepseek-v4-flash | | judge prompts | the LongMemEval benchmark's own templates, byte-copied | | store | locomo10_Hobs, 10 user stores / 2,905 fact events | | code | a pre-release build — this repository's published history begins at v0.1.0 |

No per-question artifacts are published for LoCoMo — no answers, no retrieved context, no run directory. What is published is the LoCoMo section of benchmarking/README.md: the per-category breakdown, the per-conversation spread, provenance and repro steps.


Quick start

pip install "sodamem[chroma,llm]"
from sodamem import SodaMem
from sodamem.llm import create_provider_from_env      # SODAMEM_LLM_API_KEY etc.
from sodamem.memory.ingest.extractor import FactEventExtractorV2

# Writing needs a model to extract facts with; reading never does.
mem = SodaMem.open("./data", extractor=FactEventExtractorV2(create_provider_from_env()))

mem.ingest(
    [{"role": "user", "content": "Actually I moved from Kauai to Oahu."}],
    user_id="u1", session_id="s1", session_time="2023-05-25",
)

block = mem.build_context("where am I staying?", user_id="u1", token_budget=1000)
print(block.text)        # prompt-ready — zero LLM calls
print(block.citations)   # the exact evidence behind every line of it

SodaMem.open() creates ./data if it isn't there. Only .ingest() needs the extractor — drop that argument for a read-only store and search / build_context work exactly the same.

Nothing about you leaves the machine. No telemetry, no analytics, no callback — the only outbound request the default install ever makes is a one-time download of the 90 MB MiniLM embedding model into ~/.cache/chroma/, and after that it talks to nothing but your disk. Pre-seed that cache and it runs air-gapped.


Why another memory layer

Most memory systems store what you said. The questions that break them are when it stopped being true and where it came from — and those need a data model, not a bigger vector index.

| the question | the usual answer | SodaMem | |---|---|---| | Where did this memory come from? | a similarity score and some metadata | FactEvent → SourceSpan → RawTurn, a foreign-key chain down to the exact turn | | The user changed their mind — now what? | overwrite; the old value is gone | ADD-only plus a SUPERSEDES edge; the old version closes with a valid_until and stays readable | | "I moved to Chicago last year" vs "I move next year" | one timestamp | four time axes: occurred / valid / said / stored | | What does one retrieval cost? | an LLM call per retrieval | build_context makes zero, and returns a prompt-ready block with citations | | Same query twice — same answer? | depends on the model's sampling | deterministic fusion: same store, same query, same result | | Why did it forget X? | no answer | /v1/events records every add, supersede and delete, with its reason |

Each row is expanded below, and each one is something you can check in this repository rather than take on faith.

Every memory carries its receipt

A retrieved memory is not a floating string. It points at the turn that produced it:

evidence_id  = ev_fact:fact_6ada707b…
support      = "Can you recommend a good beach on Oahu that's not too crowded?"
predicate    = user wants a not-too-crowded beach on Oahu
entities     = location=Oahu | occasion=birthday
source       = session_40 / turn_10          ← the exact turn, not "some chat"
date         = 2023-05-25

FactEvent → SourceSpan → RawTurn is a foreign-key chain, not a similarity score. When a user asks "why do you think that about me?" there is an answer. When compliance asks where a stored fact came from, there is a row.

Four time axes, not one timestamp

| field | question it answers | |---|---| | occurred_start / occurred_end | when the event happened | | valid_from / valid_until | when the fact was true | | document_time | when the user said it | | created_at | when we stored it |

One timestamp cannot separate "I moved to Chicago last year" from "I will move next year", and cannot express a fact that stopped being true.

Corrections are ADD-only: a new version plus a SUPERSEDES edge, never an in-place rewrite. PATCH /v1/memories/{id} closes the old version with a valid_until and leaves it readable — that is the whole difference from DELETE.

Two retrieval tiers, and the cheap one is genuinely free

| tier | LLM calls | for | |---|---|---| | search / build_context | zero | the default path: deterministic BM25 + vector + entity fusion | | answer | planner loop | hard multi-hop questions worth the tokens |

build_context returns a prompt-ready block with citations and makes no model call. Most systems hand you a list of records and leave the assembly — and the token budgeting, and the dedup — to you.

There is a third, in-between tier: build_context(organizer=...) runs an LLM-backed organizer (value-board, enumeration-sweep) over the retrieved set for questions like "list every X you know about me". It is Python-only on purpose — /v1/context never accepts one, so the zero-LLM guarantee on that route cannot be flipped by a request parameter.

Retrieval you can audit

Same query, same store, same result, every time. /v1/events records every add, supersede and delete with its reason, so "why did the agent forget X" is answerable after the fact instead of a shrug.


Install

| extra | what it adds | |---|---| | (base) | data model, storage, BM25 retrieval, ingest — four dependencies, none heavy | | chroma | vector search + the local ONNX embedder (SodaMem.open() needs this) | | llm | OpenAI-compatible providers (OpenAI / DeepSeek / Gemini wire format) | | anthropic | the Anthropic provider (its own SDK) | | answer | the planner + reader answer path | | server | the HTTP service (FastAPI + uvicorn — three packages, deliberately) | | mcp | MCP server surface |

Base install pulls pydantic, numpy, rank-bm25, python-dateutil — and a CI gate fails the build if that list grows by accident.


Use it from anywhere

HTTPadd / search / context / answer, plus batch write, supersede, events, metrics, token usage:

curl -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  localhost:8000/v1/context \
  -d '{"user_id":"u1","query":"what do they prefer?","token_budget":1000}'

/v1/context and /v1/search both take a JSON body; /v1/context also answers a plain GET with query params, since it is a pure read.

SDKs — TypeScript over HTTP (sdk-ts/, zero runtime dependencies, ESM + CJS). Python talks to the library directly — import sodamem and you are already past the network.

Agent frameworks — LangGraph, CrewAI, OpenAI Agents SDK, Vercel AI SDK. Scope is bound when you construct the tools and never appears in the schema the model sees: a user_id the model can choose is a user_id it can hallucinate.

MCP — 8 tools, including entity_timeline (one entity's history in order, each item still pointing at its source) and explore_memory (walk the graph outward). Six are reads and always available; the two that mutate (add_memories, delete_memory) appear only under SODAMEM_MCP_ALLOW_WRITE=true, which sodamem install writes for you into the client config it generates.

Web console — browse and inspect memories per tenant, shipped in the image.


Coding tools

sodamem daemon ensure            # the one process that owns the stores
sodamem install claude-code      # wire a client to it

Every client gets the MCP tool surface. Four also get hooks, so memory is recalled and retained without the model having to decide to call a tool — which in a coding session it mostly doesn't, because it's busy reading files.

What hooks can do is not uniform, because the hook systems aren't. This is what each client actually supports, and sodamem clients prints the same thing:

| Client | Recall | Retain | |---|---|---| | Claude Code | every prompt | every turn + session end | | GitHub Copilot CLI | every prompt | every turn | | Cursor | session start (project brief) | — | | Codex CLI | session start (project brief) | — | | Claude Desktop, VS Code, Windsurf, Zed, OpenCode | MCP tools only | MCP tools only |

Cursor's beforeSubmitPrompt can read a prompt but cannot inject anything (its docs list exactly three events that can, and that isn't one), and neither Cursor nor Codex hands a hook a transcript path — so there is nothing for a retain hook to read. Those two get a project brief at session start and write through the add_memories tool. We don't install a hook that can only ever do nothing.

Three things worth knowing before you run it:

One daemon, many editors. Per-user stores are SQLite without WAL, so exactly one process may open them (ADR 0001 §2). install therefore points every client at a running service by default rather than letting each spawn its own — and if you deliberately choose a local store (--local-store), a second client now refuses to start instead of quietly corrupting the first one's data.

Memories are scoped to the repo. install derives a project_id from the git root (a git worktree resolves to its parent repo, so one branch per task is not one memory bank per task). Narrowing, not partitioning: anything you told SodaMem outside a project still surfaces inside every project, and dropping the key answers "how did I fix this in the other repo?".

Retain needs extraction credentials. Recall is zero-LLM and works without them; storing facts does not. sodamem daemon ensure says so up front rather than accepting every write and failing the job afterwards.

sodamem install claude-code --dry-run      # print what would change
sodamem install cursor vscode zed          # several at once
sodamem daemon status                      # what is actually answering

Existing config is merged, not replaced — other MCP servers, other settings and hand-written TOML comments survive — and the first write of any file leaves a .sodamem-backup beside it.

Self-hosting

One command:

cp .env.example .env      # then set SODAMEM_API_KEY
docker compose up -d

That builds the image, starts the server on http://localhost:8000, serves the web console at http://localhost:8000/console, and persists all data in a named Docker volume (sodamem-data, mounted at /data inside the container) — nothing is written to the host filesystem directly, and nothing survives only in the container's writable layer.

The console is compiled inside the image (a dedicated console-builder stage), so nothing on the host needs Node installed. Running the server outside Docker is different: console/dist won't exist until you run npm install && npm run build in console/, and until then the API starts normally and just logs that the console isn't mounted.

Auth is on by default. docker-compose.yml never sets SODAMEM_AUTH_DISABLED — the server refuses to start if SODAMEM_API_KEY is unset (see server/settings.py), so there is no accidentally-open deployment. Set the key in .env before the first docker compose up.

Every other knob (SODAMEM_LLM_PROVIDER, SODAMEM_STORE_CACHE_MAX, SODAMEM_CORS_ORIGINS, ...) is documented with defaults in .env.example.

Run exactly one worker. --workers 1 is a correctness constraint, not a throughput setting: per-user stores are SQLite databases opened without WAL, and two processes writing the same user's store corrupt it. The shipped CMD states it explicitly, and the server takes an exclusive lock on its data root at startup — a second process pointed at the same directory refuses to start with data_root_locked rather than quietly corrupting data. Horizontal scaling needs an external job store first (docs/adr/0001-control-plane-db.md).

Calling it

# liveness — unauthenticated, touches no store
curl http://localhost:8000/health
# {"status":"ok","version":"0.0.1","schema_version":1,"auth":"enabled"}

# a real endpoint needs the API key (Authorization: Bearer, or X-API-Key)
curl http://localhost:8000/v1/search \
  -X POST -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SODAMEM_API_KEY" \
  -d '{"user_id":"alice","query":"favorite color"}'

Full route list and request/response shapes are in server/models.py and served live at /docs (Swagger UI) once the container is up.

Operating it

/v1/admin/* answers the questions that otherwise need a shell inside the container. The web console's Ops page is the same data with a UI.

# effective configuration — every secret reported as set/not-set, never masked
curl -H "Authorization: Bearer $SODAMEM_API_KEY" localhost:8000/v1/admin/config

# mint a named key; the plaintext is returned ONCE and is not recoverable
curl -X POST -H "Authorization: Bearer $SODAMEM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name":"ci-pipeline"}' localhost:8000/v1/admin/keys

# who called what, most recent first (rolling window, not an archive)
curl -H "Authorization: Bearer $SODAMEM_API_KEY" localhost:8000/v1/admin/requests

# disk + workload shape
curl -H "Authorization: Bearer $SODAMEM_API_KEY" localhost:8000/v1/admin/stats

Named keys exist for attribution, not isolation: every request records which key made it, so "who is hammering /v1/search" has an answer. There are no roles or per-key scopes — any live key can read ops data and manage other keys. SODAMEM_API_KEY keeps working exactly as before and cannot be revoked through the API, which makes it the way back in if every named key is revoked.

Latency and cost

Both are instruments, not numbers we ask you to take on faith.

# per-route latency percentiles over this process's recent requests
curl -H "Authorization: Bea

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [SodaMem](https://github.com/SodaMem)
- **Source:** [SodaMem/SodaMem](https://github.com/SodaMem/SodaMem)
- **License:** Apache-2.0
- **Homepage:** https://sodamem.com

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.