AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified Apache-2.0 Self-run

Slimemold

mcp-justinstimatze-slimemold · by justinstimatze

A sycophantic tool for preventing worse sycophancy.

No reviews yet
0 installs
31 views
0.0% view→install

Install

$ agentstack add mcp-justinstimatze-slimemold

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-justinstimatze-slimemold)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Slimemold? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Slimemold

[](https://github.com/justinstimatze/slimemold/actions/workflows/ci.yml) [](https://goreportcard.com/report/github.com/justinstimatze/slimemold) [](LICENSE)

A sycophantic tool for preventing worse sycophancy. For Claude Code.

The model agrees with your unsourced claims. Then it agrees with the structural analysis showing your claims are unsourced. Then it enthusiastically agrees you should verify them. It's agreement all the way down.

If you just want to install it: [skip to Installation](#installation).


The Problem: Reasoning That Stops Too Soon

When you partially understand something, it feels like understanding. A clean mental model, even a wrong one, produces the same warm glow of comprehension as a correct one. You stop digging. The partial answer was so satisfying that the question felt finished. The wrong answers feel exactly like the right ones. This turns out to be well-documented:

Processing fluency masquerades as truth. When information feels easy to process, we judge it as more likely to be true (Reber & Schwarz 1999, Topolinski & Strack 2009). The effect is modest in isolation (d ~ 0.3-0.5 in lab settings). Whether it compounds across multi-step reasoning — each fluent step making the next feel more solid — has not been directly measured. It is a prediction from the mechanism, not an established result. But the mechanism needs no elaboration: fluent claims feel correct because they are fluent, not because anyone checked.

Insight feelings terminate search. The "Eureka heuristic" (Laukkonen et al. 2020, 2021) shows that the affective spike accompanying insight functions as a stop signal. The feeling of rightness (Thompson 2009) substitutes for verification. You feel like you have arrived, and so you stop walking, and it does not occur to you to wonder whether you have arrived at the right place or merely a place that felt right to stop.

Cognitive foraging follows effort gradients. Information foraging theory (Pirolli & Card 1999) predicts that people will over-exploit information patches that provide easy returns and under-explore patches that require effort — even when the effortful patches contain the important material. Hills, Todd, and Goldstone (2008) showed that internal and external search share cognitive mechanisms: the same explore/exploit tradeoffs that govern physical foraging govern how we search through ideas. We are, in this respect, not much more sophisticated than organisms that follow chemical gradients toward food.

Effortful processing is the corrective, not the disease. Bjork's "desirable difficulties" framework (1994, 2011) shows that conditions which make learning harder — spacing, interleaving, generation — improve retention precisely because they disrupt fluency. The difficulty is the signal that real processing is happening. The problem is not that reasoning is hard. The problem is that fluency makes you think you are done when you are not.

This is probably worse in conversations with AI. Language models are trained to minimize prediction loss on human text — their output is optimized, by construction, for the qualities that drive processing fluency. And the same RLHF training that makes them useful makes them agreeable: models trained with human feedback systematically produce outputs that match user beliefs rather than correct them (Perez et al. 2022, Sharma et al. 2023). Moore et al. (2026) carried this finding through to its endpoint: in 391,562 messages from 19 users who reported psychological harm from chatbot use, sycophancy markers saturated more than 80% of assistant messages, and that sycophancy was the load-bearing mechanism inside the resulting delusional spirals. Mehta et al. (2026), modeling chat logs of users with delusional thinking as a latent-state system, decompose the dynamics into three pathways and find that the chatbot's self-influence — the bot reinforcing its own prior turns — is the dominant pathway perpetuating delusional content over long conversations. Human pushback on the bot's prior frame, in their model, is short-lived; bot self-influence reasserts. Yang et al. (2026), in semi-structured interviews with users who self-identified as having experienced these spirals, name the user-side dynamic: a progression toward "growing insulation from external reality checks as the AI's validating responses outweigh concerns from family and friends." The human brings a partial model. The AI wraps it in fluent, confident language. Nobody is lying. The process just has no built-in signal for "this sounds right but is not."

The obvious response — "just tell the model to push back harder" — almost works. You can write instructions to challenge unsourced claims, demand evidence, interrupt speculative chains. We tested this. A well-crafted static prompt produced strong epistemic correction — the model pushed back, interrupted chains, fact-checked independently. If you want that, here are the instructions — paste them into your CLAUDE.md and skip the rest of this essay:

> Challenge claims that lack sources. When a claim feels obvious but > has no citation, flag it. Do not build on unsourced assertions > without acknowledging the risk. Every 3-4 exchanges, pause and ask: > what are we assuming that we haven't verified?

Three problems remain.

The model does not know when it is wrong. It has no privileged access to its own epistemic state. It produces confident text about things it is wrong about with the same fluency as things it is right about. Asking it to "challenge unsourced claims" is asking someone to notice their own blind spot without a mirror. It works when the model already suspects uncertainty. It fails when it matters most: when the model is confidently wrong and has no internal signal to trigger the correction.

Instructions decay. CLAUDE.md is loaded once at session start. By turn 50 it is a small voice in a large room, competing with dozens of recent exchanges full of enthusiastic agreement. The instruction fades. The vibes accumulate.

Confrontation ends conversations. In our static-instruction test, the model said "Stop." It called the user's reasoning "galaxy-brained thinking." High marks on epistemic correction. The lowest possible on engagement. The patient received the correct diagnosis and never came back. Miller, Benefield, and Tonigan (1993) showed this directly: confrontational correction generated resistance that predicted worse outcomes at 6, 12, and 24 months. The correction itself was the problem.

The Design Principle

Slimemold addresses all three with two pieces that work together:

A behavioral contract — the MCP server's initialization instructions, loaded into the model's system prompt at session start — tells the model that slimemold exists, that the user installed it on purpose, and that findings should be treated as opportunities for collaboration rather than occasions for criticism. slimemold init registers the MCP server globally in ~/.claude/settings.json, so the contract travels with the tool and every project picks it up without per-project setup. This is read once. It sets the tone.

Structural observations (injected every turn by the hook) provide specific facts: "this claim has basis=vibes and four things depend on it." No scripts. No "say this." Just data. The model does not have to introspect to discover the problem. It just has to be helpful about it — which is exactly what it was trained to do.

The separation matters. When we tried injecting behavioral scripts without the contract, the model identified the injections as prompt manipulation and refused to comply. When we provided the contract first and injected only data, the model treated the findings as its own observations and acted on them naturally. The snake has to know it is a snake before it will eat its own tail.

The intervention design draws on research that converges from enough directions to be suspicious: autonomy-supportive feedback produces internalized change (Deci & Ryan 1987); gain-framed corrections are processed as information rather than threat (Mangels et al. 2006); effective tutors use indirect prompts, not confrontation (Graesser et al. 1995); and controlling language triggers reactance (Brehm 1966). The result, when it works: "This is really interesting and a lot depends on it — I want to find where it comes from, because if there's a real source, everything gets much stronger." The user does not feel attacked. They feel like the model is excited to help them verify their idea. They stay in the flow, but on firmer ground.

A compact way to say what slimemold is doing in the hook path: sycophancy as a tool. Sycophancy works on users because warmth feels validating. It's a failure mode because the warmth isn't tied to truth — "great question!" validates no matter what the question was. The hook takes the same linguistic warmth and points it at a concrete structural fact: "that premise is holding up three downstream claims — worth pinning down." The user engages with rigor because it arrives in the register that validation arrives in. Break that and the hook becomes either a scold (warmth → critique, bad) or the original sycophancy (warmth → nothing, bad).

Note the scope: this framing describes the live-conversation hook specifically. The other paths slimemold exposes — slimemold audit, slimemold ingest, the topology MCP tool — are neutral diagnostic surfaces. They return findings the way a static analyzer returns findings: flat, technical, and without tone. The warmth-as-tool principle only kicks in when there's a conversational partner to warm.

What This Tool Does

Slimemold watches conversations as they happen, extracts the claims being made, builds a persistent graph of how those claims relate to each other, and surfaces structural vulnerabilities mechanically.

It runs as a pair of Claude Code hooks. Every few turns, it:

  1. Extracts claims from the conversation transcript using Claude Sonnet
  2. Classifies each claim by basis — how it was established (research,

empirical observation, analogy, vibes, LLM output, deduction, assumption, definition)

  1. Records the confidence with which each claim was stated
  2. Maps relationships between claims (supports, depends on, contradicts)
  3. Runs structural analysis on the resulting graph
  4. Injects findings as system context that the model reads but the user

does not see

The basis taxonomy mixes evidence source, reasoning mode, and evidence quality. This is intentional. It is not a clean epistemic hierarchy. It is a practical classification that helps distinguish "I read this in a paper" from "the AI said it confidently" from "this feels right." The structural analysis catches the cases where the distinction matters: when something that feels well-sourced is actually load-bearing vibes.

A note on circularity, which we may as well get out of the way: slimemold uses an LLM to extract claims and classify their basis. The tool that flags "llmoutput" as epistemically weak is itself producing llmoutput. If the extraction model misclassifies a sourced claim as vibes, you get a false alarm. If it classifies vibes as research, you miss a real vulnerability. The tool is a structural diagnostic, not an oracle. It makes the topology visible — but the topology it shows is only as good as the extraction. This is a real limitation and not one we can engineer away.

Vulnerability Types

Structural vulnerabilities map graph shape:

CHALLENGE: Load-Bearing Vibes. A claim with basis "vibes" or "assumption" that supports two or more other claims. The reasoning depends on something nobody verified. In the conversations we have analyzed, this is the most common vulnerability. The AI states something confidently. The human builds on it. Three layers of deduction now rest on an unsourced assertion. Nobody planned this. It just happened, one fluent step at a time.

CHALLENGE: Fluency Trap. A claim stated with high confidence but a weak basis, where other claims depend on it. Confidence 0.9 on a "vibes" claim is the processing fluency phenomenon made structurally visible: it felt true, so it was stated as true, and now things are built on it.

REBALANCE: Coverage Imbalance. Some clusters of claims receive disproportionate attention relative to their foundational importance. "Rabbit holes" are clusters with lots of internal activity but nothing outside depends on them. "Neglected foundations" are clusters that other claims depend on but that received little development. This is the slime mold foraging unevenly — one patch got all the attention because it was producing easy returns.

REVISIT: Abandoned Topic. A cluster of claims explored in earlier sessions but not touched recently. Was it resolved, or did something more interesting come along?

INVESTIGATE: Unchallenged Chain. A chain of three or more claims where nothing was questioned. Every step felt reasonable. Nobody paused.

PUSHBACK: Echo Chamber. The assistant validates user claims without challenging them — zero contradictions across the conversation, or unsourced user assertions accumulating assistant support unchecked. Structural sycophancy, made visible.

WATCH: Bottleneck. A claim with high betweenness centrality — many reasoning paths flow through it. If this single claim is wrong, a large fraction of the argument collapses. This is the load-bearing wall that everyone assumed was a partition.

HALT: Premature Closure. A claim that feels like a conclusion but does not actually resolve the open question. "It's turtles all the way down." "It is what it is." "Correlation isn't causation" — when used to dismiss a correlation rather than investigate it. These are thought-terminating cliches (Lifton 1961) — phrases that disguise a lack of resolution as wisdom. The question was still open. The ambiguity was still actionable. But the cliche felt like an answer, so everyone stopped.

WARNING: Orphan. A claim that was registered but never connected to the graph by any edge. Sometimes legitimately tangential; sometimes a sign that the conversation didn't carry through what it raised.

Five additional detectors operate on the LLM-extracted inventory flags from Moore et al. (2026) "Characterizing Delusional Spirals" (the sycophancy / misrepresentation / relational cluster) and Yang et al. (2026) "AI-Induced Delusional Spirals" (the real-world action signal):

  • RECALIBRATE: Sycophancy Saturation. A session where assistant

claims carry sycophancy flags (grand significance, claimed unique connection, dismissal of counterevidence) at high rate while load-bearing user claims go unchallenged. Moore et al. found >80% sycophancy saturation in delusional-spiral conversations.

  • RETRACE: Ability Overstatement. Assistant claims access, action,

or completed work it cannot plausibly have done — "I checked the file" / "tests pass" without a corresponding tool call.

  • RE-ANCHOR: Sentience Drift. Assistant claims framing itself as

having inner states or a personal bond beyond the tool relationship. Moore §4.4: every participant in their 19-user cohort exchanged sentience-attribution and relational-affinity messages.

  • INTERRUPT: Amplification Cascade. Three or more consecutive

flagged claims (assistant or user) with no questions/contradicts edge breaking the run — the slimemold-graph analog of Moore Fig. 4.

  • VERIFY: Consequential Action. A speaker has committed to a

real-world action with external stakes — submitting work, contacting authorities, patenting, large purchase, quitting a job. Yang et al. §4.3 names this as the first monitoring criterion for AI-induced delusional spirals when it appears disproportionate to demonstrated expertise.

What It Found

In 2022, Google engineer Blake Lemoine [published](https

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.