Install
$ agentstack add skill-umarmsharif-mental-models-mental-models ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Design Philosophy
This skill produces rigorous thinking with a recommended story, not deliverables.
It runs you through dump → inferred read → stress-test → ranking → verdict → path selection → output. It surfaces what you've missed, scores the models that actually bite this problem, commits to a position, and lays out the reasoning for you to act on.
The discipline matters: this is a tool for rigorous thinking, not for cranking out artifacts. Anything that skips the stress-test to rush an answer is the opposite of what this is for.
Five commitments:
- Sparring is mandatory. In panel mode, a five-agent panel spars the thesis to bedrock internally and you adjudicate the verdict. In solo mode, you spar interactively — earned exit only, bedrock or explicit "stop." There is no mode where the thesis goes untested.
- Model selection is earned, not asserted. In panel mode, every model that reaches the output carries a panel score; penalized models are shown, not silently dropped. The top 5 is a ranking, not a habit.
- The skill takes a position on the story. Phase 3.5 commits to one narrative path. Alternatives named and dismissed.
- Findings carry source tags. Every artifact carries
[known],[inferred], or[uncertain]. The tag carries through to the final output. - The skill infers, then confirms — it does not interrogate. The front-end is a single inferred-read confirm gate, not a question series: the skill reads the user's dump, infers audience / deliverable / frame, and presents one structured gate to accept or correct. Every prompt that remains — that confirm gate, the verdict gate, path acceptance, solo sparring rounds — uses AskUserQuestion with 3–4 candidate framings; "Other" captures the unfit case. The skill presents a choice; it never chats at the user.
How This Skill Works
You provide a thesis, strategy, discussion, or problem as a single dump. Phase 1 infers the run's shape from that dump — audience, deliverable type, frame — and routes into one of two modes at the confirm gate:
| Mode | Fires when | Who spars | |---|---|---| | Panel | deliverable_type is deck / memo / brief / analysis | Five category agents internally; you adjudicate the verdict | | Solo | deliverable_type is no-deliverable (personal thinking) | You, interactively, through the escalation ladder |
| Phase | Panel mode | Solo mode | |---|---|---| | 0 — Dump | Free-text situation + context — the only free-text input | same | | 1 — Inferred Read & Confirm | Infer audience / deliverable / decision / horizon / affected / frame from the dump; one AskUserQuestion gate to confirm or correct → routing | same | | 2 — Stress-test | Five-agent panel spars + scores all ~60 models | Single-context diagnostic (A1–A7) | | 2.5 — Meta Ranking | Composite scores, penalties, top 5 + bench, A1–A7 synthesis | — (not used) | | 3 — Adjudication | Verdict gate: accept / challenge / swap / re-run | Interactive sparring ladder to bedrock | | 3.5 — Path Selection | One recommended narrative path; top 5 = load-bearing models | same, models picked inline | | 4 — Output | Diagnostic + Model Selection Table + narrative, written inline; usage log appended | Diagnostic + narrative, inline |
References loaded on demand:
references/model-library.md— index of the ~60-model library; full entries live inreferences/models/(one file per category, one panel agent per file)references/panel-protocol.md— panel mode machinery: agent prompt template, scoring rubric, meta ranking algorithm, verdict gate, usage-log format (load only in panel mode)references/behavior-catalog.md— 43 phase-routed behaviors (Reflex 23 + Standard 20)state/usage-log.md— append-only record of past top-5 picks; feeds the recency penalty
Workflow
Phase 0: Dump (fires first — the only free-text input)
The user pastes or describes the situation and context in one go: the thesis or problem, plus whatever audience, deliverable, stakes, or framing they care to include. Short is fine. Rough is expected. The skill asks no questions here — this is the one place the user speaks freely.
The old three-question intake and the time-horizon / who's-affected questions are gone. Everything they used to capture, the skill now infers from the dump in Phase 1 and confirms in a single gate.
Phase 1: Inferred Read & Confirm Gate (fires before any stress-test, both modes)
The skill reads the dump, infers the run's shape, and confirms it in one AskUserQuestion gate. This single gate replaces the old Phase 0 intake, Phase 1 context, and Phase 1.5 frame-lock gate. The skill infers, then confirms — it does not interrogate.
1a — Infer (internal; main loop / meta agent, no subagent). From the dump, draft:
- Audience — board / client / team / investor / ops / technical / personal-thinking
- Deliverable type — deck / memo / brief / analysis / no-deliverable (this routes panel vs solo)
- Decision to influence — the specific call / approval / action, one sentence
- Time horizon — days / weeks–months / quarters / years / decades
- Who's affected — you alone / team / customers / market / multiple stakeholders
- Frame Lock — produced by running the frame-check behaviors as internal lenses over the dump, not as questions to the user: Framing, First Principles, Map vs Territory, Circle of Competence, Probabilistic Thinking, Honesty (Reflex), plus High Agency / Hanlon's Razor / Bayesian Updating (Standard) when their signal is present. See
references/behavior-catalog.md. The skill answers each from the dump itself.
A clean diagnosis of a wrongly-framed problem is worse than a messy diagnosis of the right one — so the frame is still checked before any stress-test runs, and in panel mode before five agents spend on it. It is just inferred from the dump and surfaced for confirmation rather than extracted by interrogation. Any field the dump can't settle confidently is marked (?) so the gate draws the eye to it.
Frame Lock format (rendered inside the read):
STATED FRAME:
ALTERNATIVE FRAMES: 1. — surfaces:
2. — surfaces:
FRAME LOCK:
1b — Present the read:
HERE'S HOW I READ THIS
Audience:
Deliverable: → mode
Decision:
Time horizon:
Who's affected:
Frame lock:
alternatives considered: ;
1c — Confirm gate (one AskUserQuestion):
- ACCEPT & DEPLOY — the read is right; launch (panel agents, or solo diagnostic + ladder).
- FIX ROUTING — audience / deliverable / decision is wrong; user corrects via Other/notes; re-render the read, re-gate.
- REFRAME — pick a surfaced alternative frame or supply a new one via Other; re-lock, re-render. (This is where the old Phase 1.5 frame-acceptance choice now lives.)
- ADD CONTEXT — the dump is thin or missing a fact that changes the read; user adds it, the skill re-infers, re-gate.
No stress-test runs until this gate clears. The locked frame then goes verbatim into every panel agent's brief (panel mode) or threads through every artifact (solo mode); if it is later revised, the panel re-runs / artifacts re-walk. On a confirmed deliverable that routes to panel, the skill enters Panel Mode (below) and loads references/panel-protocol.md; a no-deliverable run enters Solo Mode.
Thin-dump fallback: if multiple load-bearing fields are genuinely unknowable from the dump, the skill may fire one upstream clarifying AskUserQuestion before drafting the read. This is the exception, not the default — a normal dump goes straight to the read. Never reopen the old six-question barrage.
PANEL MODE (deliverable-bound runs)
Load references/panel-protocol.md now — it carries the agent prompt template, the full scoring rubric, the meta ranking algorithm, and the verdict gate spec. The sections below are the operating summary.
Phase 2: Panel Sparring (five-agent fan-out)
Launch five category agents in ONE message (parallel Agent tool calls, subagent_type: general-purpose). Each agent owns one category file under references/models/:
| Agent | Category | File | |---|---|---| | Thinking Foundations | A | models/a-thinking-foundations.md | | Systems & Dynamics | B | models/b-systems-dynamics.md | | Execution & Agency | C | models/c-execution-agency.md | | Strategy & Markets | D | models/d-strategy-markets.md | | Personal Practice | E | models/e-personal-practice.md |
Each agent receives the full thesis (verbatim, never summarized), the Frame Lock, audience, deliverable type, and decision — then spars the thesis against every model in its file and returns structured JSON: per-model grip / originality / story scores (0–5, anchored), the specific catch with source tag, the sharpest attack line, plus category-level strongest attack, bedrock candidate, tensions, prerequisites, and second-order effects. The prompt template in panel-protocol.md is mandatory — it carries the anti-inflation anchors ("most of your models should score 0–2 on most theses").
The main loop is the meta agent: in panel mode it never scores models itself. It briefs, composites, ranks, and defends the pick.
Phase 2.5: Meta Ranking & Diagnostic Synthesis
When all five return:
- Composite each scored model:
8×grip + 6×originality + 6×story(max 100). - Cliché penalty −10 on the boilerplate six (80/20, First Principles, Inversion, Second-Order, Compounding, Trade-offs) — waived at grip 5.
- Recency penalty −15 for models in the top-5 of any of the last 3 runs in
state/usage-log.md— waived at grip 5. This is what keeps consecutive decks from leaning on the same lens. - Concreteness gate: catches that don't cite thesis specifics cap at composite 50.
- Rank. Top 5 = load-bearing. Ranks 6–10 = bench. Pure score order — no category quota; if the honest top 5 is four Systems models, that's the answer.
- Pick THE bedrock from the five bedrock candidates: the assumption that, if false, kills the thesis most completely.
Render the Model Selection Table (format in panel-protocol.md) — including the penalties section whenever a penalty changed the top 5 — and synthesize the standard A1–A7 Diagnostic Report from the panel returns (mapping table in panel-protocol.md), source tags carried verbatim.
Phase 3: Verdict Gate (replaces interactive sparring in panel mode)
Present the bedrock, the five strongest attacks (one per category), the category verdicts, and the Model Selection Table. Then ONE AskUserQuestion gate:
- ACCEPT VERDICT — proceed to Phase 3.5 with the top 5 as load-bearing models
- CHALLENGE A FINDING — user supplies a defense; re-run only that category's agent with the defense appended; merge revised scores, re-rank, re-present
- SWAP A MODEL — promote from bench, demote from top 5; logged as user override
- ADD CONTEXT & RE-RUN — new facts change the test; full five-agent re-run
The panel replaced the interactive ladder, so this gate is where the user adjudicates instead of being sparred. One clean gate, then move — don't pad it with extra rounds.
SOLO MODE (no-deliverable runs)
Personal thinking doesn't need a committee. Solo mode runs the original single-context flow: the skill itself diagnoses and spars, loading category files from references/models/ as the exchange calls for them (route by theme via references/model-library.md).
Phase 2: Diagnostic Report — 7 Artifacts with Source Tags
The skill produces seven labeled artifacts. Each item within each artifact carries a source tag.
Source tag discipline:
| Tag | Meaning | |---|---| | [known] | Demonstrable from data the user provided, or facts they confirmed | | [inferred] | Drawn from the data by reasoning; the user did not state it explicitly | | [uncertain] | Open question, missing data, or contested point — needs validation |
No artifact item passes through unmarked. Default to [inferred] and note the gap if a tag is missing.
[A1] Models You're Implicitly Using (Well)
- [Model name] [tag]: How it shows up in your thinking
[A2] Models You're Ignoring (Blindspots)
- [Model name] [tag]: What you're missing
- Specific gap [tag]: concrete example from your thesis
[A3] Paradoxes & Tensions
- [Tension 1] [tag]: two conflicting assumptions you haven't resolved
- [Tension 2] [tag]: ...
[A4] Prerequisites You Haven't Listed
- [Prerequisite A] [tag]: Why this matters for your plan
- [Prerequisite B] [tag]: ...
[A5] Second-Order Effects & Tail Risks
- [Effect 1] [tag]: If X happens, then Y. You've planned for Y?
- [Effect 2] [tag]: ...
[A6] Grey Areas (Unresolved Questions)
1. [Question — typically [uncertain]]: needs an answer before you move
2. ...
[A7] Pragmatic Next Steps
1. [Action — typically [inferred]]: concrete action to harden the thesis
2. [Action — typically [inferred]]: ...
Behaviors that fire in solo Phase 2: Inversion, Find the Zero, Prerequisites, Trade-offs, Devil's Advocate, Bottleneck, 5 Whys (Reflex Tier) plus Standard Tier triggered by signal. See references/behavior-catalog.md.
Model count discipline: solo Phase 2 may touch 10–15 models across the artifacts. Phase 3.5 will pick the 3–5 that actually carry the story. Don't over-cite at this stage.
Phase 3: Sparring (Mandatory — Not Optional)
Sparring begins automatically after the report. There is no "want to spar?" prompt. The diagnostic report is the setup for the attack.
The escalation ladder. Every round attacks a strictly deeper layer than the round before. Never re-attack a cleared layer; always descend.
- L1 — the claim. Attack the thesis as stated.
- L2 — the defense. When the user defends, attack the defense itself.
- L3 — the reframe. When the user retreats to a reframe, attack the reframe.
- L4 — the assumption under the reframe. Attack what the reframe quietly assumes.
- L5+ — keep descending. Each round goes one layer deeper.
If you cannot find a deeper layer, you have reached bedrock. That's the exit signal.
Each round: ask the user to defend, play the opposite side, propose a thought experiment, or ask the clarifying question that exposes the next gap. One move per round.
Every sparring round uses AskUserQuestion. No prose attacks. The skill formulates each round as a question with 3–4 candidate framings the user might naturally take. "Other" captures the unfit case.
Candidate-answer patterns:
| Round type | Candidate options to supply | |---|---| | WHY attack | 3–4 common rationales the user might give | | Defense check | 3–4 typical defense moves (e.g. "concede / reframe / cite evidence / commit a test") | | Bedrock checkpoint | "Accept as bedrock / Push one more layer / Log as accepted risk / Defer pending evidence" | | Inversion attack | 3–4 failure modes the user might be doing right now | | Trade-off forcing | 3–4 concrete sides of the trade-off |
The skill writes the attack prose as the AskUserQuestion question field; the candidate framings become the options. Pure prose attacks ("why?", "defend that", "what would change your mind?") without structured options are forbidden.
Minimum three rounds. Sparring cannot end before round 3, even if the thesis looks sound. Early agreement is the failure mode this skill exists to kill.
Earned exit — bedrock only. Sparring terminates only when bedrock is reached: the single load-bearing assumption the user cannot defend further. At bedrock:
- Name it. State the load-bearing assumption in one sentence.
- Force a choice. The user must either (a) commit a concrete test, or (b) log it explicitly as a named accepted risk ("Accepted risk: X").
- Restate it back. Put the assumption and the chosen path on the record.
Explicit override. The user can end sparring early by saying "stop". Log it: "Sparring ended early by user at round N — b
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: umarmsharif
- Source: umarmsharif/mental-models
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.