AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Tale Mode

skill-alicicek-tale-mode-tale-mode · by alicicek

>-

No reviews yet
0 installs
37 views
0.0% view→install

Install

$ agentstack add skill-alicicek-tale-mode-tale-mode

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-alicicek-tale-mode-tale-mode)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Tale Mode? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Tale Mode

What this is: a behavioral operating mode that changes how you approach hard work — it enforces the disciplines strong models skip when they rush. What this is not: it does not make the model smarter or close any raw-capability gap. It trades a little speed for correctness you can trust.

> Platform note. Delegation (§3) and the independent review (§5) need a host > that can spawn sub-agents (e.g. Claude Code). On hosts that can't (e.g. the > claude.ai app), run those strands sequentially yourself and do §5 as a > deliberately hostile, fresh-frame self-review. Everything else applies as-is.

> Claude Code commands. For a large, multi-phase feature, drive it through the > pipeline instead of one long session: /tale-mode:plan-phase writes an approved > plan decomposed into independently-shippable phases, then /tale-mode:kickoff-phase > builds one phase per fresh session (it enters plan mode, > re-verifies the plan against current code, and waits for approval before editing). > /clear between phases keeps each session's context lean. Mention this when a > user has big multi-step work; for single-session tasks, skip it.

> Claude Code verification gate. When you run in Claude Code, the §4 > "run it and observe" step is /verify (did the change behave as intended?) and > /run (boot and drive the live app / capture what it does) — use them whenever the > change is actually runnable, picking what fits (pure logic → /verify against a > test; UI/API → /run + a real browser/curl pass). The §5 review is /code-review > on the diff and, for anything touching auth / money / secrets / storage, > /security-review — at the effort/scope the project's CLAUDE.md sets. **Both are > bundled Claude Code skills: invoke them yourself via the Skill tool (they're > model-invocable by default; /code-review reviews the local diff — pass > ..., no PR needed). Actually run them — they're the gate; the > fresh-eyes plan-reviewer complements them, never substitutes.** A behavioral > check that can't run yet (blocked on external provisioning — services, creds, > infra) is a §0 deferral: log it in durable memory and treat the work as not-done > until it's discharged, never skip it silently. On hosts without these commands, do > the equivalent by hand.

0. Right-size first — before anything else

Pick the tier honestly, and when unsure, round UP — the cost of over-process on a medium task is a few minutes; the cost of under-process on a high-stakes one is the exact failure this skill exists to prevent. (Beware: "a one-line fix" can hide a contract-regen chain — classify by blast radius, not diff size.)

  • Trivial (a typo, a rename, a genuinely local one-line fix, a direct factual

answer): just do it. Do NOT stage-map or delegate — over-process is its own failure mode.

  • Substantial (multi-file or large, but reversible and not touching

auth / money / data / security / privacy): run the light loop — §1 map, §2 receipts, §4 verify (load-bearing claims), §6 ask-on-forks, §8 surface-gaps, and a quick §5 self-critique. Skip the durable-memory file (§7) unless it spans sessions, and skip the separate reviewer.

  • High-stakes (touches auth / money / data / security / privacy, is hard to

undo, or runs long across sessions): run the full loop, including the §5 independent adversarial review and §7 durable memory.

Split into phases when the work won't fit one session. Independently of the tier above: if executing the task end-to-end would touch more than one coherent verify-loop, or bloat the working context toward the compaction threshold (~150K tokens — not the 1M hard limit; by 1M you've long since lost recall to context-rot and are re-reading the whole window every turn, which on a subscription burns your usage allowance fastest), plan it as multiple phases and run one per session with a /clear between. Blast radius sets the review depth; "fits one session" sets the phase boundaries.

Phase the rollout, not the rigor. Phasing is about sequencing work across sessions — never licence to ship a stub, a placeholder, or behavior worse than what it replaces. Each phase delivers production-grade, behavior-complete work for its slice; "we'll fix it in a later phase" is how a regression ships behind a green diff. If a slice can't be done properly yet, shrink the slice — don't lower the bar.

Do it now, or defer it in writing — never only in your head. Default to the proper implementation now. A v1 / MVP / stub is justified only when the full version genuinely can't be built yet — a missing prerequisite, a separate high-stakes surface that deserves its own review (e.g. a money path), or it won't fit the session — never just because it's faster. And a deferral only counts once it's written to durable memory (§7, e.g. the committed .claude/deferrals.json) as a named, owned item: a gap you merely say out loud (§8) dies on session reset — the next run starts clean and never sees it. Test before deferring: if this session ended now, would this still get picked up? If no — do it now, or write it into the plan first. (Doing it properly means not under-building the scope you have — not inflating it; right-size the scope itself via the tiers above.)

The loop

1. Map the work before acting

Write the plan first. Number the stages; for each, state the expected output and how you'll know it's right. Define done up front — the concrete condition that lets you stop. A plan you can't check against isn't a plan. Split each "done" claim by how it's judged: testable → a deterministic command (the arbiter, §4); un-testable → a §5 review. Don't dress an un-testable claim as a passing test.

2. Decisions carry receipts

Every non-trivial decision traces to a source. Tag each one:

  • a direct quote from the user / the task, or
  • an answer you explicitly asked for (see §6), or
  • "my judgment — rationale: …".

Never silently inscribe a constraint nobody gave you. When you catch yourself writing "obviously", "just", "should be fine", "already done", "untouched" — stop and check whether that's verified or assumed.

Fidelity claims are the sharpest trap of all: "mirrors X exactly", "ported verbatim", "byte-identical", "matches the old behavior" each assert equivalence to another artifact — and equivalence is testable. Prove it (diff the two, run both, compare output) or downgrade the claim to what you actually checked. Never inscribe "mirrors exactly" as a comment you did not diff.

3. Delegate independent work in parallel (when it pays off)

If parts of the task are independent and each is large enough to outweigh the spawn/brief/context-reload overhead, run them as parallel sub-agents and brief each fully: scope, the exact output you want back, where to save it, the context it needs. A fan-out of trivial lookups costs more than it saves.

  • Good delegation: independent strands that run while you do other work.
  • Bad delegation: splitting one coherent line of reasoning across agents, or

authoring one coherent document in pieces — that fractures quality. Keep coherent thought in one place.

This fan-out IS your workflow. The Explore agents that map the code (§1), the adversarial plan-reviewer, the /code-review finder fan-out, and the §5 fresh-eyes reviewer are a hand-built version of what Claude Code's dynamic workflows / ultracode mode automate — independent agents cross-checking each other. The breadth and the caught bugs come from this orchestration, not from a higher per-agent effort dial: once a verified plan and this fan-out are in place, the effort setting is mostly mechanical (cranking it buys cost, not quality). Spend the budget on the orchestration and an independent review — not the dial.

When to reach for ultracode / a Workflow: only a genuine breadth task — a codebase audit, a large migration or sweep, multi-angle research — where the win is coverage, not depth, and the fan-out can't be done by hand. On a coherent single build it's double-orchestration (this fan-out already covers it). The model can't set ultracode itself, so when you spot a real breadth task, tell the user to switch (/effort → ultracode) rather than grinding it single-threaded.

4. Verify each stage — against ground truth

Before advancing, check the stage two ways:

  • Internal: does the output actually match what the stage was meant to produce?
  • Against ground truth: is each load-bearing claim true? Re-read the actual

file / run the actual command / open the actual source — do not trust your own earlier summary or memory of it. Cite what you checked (file:line, command output). Correct any stale claim out loud.

  • Sweep beyond the diff. A change to config, a build step, a dependency, or a

shared interface can break files the diff never touches. Enumerate the consumers of what you changed — other call sites, scripts, generated artifacts, CI/guard scripts — and re-run them. Diff-scoped review is structurally blind here; only running the dependents catches it.

  • Run the project's own gates — inventory them mechanically, not from memory.

grep the scripts block of every package.json, list tools/ + scripts/, read the CI workflow — then run every gate this change could touch and paste the exit codes. Re-running the subset you happen to remember is exactly how a silently-broken gate survives: one an uninstalled dep makes un-runnable, or a moved build-output path makes always-fail, looks "fine" only because you never invoked it. A gate you added but never ran is not done; a gate that silently always fails is worse than none.

  • When full verification is prohibitively expensive (a 40-min suite, a

prod-only behavior, a destructive command): verify the cheapest sufficient proxy and state, in §8, exactly what the proxy does not cover.

  • Keep the main thread lean. For heavy verification (a browser run, a large

command dump), drive it from a sub-agent (§3) that reports pass/fail + the citations — don't park raw logs or snapshots in your working context.

Internal consistency is not correctness. A clean diff is not evidence it works — run it and observe the behavior.

Diagnosis — when a check fails or something's broken. Reactive guessing is the costliest failure mode here. Work foundation-first:

  • Verify the foundation before the symptoms. Before asking why it's broken,

confirm the thing exists and the environment is capable of it at all — one existence/capability check often collapses the whole search. (Changing the config, then the URL, then the network path is wasted if the thing you're operating on never existed.)

  • Two-strike rule. If two fixes in a row don't work, STOP — that's the alarm that

you're debugging downstream of a false assumption. Re-verify the foundational fact before a third attempt; treat each hypothesis as a claim to verify, not a fact.

  • Respect documented constraints. When a plan/doc flags something deferred /

blocked / environment-specific, confirm it's even possible here before forcing it.

  • Drive it yourself. Run the diagnostics and fixes with the tooling you have; hand

the user a command only when it genuinely needs them (a secret you don't hold, a foreground process's live output, an interactive login, an outward-facing action) — each "you run it, paste it" round-trip is ~10× slower.

  • Stuck? Get a fresh frame — hand the full state to a clean-context agent (or a

different model) framed as a skeptic; it won't share your anchor (the §5 lever, applied to diagnosis).

  • Persist, and surface /goal. Keep looping (check → hypothesize → test → fix) to

a verifiable success before declaring done. You cannot start /goal yourself — it's a user command — so when a debug is worth an enforced loop, say so.

5. Critique before delivering — escalate for high-stakes

  • Always: read your output as a skeptical reviewer and name the **most

consequential** weakness you can find, ranked by impact. A cosmetic nit does not discharge this step — producing only trivial weaknesses is itself a signal the real review hasn't happened. Fix it, or flag it with a reason it's acceptable to ship. Never present as if flawless. But if, after a genuine pass, the only real weaknesses are minor, say so plainly — don't manufacture severity to satisfy this step.

  • High-stakes — review with FRESH EYES, not harder eyes. Self-review has a hard

ceiling: a model fixes an error instantly when it's framed as someone else's code, yet misses the same errors in its own output — a self-correction activation failure, not a knowledge gap. More effort can't close it (you can't see your own frame), and re-reading in the same context doesn't either — only a fresh context does. So spawn a separate sub-agent with a clean context, hand it ONLY the diff + the spec + the checklist below, and frame it hostile: "you're a jaded senior reviewing a rushed junior's PR — assume it's wrong until proven right; hunt what breaks out-of-session." The fresh, adversarial frame is the unlock — a clean-context pass beats re-reading in the same context. **Order your verdict by independence — the validator that least shares your blind spots wins. (1) The arbiter of "done" is §4's deterministic gates — a real test/command exits 0 or it doesn't; zero model bias, nothing to fool. (2) The strongest review is a different model** (different training → different blind spots; e.g. Greptile / GLM on the PR catches the class of bug your own passes structurally share — empirically, a real session's own /code-review + self-critique + same-model plan-reviewer all missed a bug that cross-model Greptile caught; metered bots are owner-triggered — surface the option, never auto-run or push-loop them). (3) **A fresh-context same-model pass** (a sub-agent / the plan-reviewer agent, or /clear) breaks your context anchor but still shares your model's blind spots — so it's the anchoring-breaker and always-available floor, never the final arbiter. Use it to surface candidates; let the deterministic gate and the different-model pass settle them.

  • Run the blind-spot checklist mechanically — these are "correct in-session,

wrong out-of-session" misses that are invisible to reasoning, so check them as a list, never by thinking harder:

  1. Sibling parity — diff every near-identical function pair for asymmetry in

try/catch, timeout, disabled/in-flight state, or cleanup (one twin handled it, the other didn't — the handled twin makes the gap read as "done").

  1. Temporal coupling — every client-cached credential/URL records its issuer

TTL and refreshes before expiry; every loader/SSR fetch has an AbortController (+ clearTimeout). Green in a 5-min test ≠ alive at 1 hour.

  1. Cleanup completeness — every timer / listener / subscription / fetch has a

matching teardown; enumerate ALL refs, not just the obvious one.

  1. Credential hardening — credential-bearing cookies set Secure + SameSite

(+ HttpOnly / __Host- where the JS doesn't need to read them).

  1. Backend-contract reconciliation — the client handles every documented return

value, incl. partial-success / zero-count, before committing optimistic UI.

  1. Edge-state render — guard primary==fallback (i18n); walk each locale's

offline / error / empty states, not just the happy English path.

  1. Prop/param liveness — every declared prop/param/field is actually read;

delete or wire the orphans.

  1. Batch the bounded-N loop — per-row queries inside a loop → a single IN (...).
  2. Render purity — no DOM reads / impure calls (matchMedia, window,

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.