AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Plan It

skill-devotts-plan-it-plan-it · by DevOtts

>-

— No reviews yet
0 installs
31 views
0.0% view→install

Install

$ agentstack add skill-devotts-plan-it-plan-it

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-devotts-plan-it-plan-it)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Plan It? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

plan-it

Take a fuzzy demand → ship a buildable delivery package. This is the planning conductor: the disciplined front-half of the lifecycle that ends exactly where /fable-it begins. It does discovery (research the ground truth), spec (author the design docs), and agile split (PRDs, epics, tests, the shared contract) — then hands off.

  /plan-it  ─────────────►  docs/ + delivery/  ─────────────►  /fable-it
  (discovery → spec → plan)   (the buildable package)            (builds it)

Usable by a human directly, and by any orchestrating conductor agent that receives a new demand and must turn it into a delivery package before dispatching workers.


The five non-negotiable rules (enforce these — don't just suggest them)

These are the load-bearing rules reverse-engineered from every successful run. If you violate one, the build downstream drifts or silently fails.

  1. Freeze a shared CONTRACT before any parallel planning. The CONTRACT is the

law: canonical entities, schema, API/interface, enums, repo/branch map, and the definition of "shipped." Squads write to it; any cross-cutting discovery folds back into it as a dated amendment (v1.0 → v1.1 …). No frozen contract → no parallel squads.

  1. Batch every human-only decision into ONE gate. Do not pre-decide anything

irreversible (repo topology, hosting, product name, architectural mode, build-vs-buy). Surface them together, each with a recommendation attached, and let the human answer numbered. This gate is where the human injects vision, not just picks options — leave room for them to add a concept you didn't propose. Lock each answer with owner + date.

  1. Verify every agent's output on disk — "idle ≠ delivered." A team going idle

does NOT mean it wrote files. After any fan-out, check the actual paths exist and are non-empty before proceeding. If a team held its output as a message, direct it to Write to the exact absolute path. Never trust a "done."

  1. Ground the plan against the LIVE system, not the repo — before you freeze.

For any plan touching a running system, the repo is a hypothesis; the deployed reality is the truth, and they drift. Before freezing the CONTRACT, verify against the actual system and write the observations in (not repo-derived guesses). Battle-tested: in one program EVERY mid-flight correction traced to a repo-inferred assumption reality contradicted — wrong canonical identifiers (manifests had drifted from the deployed catalog), a config value that was present-but-pointing-at-a-dead-host, "to-be-built" components that were already deployed, and a found credential that was the wrong one. See Phase 3's live-grounding gate for the concrete checks.

  1. Run the machine, not the prose. The pipeline's control flow lives in

machine.json (the explicit statechart), not in this document — this prose explains the machine. On invocation, read or initialize .plan-it/state.json (in the target project) and resume from its state; write it on every transition. At every guarded transition, run the guard's mapped subcommand (node scripts/gate-check.mjs …) and never advance on a non-zero exit — fix, re-run, then transition. If Node is unavailable, perform the same checks manually and record them in the state file (degrade, never break). Full protocol: references/machine.md.


The deterministic core (v2) — why a machine

Control flow written as prose ("do step 1, never skip the gate") is what the determinism literature calls prose control flow: it relies on the model's discipline across a long, summarization-prone context, and sometimes the model won't follow it. v2 inverts that at the right altitude — non-determinism at the edges, determinism at the core:

  • machine.json — XState v5-compatible statechart of the pipeline: 15 states,

the three human gates (meta.gate + meta.human), guarded transitions, and an AMENDMENT self-loop on parallelPlanning. Paste into stately.ai/viz to see it.

  • .plan-it/state.json — the persisted run: current state, gate approvals

(owner + date), contract version, verified-artifact registry, history. This is what makes a run survive a crash or a fresh session.

  • scripts/gate-check.mjs — the guards as exit codes: verify (Rule 3,

idle ≠ delivered), freeze (Rule 1, no contract → no squads), handoff (the mechanizable half of playbooks §F), state (Rule 2, gates recorded), adversary (Rule 6 / D4 — failure-mode depth: a modelled machine must cover-or- waive the five cascade classes; N/A for linear workflows). The LINT_CLEAN and ADVERSARY_CLEAN transitions (verify → adversaryGate → handoff) gate on the last two.

The fuzzy phases — discovery, synthesis, spec authoring, judgment — stay LLM-at-the-node (tagged llmAtTheNode in the machine). Do not formalize them; modeling is not ceremony only when it replaces confusion. Details, state-file schema, and the resume protocol: references/machine.md.

Hard enforcement (v2.1, plugin installs on Claude Code only): a PreToolUse hook (scripts/hooks/planit-guard.mjs) denies Write/Edit calls on PRD/epic deliverables while the run's contract is unfrozen — Rule 1 stops being an instruction and becomes something the harness refuses. Fail-open: it never touches non-plan-it work. Skill-only installs rely on Rule 5 discipline instead.


The Test Contract — the quality differentiator (make this non-negotiable)

The single thing that most raises delivered code quality: every PRD/epic ends by generating its own test contract — up to ~20 concrete use-cases/scenarios that stress the implementation — and the feature is NOT "done" until 100% of them pass. The build agent cannot just deliver the feature; it must satisfy the contract, and /iterate until green.

This is a named, proven discipline: Specification by Example (Gojko Adzic) + ATDD/BDD for code (concrete examples become executable acceptance tests and living documentation), and Eval-Driven Development for skills/LLM features (register goldens with expected outputs, iterate until they pass). Authoring the cases at planning time is the whole point — they become a binding contract, not an afterthought.

Rules of the Test Contract:

  1. Authored at planning time — the LAST step of writing each epic/PRD is its

test contract. Expected outputs are registered now ("designed first; expected outputs registered"), not discovered mid-build.

  1. Up to ~20 high-quality cases per feature — enough to stress real behavior +

edges; not thousands of shallow ones (quality > quantity — auto-bulk = "slop", per the eval literature). Draw cases from real/likely failure modes. Per-shape count rule: large Shape-1 multi-squad programs hold a ≥10 cases-per-epic floor (each epic is a big feature); small shapes (single skill/feature, S/M) author ~20 cases total across the package (a handful per epic). Never both at once — pick by shape so a reviewer doesn't flag a correct small package as under-tested.

  1. Binding — DoD = 100% of the contract passes; until then, /iterate. No

partial ship; no VERIFIED-on-a-mock (a [REAL] case whose target is unreachable → IMPLEMENTED-NOT-VERIFIED, never a fake green).

  1. Pick the test types by implementation (one or more of unit / e2e / use-cases

/ stress):

| Implementation | Test types | How | |----|----|----| | CRUD / REST API | use-cases (happy+edge) + e2e | run every scenario via API and via UI with Chrome CDP (chrome-cdp-control); unit-test the logic | | Skill / prompt / LLM function | use-cases w/ expected output | run it, compare real vs expected — exact match for closed outputs, rubric / LLM-as-judge for open ones (make-eval, promptfoo, DeepEval G-Eval) | | Agent / stateful / multi-step | stress scenarios + use-cases | six axes: async, fan-out, escalation, human-gate, recursion, cycle-guard (Setup/Expected/Pass) | | Pure logic / library | unit + property-based | enumerated cases + invariants | | Data pipeline / migration | golden-value + e2e | hand-computed expected values; idempotency/rollback | | Anything with load/abuse surface | stress / adversarial | concurrency, rate, malformed input, red-team |

  1. Execution path: /full-qa runs the contract, /iterate loops it to 100%,

chrome-cdp-control drives UI scenarios. The contract is the bridge from plan-it → fable-it: /fable-it's Definition of Done = this contract.

Grammars and the contract header format: references/formats.md (the Test Contract block + §4–5).


Autonomy posture (guided with autonomous bursts)

Run research and authoring autonomously at high effort, but stop at three gates:

| Gate | When | What you ask | |------|------|--------------| | G1 — Scope | after intake (Phase 2) | confirm the sizing (feature vs program) + the numbered DoD before burning effort | | G2 — Decisions | after specs drafted (Phase 7) | the batched "decisions only you can make," each with a recommendation | | G3 — Delivery | before the agile split (Phase 8) | "specs look aligned — proceed to PRDs/epics?" |

Everything between gates runs unattended. Recommend /effort xhigh at the start (you cannot set it yourself — tell the user to run /effort xhigh if they haven't).


Phase 0 — Intake

Machine first (Rule 5): if .plan-it/state.json exists in the target project, run node scripts/gate-check.mjs state .plan-it/state.json and resume from the printed state — do not restart phases already in history. If it doesn't exist, create it now in state intake (schema in references/machine.md) and keep it updated on every transition for the rest of the run.

Accept the demand in whatever form it arrives: a brain-dump, a pasted transcription, a list of wants, or a one-liner. Expect pointers, not content — session names (/read-chat ""), repo paths, doc folders. Your job is to go fetch the ground truth, not to be handed it.

Capture up front:

  • The raw vision in the user's own words (you'll quote it back in 02 §1).
  • Pointers to prior sessions / repos / docs to research.
  • Use-case (auto-detect — this drives the packaging shape at Gate G1):
  • new single app, greenfield · feature on a large existing repo ·

from-scratch multi-subsystem program · multi-app platform (many PRDs) · refactor / migration / debt · research spike (no build yet) · PM/board automation · document/audit an already-built system.

  • Research method: default to parallel Claude teams at xhigh.

If the demand is genuinely one fuzzy paragraph with no pointers and an existing repo, that's fine — pre-grounding (Phase 3) will find the targets.


Phase 1 — DoD lock

Restructure the fuzzy prose into a numbered, individually-verifiable Definition of Done + a short list of stated assumptions. This is your contract with the user for the planning job itself. Example shape:

DoD for this planning run:
  1. Ground-truth findings doc (every claim → path:line or table)
  2. Vision + architecture doc that solves each finding/contradiction
  3. Data/interface contract
  4. … (auto-sized — see Phase 2)
  N. Handoff: contract frozen, PRDs+epics with ≥10 tests each, kickoff prompt
Assumptions: 

Phase 2 — Scope & shape governor ⏸ GATE G1

Pick size (how much) and shape (what form) before spending effort. Confirm both with the user.

Size scales the artifact count:

| Signal | Size | |--------|------| | Single feature, 1 subsystem | S | | Multi-feature / new subsystem, 1–2 repos | M | | From-scratch program / many subsystems | L |

Shape is chosen by use-case (full definitions + the use-case→shape table in references/templates.md PART D):

  1. Multi-doc + delivery/ (baseline) — from-scratch program, parallel squads.
  2. Single-file PRD-as-everything — greenfield single app; CONTRACT inlined as G-rules.
  3. Research → locked-architecture → master+phase PRDs — feature on a large existing repo.
  4. implementation// with numbered PRD-NN — multi-app platform, many PRDs.
  5. Refactor/debt workstream catalog — brownfield in-place.

(+ research-spike, executable-board, and reverse-doc modes — see PART D.)

Same phases regardless — only the artifact count and form change. Don't give a feature a 4-squad org; don't give a brownfield refactor a greenfield vision doc. Present the chosen size + shape + the numbered DoD and get a yes before proceeding. On yes, record G1 = {approved, owner, date} plus the chosen size/shape in .plan-it/state.json — the G1_APPROVED transition is guarded by gateRecorded.


Phase 3 — Pre-ground

Before any fan-out, locate exact targets so agents get precise paths, not vague instructions:

  • Resolve pointer sessions with a session-reading skill (/read-chat or similar)

if one is installed; otherwise ask for the relevant transcript or summary.

  • ls/grep the named repos; find the docs/, schema files, entry points.
  • Read the wiki hot-cache/index if the repo is vault-wired (per its CLAUDE.md).

Output: a list of {subsystem → exact paths to read} that seeds the research teams.

Live-grounding gate (rule 4) — when the plan touches a RUNNING system

The repo tells you what should be true; the live system tells you what is. Do this before Phase 8 freeze, and write the findings into the CONTRACT as observed facts:

  • Config reachability, not presence. For every env/config lever the plan depends on

(base URLs, API hosts, feature flags), hit it — don't just confirm it's set. A value can be set to a dead host and pass a "does it exist" check while silently breaking prod.

  • Live registry / canonical identifiers. Dump the actual slugs / IDs / enum values the

running system uses (integration slugs, catalog rows, schema). Repo manifests drift from the deployed catalog — bind the CONTRACT's canon to the LIVE value, not the manifest.

  • Deployed vs installed vs in-use. Inventory what's already running, what's registered,

and what's actually used by a tenant/user. "Build & deploy X" is often "redeploy X," and a component that's deployed-but-uninstalled cannot be live-verified per-item — which changes both verification depth and priority. Tag each target.

  • Credential validity, not existence. A found key/token ≠ a correct one. If a plan

relies on a credential, verify it actually works (a live call), and confirm ownership — a leftover token from another account will pass structural checks and fail at runtime.

  • Derive dependency sets from ACTUAL usage, cross-checked live. Build the integration/

dependency list from code-grep ∪ every component's declared requirements, intersected with the live registry — then list the gaps as explicit work. A list derived from a subset or from repo manifests will miss things the running system actually needs.

  • Separate "our code change" from "external connection/config/data seeding." The code

change is usually the easy, ownable part; the real blocker is often an external OAuth connection, a seeded secret VALUE, or an ops action — and it belongs on the critical path with an owner, not buried as a footnote.

Two build-time corollaries worth encoding in the CONTRACT/epics so the builder inherits them: map a caller/consumer surface before planning any auth-guard or interface change (a guard that breaks N callers is worse than no guard); and any repo import must be a full-tree secret-scan + runtime-only subset, never a history mirror.


Phase 4 — Discovery (autonomous burst)

Pick a discovery mode by use-case (full playbook in references/playbooks.md §A):

  • Solo read — small feature, one subsystem.
  • Parallel research streams — research-hea

…

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.