Install
$ agentstack add skill-devotts-plan-it-plan-it ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
plan-it
Take a fuzzy demand → ship a buildable delivery package. This is the planning conductor: the disciplined front-half of the lifecycle that ends exactly where /fable-it begins. It does discovery (research the ground truth), spec (author the design docs), and agile split (PRDs, epics, tests, the shared contract) — then hands off.
/plan-it ─────────────► docs/ + delivery/ ─────────────► /fable-it
(discovery → spec → plan) (the buildable package) (builds it)
Usable by a human directly, and by any orchestrating conductor agent that receives a new demand and must turn it into a delivery package before dispatching workers.
The five non-negotiable rules (enforce these — don't just suggest them)
These are the load-bearing rules reverse-engineered from every successful run. If you violate one, the build downstream drifts or silently fails.
- Freeze a shared CONTRACT before any parallel planning. The CONTRACT is the
law: canonical entities, schema, API/interface, enums, repo/branch map, and the definition of "shipped." Squads write to it; any cross-cutting discovery folds back into it as a dated amendment (v1.0 → v1.1 …). No frozen contract → no parallel squads.
- Batch every human-only decision into ONE gate. Do not pre-decide anything
irreversible (repo topology, hosting, product name, architectural mode, build-vs-buy). Surface them together, each with a recommendation attached, and let the human answer numbered. This gate is where the human injects vision, not just picks options — leave room for them to add a concept you didn't propose. Lock each answer with owner + date.
- Verify every agent's output on disk — "idle ≠ delivered." A team going idle
does NOT mean it wrote files. After any fan-out, check the actual paths exist and are non-empty before proceeding. If a team held its output as a message, direct it to Write to the exact absolute path. Never trust a "done."
- Ground the plan against the LIVE system, not the repo — before you freeze.
For any plan touching a running system, the repo is a hypothesis; the deployed reality is the truth, and they drift. Before freezing the CONTRACT, verify against the actual system and write the observations in (not repo-derived guesses). Battle-tested: in one program EVERY mid-flight correction traced to a repo-inferred assumption reality contradicted — wrong canonical identifiers (manifests had drifted from the deployed catalog), a config value that was present-but-pointing-at-a-dead-host, "to-be-built" components that were already deployed, and a found credential that was the wrong one. See Phase 3's live-grounding gate for the concrete checks.
- Run the machine, not the prose. The pipeline's control flow lives in
machine.json (the explicit statechart), not in this document — this prose explains the machine. On invocation, read or initialize .plan-it/state.json (in the target project) and resume from its state; write it on every transition. At every guarded transition, run the guard's mapped subcommand (node scripts/gate-check.mjs …) and never advance on a non-zero exit — fix, re-run, then transition. If Node is unavailable, perform the same checks manually and record them in the state file (degrade, never break). Full protocol: references/machine.md.
The deterministic core (v2) — why a machine
Control flow written as prose ("do step 1, never skip the gate") is what the determinism literature calls prose control flow: it relies on the model's discipline across a long, summarization-prone context, and sometimes the model won't follow it. v2 inverts that at the right altitude — non-determinism at the edges, determinism at the core:
machine.json— XState v5-compatible statechart of the pipeline: 15 states,
the three human gates (meta.gate + meta.human), guarded transitions, and an AMENDMENT self-loop on parallelPlanning. Paste into stately.ai/viz to see it.
.plan-it/state.json— the persisted run: current state, gate approvals
(owner + date), contract version, verified-artifact registry, history. This is what makes a run survive a crash or a fresh session.
scripts/gate-check.mjs— the guards as exit codes:verify(Rule 3,
idle ≠ delivered), freeze (Rule 1, no contract → no squads), handoff (the mechanizable half of playbooks §F), state (Rule 2, gates recorded), adversary (Rule 6 / D4 — failure-mode depth: a modelled machine must cover-or- waive the five cascade classes; N/A for linear workflows). The LINT_CLEAN and ADVERSARY_CLEAN transitions (verify → adversaryGate → handoff) gate on the last two.
The fuzzy phases — discovery, synthesis, spec authoring, judgment — stay LLM-at-the-node (tagged llmAtTheNode in the machine). Do not formalize them; modeling is not ceremony only when it replaces confusion. Details, state-file schema, and the resume protocol: references/machine.md.
Hard enforcement (v2.1, plugin installs on Claude Code only): a PreToolUse hook (scripts/hooks/planit-guard.mjs) denies Write/Edit calls on PRD/epic deliverables while the run's contract is unfrozen — Rule 1 stops being an instruction and becomes something the harness refuses. Fail-open: it never touches non-plan-it work. Skill-only installs rely on Rule 5 discipline instead.
The Test Contract — the quality differentiator (make this non-negotiable)
The single thing that most raises delivered code quality: every PRD/epic ends by generating its own test contract — up to ~20 concrete use-cases/scenarios that stress the implementation — and the feature is NOT "done" until 100% of them pass. The build agent cannot just deliver the feature; it must satisfy the contract, and /iterate until green.
This is a named, proven discipline: Specification by Example (Gojko Adzic) + ATDD/BDD for code (concrete examples become executable acceptance tests and living documentation), and Eval-Driven Development for skills/LLM features (register goldens with expected outputs, iterate until they pass). Authoring the cases at planning time is the whole point — they become a binding contract, not an afterthought.
Rules of the Test Contract:
- Authored at planning time — the LAST step of writing each epic/PRD is its
test contract. Expected outputs are registered now ("designed first; expected outputs registered"), not discovered mid-build.
- Up to ~20 high-quality cases per feature — enough to stress real behavior +
edges; not thousands of shallow ones (quality > quantity — auto-bulk = "slop", per the eval literature). Draw cases from real/likely failure modes. Per-shape count rule: large Shape-1 multi-squad programs hold a ≥10 cases-per-epic floor (each epic is a big feature); small shapes (single skill/feature, S/M) author ~20 cases total across the package (a handful per epic). Never both at once — pick by shape so a reviewer doesn't flag a correct small package as under-tested.
- Binding — DoD = 100% of the contract passes; until then,
/iterate. No
partial ship; no VERIFIED-on-a-mock (a [REAL] case whose target is unreachable → IMPLEMENTED-NOT-VERIFIED, never a fake green).
- Pick the test types by implementation (one or more of unit / e2e / use-cases
/ stress):
| Implementation | Test types | How | |----|----|----| | CRUD / REST API | use-cases (happy+edge) + e2e | run every scenario via API and via UI with Chrome CDP (chrome-cdp-control); unit-test the logic | | Skill / prompt / LLM function | use-cases w/ expected output | run it, compare real vs expected — exact match for closed outputs, rubric / LLM-as-judge for open ones (make-eval, promptfoo, DeepEval G-Eval) | | Agent / stateful / multi-step | stress scenarios + use-cases | six axes: async, fan-out, escalation, human-gate, recursion, cycle-guard (Setup/Expected/Pass) | | Pure logic / library | unit + property-based | enumerated cases + invariants | | Data pipeline / migration | golden-value + e2e | hand-computed expected values; idempotency/rollback | | Anything with load/abuse surface | stress / adversarial | concurrency, rate, malformed input, red-team |
- Execution path:
/full-qaruns the contract,/iterateloops it to 100%,
chrome-cdp-control drives UI scenarios. The contract is the bridge from plan-it → fable-it: /fable-it's Definition of Done = this contract.
Grammars and the contract header format: references/formats.md (the Test Contract block + §4–5).
Autonomy posture (guided with autonomous bursts)
Run research and authoring autonomously at high effort, but stop at three gates:
| Gate | When | What you ask | |------|------|--------------| | G1 — Scope | after intake (Phase 2) | confirm the sizing (feature vs program) + the numbered DoD before burning effort | | G2 — Decisions | after specs drafted (Phase 7) | the batched "decisions only you can make," each with a recommendation | | G3 — Delivery | before the agile split (Phase 8) | "specs look aligned — proceed to PRDs/epics?" |
Everything between gates runs unattended. Recommend /effort xhigh at the start (you cannot set it yourself — tell the user to run /effort xhigh if they haven't).
Phase 0 — Intake
Machine first (Rule 5): if .plan-it/state.json exists in the target project, run node scripts/gate-check.mjs state .plan-it/state.json and resume from the printed state — do not restart phases already in history. If it doesn't exist, create it now in state intake (schema in references/machine.md) and keep it updated on every transition for the rest of the run.
Accept the demand in whatever form it arrives: a brain-dump, a pasted transcription, a list of wants, or a one-liner. Expect pointers, not content — session names (/read-chat ""), repo paths, doc folders. Your job is to go fetch the ground truth, not to be handed it.
Capture up front:
- The raw vision in the user's own words (you'll quote it back in
02 §1). - Pointers to prior sessions / repos / docs to research.
- Use-case (auto-detect — this drives the packaging shape at Gate G1):
- new single app, greenfield · feature on a large existing repo ·
from-scratch multi-subsystem program · multi-app platform (many PRDs) · refactor / migration / debt · research spike (no build yet) · PM/board automation · document/audit an already-built system.
- Research method: default to parallel Claude teams at xhigh.
If the demand is genuinely one fuzzy paragraph with no pointers and an existing repo, that's fine — pre-grounding (Phase 3) will find the targets.
Phase 1 — DoD lock
Restructure the fuzzy prose into a numbered, individually-verifiable Definition of Done + a short list of stated assumptions. This is your contract with the user for the planning job itself. Example shape:
DoD for this planning run:
1. Ground-truth findings doc (every claim → path:line or table)
2. Vision + architecture doc that solves each finding/contradiction
3. Data/interface contract
4. … (auto-sized — see Phase 2)
N. Handoff: contract frozen, PRDs+epics with ≥10 tests each, kickoff prompt
Assumptions:
Phase 2 — Scope & shape governor ⏸ GATE G1
Pick size (how much) and shape (what form) before spending effort. Confirm both with the user.
Size scales the artifact count:
| Signal | Size | |--------|------| | Single feature, 1 subsystem | S | | Multi-feature / new subsystem, 1–2 repos | M | | From-scratch program / many subsystems | L |
Shape is chosen by use-case (full definitions + the use-case→shape table in references/templates.md PART D):
- Multi-doc +
delivery/(baseline) — from-scratch program, parallel squads. - Single-file PRD-as-everything — greenfield single app; CONTRACT inlined as G-rules.
- Research → locked-architecture → master+phase PRDs — feature on a large existing repo.
implementation//with numbered PRD-NN — multi-app platform, many PRDs.- Refactor/debt workstream catalog — brownfield in-place.
(+ research-spike, executable-board, and reverse-doc modes — see PART D.)
Same phases regardless — only the artifact count and form change. Don't give a feature a 4-squad org; don't give a brownfield refactor a greenfield vision doc. Present the chosen size + shape + the numbered DoD and get a yes before proceeding. On yes, record G1 = {approved, owner, date} plus the chosen size/shape in .plan-it/state.json — the G1_APPROVED transition is guarded by gateRecorded.
Phase 3 — Pre-ground
Before any fan-out, locate exact targets so agents get precise paths, not vague instructions:
- Resolve pointer sessions with a session-reading skill (
/read-chator similar)
if one is installed; otherwise ask for the relevant transcript or summary.
ls/grepthe named repos; find thedocs/, schema files, entry points.- Read the wiki hot-cache/index if the repo is vault-wired (per its
CLAUDE.md).
Output: a list of {subsystem → exact paths to read} that seeds the research teams.
Live-grounding gate (rule 4) — when the plan touches a RUNNING system
The repo tells you what should be true; the live system tells you what is. Do this before Phase 8 freeze, and write the findings into the CONTRACT as observed facts:
- Config reachability, not presence. For every env/config lever the plan depends on
(base URLs, API hosts, feature flags), hit it — don't just confirm it's set. A value can be set to a dead host and pass a "does it exist" check while silently breaking prod.
- Live registry / canonical identifiers. Dump the actual slugs / IDs / enum values the
running system uses (integration slugs, catalog rows, schema). Repo manifests drift from the deployed catalog — bind the CONTRACT's canon to the LIVE value, not the manifest.
- Deployed vs installed vs in-use. Inventory what's already running, what's registered,
and what's actually used by a tenant/user. "Build & deploy X" is often "redeploy X," and a component that's deployed-but-uninstalled cannot be live-verified per-item — which changes both verification depth and priority. Tag each target.
- Credential validity, not existence. A found key/token ≠ a correct one. If a plan
relies on a credential, verify it actually works (a live call), and confirm ownership — a leftover token from another account will pass structural checks and fail at runtime.
- Derive dependency sets from ACTUAL usage, cross-checked live. Build the integration/
dependency list from code-grep ∪ every component's declared requirements, intersected with the live registry — then list the gaps as explicit work. A list derived from a subset or from repo manifests will miss things the running system actually needs.
- Separate "our code change" from "external connection/config/data seeding." The code
change is usually the easy, ownable part; the real blocker is often an external OAuth connection, a seeded secret VALUE, or an ops action — and it belongs on the critical path with an owner, not buried as a footnote.
Two build-time corollaries worth encoding in the CONTRACT/epics so the builder inherits them: map a caller/consumer surface before planning any auth-guard or interface change (a guard that breaks N callers is worse than no guard); and any repo import must be a full-tree secret-scan + runtime-only subset, never a history mirror.
Phase 4 — Discovery (autonomous burst)
Pick a discovery mode by use-case (full playbook in references/playbooks.md §A):
- Solo read — small feature, one subsystem.
- Parallel research streams — research-hea
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: DevOtts
- Source: DevOtts/plan-it
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.