# Plan It

> >-

- **Type:** Skill
- **Install:** `agentstack add skill-devotts-plan-it-plan-it`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [DevOtts](https://agentstack.voostack.com/s/devotts)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [DevOtts](https://github.com/DevOtts)
- **Source:** https://github.com/DevOtts/plan-it/tree/main/plugins/plan-it/skills/plan-it

## Install

```sh
agentstack add skill-devotts-plan-it-plan-it
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# plan-it

**Take a fuzzy demand → ship a buildable delivery package.** This is the planning
conductor: the disciplined front-half of the lifecycle that ends exactly where
`/fable-it` begins. It does *discovery* (research the ground truth), *spec*
(author the design docs), and *agile split* (PRDs, epics, tests, the shared
contract) — then hands off.

```
  /plan-it  ─────────────►  docs/ + delivery/  ─────────────►  /fable-it
  (discovery → spec → plan)   (the buildable package)            (builds it)
```

Usable by a human directly, and by any orchestrating **conductor** agent that receives a new
demand and must turn it into a delivery package before dispatching workers.

---

## The five non-negotiable rules (enforce these — don't just suggest them)

These are the load-bearing rules reverse-engineered from every successful run.
If you violate one, the build downstream drifts or silently fails.

1. **Freeze a shared CONTRACT before any parallel planning.** The CONTRACT is the
   law: canonical entities, schema, API/interface, enums, repo/branch map, and the
   definition of "shipped." Squads write *to* it; any cross-cutting discovery folds
   *back* into it as a dated amendment (v1.0 → v1.1 …). No frozen contract → no
   parallel squads.

2. **Batch every human-only decision into ONE gate.** Do not pre-decide anything
   irreversible (repo topology, hosting, product name, architectural mode, build-vs-buy).
   Surface them together, each with a *recommendation attached*, and let the human
   answer numbered. This gate is where the human injects **vision**, not just picks
   options — leave room for them to add a concept you didn't propose. Lock each
   answer with **owner + date**.

3. **Verify every agent's output on disk — "idle ≠ delivered."** A team going idle
   does NOT mean it wrote files. After any fan-out, check the actual paths exist and
   are non-empty before proceeding. If a team held its output as a message, direct
   it to `Write` to the exact absolute path. Never trust a "done."

4. **Ground the plan against the LIVE system, not the repo — before you freeze.**
   For any plan touching a running system, the repo is a *hypothesis*; the deployed
   reality is the truth, and they drift. Before freezing the CONTRACT, verify against
   the actual system and write the *observations* in (not repo-derived guesses).
   Battle-tested: in one program EVERY mid-flight correction traced to a repo-inferred
   assumption reality contradicted — wrong canonical identifiers (manifests had drifted
   from the deployed catalog), a config value that was present-but-pointing-at-a-dead-host,
   "to-be-built" components that were already deployed, and a found credential that was
   the *wrong* one. See Phase 3's live-grounding gate for the concrete checks.

5. **Run the machine, not the prose.** The pipeline's control flow lives in
   `machine.json` (the explicit statechart), not in this document — this prose
   *explains* the machine. On invocation, read or initialize `.plan-it/state.json`
   (in the target project) and resume from its `state`; write it on **every**
   transition. At every guarded transition, run the guard's mapped subcommand
   (`node scripts/gate-check.mjs  …`) and **never advance on a non-zero
   exit** — fix, re-run, then transition. If Node is unavailable, perform the same
   checks manually and record them in the state file (degrade, never break). Full
   protocol: `references/machine.md`.

---

## The deterministic core (v2) — why a machine

Control flow written as prose ("do step 1, never skip the gate") is what the
determinism literature calls **prose control flow**: it relies on the model's
discipline across a long, summarization-prone context, and sometimes the model
won't follow it. v2 inverts that at the right altitude — *non-determinism at the
edges, determinism at the core*:

- **`machine.json`** — XState v5-compatible statechart of the pipeline: 15 states,
  the three human gates (`meta.gate` + `meta.human`), guarded transitions, and an
  `AMENDMENT` self-loop on `parallelPlanning`. Paste into stately.ai/viz to see it.
- **`.plan-it/state.json`** — the persisted run: current state, gate approvals
  (owner + date), contract version, verified-artifact registry, history. This is
  what makes a run survive a crash or a fresh session.
- **`scripts/gate-check.mjs`** — the guards as exit codes: `verify` (Rule 3,
  idle ≠ delivered), `freeze` (Rule 1, no contract → no squads), `handoff` (the
  mechanizable half of playbooks §F), `state` (Rule 2, gates recorded),
  `adversary` (Rule 6 / D4 — failure-mode depth: a modelled machine must cover-or-
  waive the five cascade classes; N/A for linear workflows). The `LINT_CLEAN` and
  `ADVERSARY_CLEAN` transitions (verify → adversaryGate → handoff) gate on the
  last two.

The fuzzy phases — discovery, synthesis, spec authoring, judgment — stay
LLM-at-the-node (tagged `llmAtTheNode` in the machine). Do not formalize them;
modeling is not ceremony only when it replaces confusion. Details, state-file
schema, and the resume protocol: `references/machine.md`.

**Hard enforcement (v2.1, plugin installs on Claude Code only):** a `PreToolUse`
hook (`scripts/hooks/planit-guard.mjs`) *denies* Write/Edit calls on PRD/epic
deliverables while the run's contract is unfrozen — Rule 1 stops being an
instruction and becomes something the harness refuses. Fail-open: it never
touches non-plan-it work. Skill-only installs rely on Rule 5 discipline instead.

---

## The Test Contract — the quality differentiator (make this non-negotiable)

The single thing that most raises delivered code quality: **every PRD/epic ends by
generating its own test contract — up to ~20 concrete use-cases/scenarios that
stress the implementation — and the feature is NOT "done" until 100% of them pass.**
The build agent cannot just deliver the feature; it must satisfy the contract, and
`/iterate` until green.

This is a named, proven discipline: **Specification by Example** (Gojko Adzic) +
**ATDD/BDD** for code (concrete examples become executable acceptance tests and
living documentation), and **Eval-Driven Development** for skills/LLM features
(register goldens with expected outputs, iterate until they pass). Authoring the
cases *at planning time* is the whole point — they become a binding contract, not
an afterthought.

Rules of the Test Contract:
1. **Authored at planning time** — the LAST step of writing each epic/PRD is its
   test contract. Expected outputs are **registered now** ("designed first; expected
   outputs registered"), not discovered mid-build.
2. **Up to ~20 high-quality cases** per feature — enough to stress real behavior +
   edges; *not* thousands of shallow ones (quality > quantity — auto-bulk = "slop",
   per the eval literature). Draw cases from real/likely failure modes.
   **Per-shape count rule:** large Shape-1 multi-squad programs hold a **≥10
   cases-per-epic floor** (each epic is a big feature); small shapes (single
   skill/feature, S/M) author **~20 cases total across the package** (a handful per
   epic). Never both at once — pick by shape so a reviewer doesn't flag a correct
   small package as under-tested.
3. **Binding** — DoD = **100% of the contract passes**; until then, `/iterate`. No
   partial ship; no VERIFIED-on-a-mock (a `[REAL]` case whose target is unreachable
   → IMPLEMENTED-NOT-VERIFIED, never a fake green).
4. **Pick the test types by implementation** (one or more of unit / e2e / use-cases
   / stress):

| Implementation | Test types | How |
|----|----|----|
| CRUD / REST API | use-cases (happy+edge) + e2e | run every scenario **via API** *and* **via UI with Chrome CDP** (`chrome-cdp-control`); unit-test the logic |
| Skill / prompt / LLM function | use-cases w/ **expected output** | run it, **compare real vs expected** — exact match for closed outputs, rubric / LLM-as-judge for open ones (`make-eval`, promptfoo, DeepEval G-Eval) |
| Agent / stateful / multi-step | **stress** scenarios + use-cases | six axes: async, fan-out, escalation, human-gate, recursion, cycle-guard (Setup/Expected/Pass) |
| Pure logic / library | unit + **property-based** | enumerated cases + invariants |
| Data pipeline / migration | golden-value + e2e | hand-computed expected values; idempotency/rollback |
| Anything with load/abuse surface | **stress / adversarial** | concurrency, rate, malformed input, red-team |

5. **Execution path:** `/full-qa` runs the contract, `/iterate` loops it to 100%,
   `chrome-cdp-control` drives UI scenarios. **The contract is the bridge from
   plan-it → fable-it: `/fable-it`'s Definition of Done = this contract.**

Grammars and the contract header format: `references/formats.md` (the Test Contract
block + §4–5).

---

## Autonomy posture (guided with autonomous bursts)

Run research and authoring autonomously at high effort, but **stop at three gates**:

| Gate | When | What you ask |
|------|------|--------------|
| **G1 — Scope** | after intake (Phase 2) | confirm the sizing (feature vs program) + the numbered DoD before burning effort |
| **G2 — Decisions** | after specs drafted (Phase 7) | the batched "decisions only you can make," each with a recommendation |
| **G3 — Delivery** | before the agile split (Phase 8) | "specs look aligned — proceed to PRDs/epics?" |

Everything between gates runs unattended. Recommend `/effort xhigh` at the start
(you cannot set it yourself — tell the user to run `/effort xhigh` if they haven't).

---

## Phase 0 — Intake

**Machine first (Rule 5):** if `.plan-it/state.json` exists in the target project,
run `node scripts/gate-check.mjs state .plan-it/state.json` and **resume from the
printed state** — do not restart phases already in `history`. If it doesn't exist,
create it now in state `intake` (schema in `references/machine.md`) and keep it
updated on every transition for the rest of the run.

Accept the demand in whatever form it arrives: a brain-dump, a pasted
transcription, a list of wants, or a one-liner. **Expect pointers, not content** —
session names (`/read-chat ""`), repo paths, doc folders. Your job is to go
fetch the ground truth, not to be handed it.

Capture up front:
- **The raw vision** in the user's own words (you'll quote it back in `02 §1`).
- **Pointers** to prior sessions / repos / docs to research.
- **Use-case** (auto-detect — this drives the packaging shape at Gate G1):
  - new single app, greenfield · feature on a large existing repo ·
    from-scratch multi-subsystem program · multi-app platform (many PRDs) ·
    refactor / migration / debt · research spike (no build yet) · PM/board
    automation · document/audit an already-built system.
- **Research method**: default to parallel Claude teams at xhigh.

If the demand is genuinely one fuzzy paragraph with no pointers and an existing
repo, that's fine — pre-grounding (Phase 3) will find the targets.

---

## Phase 1 — DoD lock

Restructure the fuzzy prose into a **numbered, individually-verifiable Definition
of Done** + a short list of stated assumptions. This is your contract with the
user for the planning job itself. Example shape:

```
DoD for this planning run:
  1. Ground-truth findings doc (every claim → path:line or table)
  2. Vision + architecture doc that solves each finding/contradiction
  3. Data/interface contract
  4. … (auto-sized — see Phase 2)
  N. Handoff: contract frozen, PRDs+epics with ≥10 tests each, kickoff prompt
Assumptions: 
```

---

## Phase 2 — Scope & shape governor  ⏸ GATE G1

Pick **size** (how much) *and* **shape** (what form) before spending effort.
Confirm both with the user.

**Size** scales the artifact count:

| Signal | Size |
|--------|------|
| Single feature, 1 subsystem | **S** |
| Multi-feature / new subsystem, 1–2 repos | **M** |
| From-scratch program / many subsystems | **L** |

**Shape** is chosen by use-case (full definitions + the use-case→shape table in
`references/templates.md` PART D):

1. **Multi-doc + `delivery/`** (baseline) — from-scratch program, parallel squads.
2. **Single-file PRD-as-everything** — greenfield single app; CONTRACT inlined as G-rules.
3. **Research → locked-architecture → master+phase PRDs** — feature on a large existing repo.
4. **`implementation//` with numbered PRD-NN** — multi-app platform, many PRDs.
5. **Refactor/debt workstream catalog** — brownfield in-place.
   (+ research-spike, executable-board, and reverse-doc modes — see PART D.)

Same phases regardless — only the artifact count and form change. **Don't give a
feature a 4-squad org; don't give a brownfield refactor a greenfield vision doc.**
Present the chosen size + shape + the numbered DoD and get a yes before proceeding.
On yes, record `G1 = {approved, owner, date}` plus the chosen size/shape in
`.plan-it/state.json` — the `G1_APPROVED` transition is guarded by `gateRecorded`.

---

## Phase 3 — Pre-ground

Before any fan-out, locate exact targets so agents get precise paths, not vague
instructions:
- Resolve pointer sessions with a session-reading skill (`/read-chat` or similar)
  if one is installed; otherwise ask for the relevant transcript or summary.
- `ls`/`grep` the named repos; find the `docs/`, schema files, entry points.
- Read the wiki hot-cache/index if the repo is vault-wired (per its `CLAUDE.md`).

Output: a list of `{subsystem → exact paths to read}` that seeds the research teams.

### Live-grounding gate (rule 4) — when the plan touches a RUNNING system

The repo tells you what *should* be true; the live system tells you what *is*. Do this
before Phase 8 freeze, and write the findings into the CONTRACT as observed facts:

- **Config reachability, not presence.** For every env/config lever the plan depends on
  (base URLs, API hosts, feature flags), *hit it* — don't just confirm it's set. A value
  can be set to a dead host and pass a "does it exist" check while silently breaking prod.
- **Live registry / canonical identifiers.** Dump the actual slugs / IDs / enum values the
  running system uses (integration slugs, catalog rows, schema). Repo manifests drift from
  the deployed catalog — bind the CONTRACT's canon to the LIVE value, not the manifest.
- **Deployed vs installed vs in-use.** Inventory what's already running, what's registered,
  and what's actually used by a tenant/user. "Build & deploy X" is often "redeploy X," and
  a component that's deployed-but-uninstalled cannot be live-verified per-item — which
  changes both verification depth and priority. Tag each target.
- **Credential validity, not existence.** A found key/token ≠ a correct one. If a plan
  relies on a credential, verify it actually works (a live call), and confirm ownership —
  a leftover token from another account will pass structural checks and fail at runtime.
- **Derive dependency sets from ACTUAL usage, cross-checked live.** Build the integration/
  dependency list from `code-grep ∪ every component's declared requirements`, intersected
  with the live registry — then list the *gaps* as explicit work. A list derived from a
  subset or from repo manifests will miss things the running system actually needs.
- **Separate "our code change" from "external connection/config/data seeding."** The code
  change is usually the easy, ownable part; the real blocker is often an external OAuth
  connection, a seeded secret VALUE, or an ops action — and it belongs on the critical path
  with an owner, not buried as a footnote.

Two build-time corollaries worth encoding in the CONTRACT/epics so the builder inherits them:
map a caller/consumer surface **before** planning any auth-guard or interface change (a
guard that breaks N callers is worse than no guard); and any repo *import* must be a
full-tree secret-scan + runtime-only subset, never a history mirror.

---

## Phase 4 — Discovery (autonomous burst)

Pick a **discovery mode** by use-case (full playbook in `references/playbooks.md` §A):
- **Solo read** — small feature, one subsystem.
- **Parallel research streams** — research-hea

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [DevOtts](https://github.com/DevOtts)
- **Source:** [DevOtts/plan-it](https://github.com/DevOtts/plan-it)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-devotts-plan-it-plan-it
- Seller: https://agentstack.voostack.com/s/devotts
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
