AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Incremental Implementation

skill-celestialdust-achilles-skills-incremental-implementation · by celestialdust

Builds one assigned slice as thin, individually-tested vertical increments — skeleton-first (stub→mock→wire→fill), simplicity-first, test-first. Use the moment you start writing code for a planned slice, and ESPECIALLY when you're tempted to write more than ~100 lines before running a test, land a whole feature in one pass, "clean up" code outside the slice, or weaken a test to get green. Runs in…

No reviews yet
0 installs
11 views
0.0% view→install

Install

$ agentstack add skill-celestialdust-achilles-skills-incremental-implementation

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-celestialdust-achilles-skills-incremental-implementation)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Incremental Implementation? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Incremental Implementation

Purpose

Stage: Implement (agent — per slice, inside a worktree). THE implementer workhorse. It applies test-driven-development as its rigid core loop and source-driven-development as a referenced discipline; the worktree isolation mechanism is owned by the orchestrator — implement runs in the worktree it is handed, it does not make its own.

Build in thin vertical slices — implement one piece, test it, verify it, then expand. Avoid implementing an entire feature in one pass. Each increment should leave the system in a working, testable state. This is the execution discipline that makes large features manageable.

When to use / when to skip

  • Implementing any multi-file change
  • Building a new feature from a task breakdown
  • Refactoring existing code
  • Any time you're tempted to write more than ~100 lines before testing

When NOT to use: Single-file, single-function changes where the scope is already minimal.

Inputs

Implement runs per slice, inside the worktree the orchestrator hands it (it does not create its own isolation; that is worktree, owned upstream). It consumes the Plan-stage contract cold and refuses to run if a load-bearing input is missing.

| Input | Source (skill) | Stable sections it reads | Refuse-to-run if absent | |---|---|---|---| | plan.md + slices | plan-breakdown | the assigned slice's row keyed by Slice id — its Story-ref · Files (owned) · Regression surface · Checkpoint · Blocked-by columns — plus the line-level steps and exact tests in the plan body | no plan.md, or no concrete steps for the assigned slice | | assigned slice id | STATE.md (orchestrator) | the PRD-namespaced slice row in state impl, gate agent | no slice assigned, or it is not in impl/agent | | acceptance.md | acceptance-criteria (Spec, signed) | the behavioral Given/When/Then scenario ids (e.g. PWR-A2) this slice realizes — the frozen oracle test-driven-development turns RED | the slice references an acceptance id that does not exist | | clean worktree | worktree (orchestrator) | a provisioned, preflight-green baseline branch for this slice | not running inside the handed worktree |

Disciplines it applies (consulted, not stages it consumes): test-driven-development (RED-GREEN-REFACTOR; test-first order is hook-enforced), source-driven-development (ground any framework/library decision in fetched official docs, as needed), debugging-and-error-recovery (five-step triage when a slice's tests break). It does not re-derive the plan, re-open the spec, or re-slice — the plan handed to it is already vertical.

The Increment Cycle

┌──────────────────────────────────────┐
│                                      │
│   Implement ──→ Test ──→ Verify ──┐  │
│       ▲                           │  │
│       └───── Commit ◄─────────────┘  │
│              │                       │
│              ▼                       │
│          Next slice                  │
│                                      │
└──────────────────────────────────────┘

For each slice:

  1. Implement the smallest complete piece of functionality
  2. Test — run the test suite (or write a test if none exists)
  3. Verify — confirm the slice works as expected (tests pass, build succeeds, manual check)
  4. Commit -- save your progress with a descriptive message (see the git-workflow skill for atomic commit guidance)
  5. Move to the next slice — carry forward, don't restart

Slicing Strategies

Vertical Slices (Preferred)

Build one complete path through the stack:

Slice 1: Create a task (DB + API + basic UI)
    → Tests pass, user can create a task via the UI

Slice 2: List tasks (query + API + UI)
    → Tests pass, user can see their tasks

Slice 3: Edit a task (update + API + UI)
    → Tests pass, user can modify tasks

Slice 4: Delete a task (delete + API + UI + confirmation)
    → Tests pass, full CRUD complete

Each slice delivers working end-to-end functionality.

Skeleton-First (stub → mock → wire → fill)

Within a vertical slice, build the skeleton end-to-end first, then fill it in — absorbed from cr-structure's build-order. Each step is independently observable:

  • Stub — every layer the slice touches returns a hardcoded value; the end-to-end path already runs.
  • Mock — swap stubs for mocks at the real boundaries; the shape of the data flows through.
  • Wire — replace mocks with the real calls, one boundary at a time.
  • Fill — handle the edge cases and error paths.

This is the antidote to horizontal building (all DB, then all API, then all UI), which yields code that does not work end-to-end until the last step and gives you nothing to debug from in between. The plan you were handed is already sliced vertically; skeleton-first is how you build each slice without silently re-horizontalizing it. The Increment Cycle's "Verify" step is the slice's checkpoint — a specific observable fact ("submitting the form shows the inline error"), never "it compiles".

Contract-First Slicing

When backend and frontend need to develop in parallel:

Slice 0: Define the API contract (types, interfaces, OpenAPI spec)
Slice 1a: Implement backend against the contract + API tests
Slice 1b: Implement frontend against mock data matching the contract
Slice 2: Integrate and test end-to-end

Risk-First Slicing

Tackle the riskiest or most uncertain piece first:

Slice 1: Prove the WebSocket connection works (highest risk)
Slice 2: Build real-time task updates on the proven connection
Slice 3: Add offline support and reconnection

If Slice 1 fails, you discover it before investing in Slices 2 and 3.

Implementation Rules

Rule 0: Simplicity First

Before writing any code, ask: "What is the simplest thing that could work?"

After writing code, review it against these checks:

  • Can this be done in fewer lines?
  • Are these abstractions earning their complexity?
  • Would a staff engineer look at this and say "why didn't you just..."?
  • Am I building for hypothetical future requirements, or the current task?
SIMPLICITY CHECK:
✗ Generic EventBus with middleware pipeline for one notification
✓ Simple function call

✗ Abstract factory pattern for two similar components
✓ Two straightforward components with shared utilities

✗ Config-driven form builder for three forms
✓ Three form components

Three similar lines of code is better than a premature abstraction. Implement the naive, obviously-correct version first. Optimize only after correctness is proven with tests.

Rule 0.5: Scope Discipline

Touch only what the task requires.

Do NOT:

  • "Clean up" code adjacent to your change
  • Refactor imports in files you're not modifying
  • Remove comments you don't fully understand
  • Add features not in the spec because they "seem useful"
  • Modernize syntax in files you're only reading

If you notice something worth improving outside your task scope, note it — don't fix it:

NOTICED BUT NOT TOUCHING:
- src/utils/format.ts has an unused import (unrelated to this task)
- The auth middleware could use better error messages (separate task)
→ Want me to create tasks for these?

Rule 1: One Thing at a Time

Each increment changes one logical thing. Don't mix concerns:

Bad: One commit that adds a new component, refactors an existing one, and updates the build config.

Good: Three separate commits — one for each change.

Rule 2: Keep It Compilable

After each increment, the project must build and existing tests must pass. Don't leave the codebase in a broken state between slices.

Rule 3: Feature Flags for Incomplete Features

If a feature isn't ready for users but you need to merge increments:

// Feature flag for work-in-progress
const ENABLE_TASK_SHARING = process.env.FEATURE_TASK_SHARING === 'true';

if (ENABLE_TASK_SHARING) {
  // New sharing UI
}

This lets you merge small increments to the main branch without exposing incomplete work.

Rule 4: Safe Defaults

New code should default to safe, conservative behavior:

// Safe: disabled by default, opt-in
export function createTask(data: TaskInput, options?: { notify?: boolean }) {
  const shouldNotify = options?.notify ?? false;
  // ...
}

Rule 5: Rollback-Friendly

Each increment should be independently revertable:

  • Additive changes (new files, new functions) are easy to revert
  • Modifications to existing code should be minimal and focused
  • Database migrations should have corresponding rollback migrations
  • Avoid deleting something in one commit and replacing it in the same commit — separate them

Working with Agents

When directing an agent to implement incrementally:

"Let's implement Task 3 from the plan.

Start with just the database schema change and the API endpoint.
Don't touch the UI yet — we'll do that in the next increment.

After implementing, run `npm test` and `npm run build` to verify
nothing is broken."

Be explicit about what's in scope and what's NOT in scope for each increment.

Increment Checklist

After each increment, verify:

  • [ ] The change does one thing and does it completely
  • [ ] All existing tests still pass (npm test)
  • [ ] The build succeeds (npm run build)
  • [ ] Type checking passes (npx tsc --noEmit)
  • [ ] Linting passes (npm run lint)
  • [ ] The new functionality works as expected
  • [ ] The change is committed with a descriptive message

Note: Run each verification command after a change that could affect it. After a successful run, don't repeat the same command unless the code has changed since — re-running on unchanged code adds no information.

Rationalizations

| Rationalization | Reality | |---|---| | "I'll test it all at the end" | Bugs compound. A bug in Slice 1 makes Slices 2-5 wrong. Test each slice. | | "It's faster to do it all at once" | It feels faster until something breaks and you can't find which of 500 changed lines caused it. | | "These changes are too small to commit separately" | Small commits are free. Large commits hide bugs and make rollbacks painful. | | "I'll add the feature flag later" | If the feature isn't complete, it shouldn't be user-visible. Add the flag now. | | "This refactor is small enough to include" | Refactors mixed with features make both harder to review and debug. Separate them. | | "Let me run the build command again just to be sure" | After a successful run, repeating the same command adds nothing unless the code has changed since. Run it again after subsequent edits, not as reassurance. |

Red flags

  • More than 100 lines of code written without running tests
  • Multiple unrelated changes in a single increment
  • "Let me just quickly add this too" scope expansion
  • Skipping the test/verify step to move faster
  • Build or tests broken between increments
  • Large uncommitted changes accumulating
  • Building abstractions before the third use case demands it
  • Touching files outside the task scope "while I'm here"
  • Creating new utility files for one-time operations
  • Running the same build/test command twice in a row without any intervening code change

Verification (ending criteria)

After completing all increments for a task:

  • [ ] Each increment was individually tested and committed
  • [ ] The full test suite passes
  • [ ] The build is clean
  • [ ] The feature works end-to-end as specified
  • [ ] No uncommitted changes remain

See Also

Per-increment verification is the local check. Before declaring a task done, apply the project-wide Definition of Done as the final gate, the standing bar every increment clears regardless of the task. See ../../references/definition-of-done.md.

Outputs & handoff contract

  • Emits: a diff on the slice's worktree branch — the implemented slice as a sequence of atomic,

individually-tested commits (code + the tests that prove it). This is the artifact quality-verification (Verify) consumes cold, alongside the running app and acceptance.md.

  • Stable guarantees the consumer depends on:
  • The diff stays inside the slice's declared Regression surface (from plan.md). Narrowing or

widening that surface to pass is a gate-erosion HALT, not a fix.

  • Tests land test-first (the test-driven-development order hook enforces it) and assert observable behavior, never

mock calls (testing-strategy AP1).

  • The worktree is compilable and green at every increment checkpoint — never left broken between slices.
  • The slice diff stays within the ≤400 LOC cluster cap; if it cannot, the

slice was mis-sliced — stop and surface, do not stretch the cap.

  • Frozen-under-retry (silent-false-green defense — non-negotiable): during this slice's bounded retry

rounds, acceptance.md, the RED tests, and the declared Regression surface are immutable. A retry diff that weakens an assertion, deletes a test, or narrows the surface = HALT (gate-erosion + reward-hack tripwire: the failure signature must not move only because a test/acceptance was edited while impl is materially unchanged). The way to green is to fix the impl (via debugging-and-error-recovery), never to move the goalposts. No --no-verify, no hook edits, no SKIP_HOOKS (security.md / CLAUDE.md).

  • STATE.md update: on a green slice checkpoint, flip the slice impl → verify (gate stays agent)

and hand off to quality-verification. If the slice cannot pass after the bounded rounds (3 implement→verify→review cycles), flip it impl → halted and flip its gate agent → you — the failure-escalation path is the only place a human gate survives the autonomous run.

  • Consumer: quality-verification (Verify). Change the shape of what you emit → update quality-verification in the same commit.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.