AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Planner Evaluator Harness

skill-archive228-lab-skills-planner-evaluator-harness · by Archive228

Multi-agent harness pattern (planner / generator / evaluator) for long-running application builds, distilled from Anthropic's engineering post on harness design. Use when a build will run for hours or sits beyond what the current model ships reliably solo: covers sprint contracts, skeptical-evaluator calibration, hard-threshold QA grading via browser automation, context resets vs compaction, and…

— No reviews yet
0 installs
21 views
0.0% view→install

Install

$ agentstack add skill-archive228-lab-skills-planner-evaluator-harness

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-archive228-lab-skills-planner-evaluator-harness)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Planner Evaluator Harness? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Planner / Generator / Evaluator Harness for Long Application Builds

This skill gives you Anthropic's multi-agent harness pattern for long-running application development: a planner that expands a short prompt into an ambitious spec, a generator that builds in bounded sprints, and a separately-prompted skeptical evaluator that grades the running application through browser automation. The split exists because agents asked to judge their own work confidently praise it even when quality is obviously mediocre; "far more tractable than making a generator critical of its own work" is how the source describes tuning a standalone evaluator to be skeptical. On a retro-game-maker build with Opus 4.5, the harness ($200, 6 hours) shipped a game whose core play loop actually worked — the entity moved and responded to input, with rough edges remaining — where a solo run ($9, 20 minutes) shipped an interface that looked functional but whose core gameplay was silently broken.

When to use

Use the harness when:

  • A build will run for hours, not minutes, from a short (1–4 sentence) prompt to a full application.
  • The task sits beyond what the current model completes reliably in a solo run — per the source, the evaluator earns its cost only when the task is past the model's solo-reliability edge.
  • Quality targets are subjective (e.g., frontend design), where binary self-verification is impossible and a calibrated evaluator makes subjectivity gradable.
  • A previous solo attempt produced work that looked done on the surface but failed under real interaction.

Do NOT use the full harness when:

  • The task is within the current model's baseline solo capability — there the evaluator is unnecessary overhead and a solo run is ~20x cheaper (the $9-vs-$200 comparison above).
  • You are on a stronger model than the harness was designed for. Every harness component encodes an assumption about what the model cannot do alone; those assumptions go stale. Re-test before paying the overhead (see rules 11–13).

Rules

  1. Separate the generator from the evaluator. Never let the agent that built the work grade it as the gate. Self-evaluation is blind; a standalone evaluator prompted to be skeptical is the tractable lever.
  1. Planner: expand, don't over-specify. The planner turns the 1–4 sentence user prompt into a full product specification. Make it deliberately ambitious about scope, and keep it at product context and high-level technical design. If the planner fixes granular technical details upfront and gets one wrong, the error cascades into every downstream implementation step. Instruct it to weave AI features into the product spec. Reference scale: "Create a 2D retro game maker" expanded to a 16-feature spec across 10 sprints.
  1. Generator: one feature per sprint. The generator implements the spec sprint by sprint, one feature at a time — sprint decomposition keeps the model coherent over long builds. Reference stack from the source: React, Vite, FastAPI, SQLite (later PostgreSQL), with git for version control. The generator self-checks its work at sprint end before handing off to QA; self-check informs, the evaluator decides.
  1. Negotiate a sprint contract before writing code. Before each sprint, generator and evaluator agree on what "done" looks like for that chunk. Have the generator put forward the sprint's scope plus a plan for verifying success, let the evaluator push back, and keep iterating until both sides accept it. This turns high-level user stories into testable commitments.
  1. Evaluator: test the running app like a user. Give the evaluator Playwright MCP (or equivalent browser automation) and have it click through the live application — UI features, API endpoints, and database states — not just read code or score static screenshots. For design work, have it navigate the live page and screenshot before scoring.
  1. Grade against hard per-criterion thresholds. The evaluator grades each sprint against discovered bugs plus criteria covering product depth, functionality, visual design, and code quality. Each criterion gets a hard threshold; if any single one falls below it, the sprint fails and the generator receives detailed feedback. Demand file-and-line specificity in failure reports — the source's evaluator produced findings like a fillRectangle not triggered on mouseUp, a delete-key condition bug at LevelEditor.tsx:892, and a PUT /frames/reorder route defined after /{frame_id} routes so FastAPI matched "reorder" as an integer and returned 422.
  1. Calibrate the evaluator; out of the box it is a poor QA agent. Untuned, the evaluator finds legitimate issues, then talks itself into deciding they aren't a big deal and approves anyway; it also tests superficially without probing edge cases, letting subtle bugs slip through. Fix by (a) few-shot scored examples with detailed breakdowns, which anchor its judgment to yours and cut score drift across iterations, and (b) several rounds of QA-prompt iteration driven by concrete cases where its judgment diverged from yours.
  1. Communicate between agents via files. Specs, sprint contracts, and feedback all travel as files: each agent reads its counterpart's file and answers either in place or in a fresh file for the other to pick up.
  1. Prefer context resets over compaction for long runs. Compaction keeps continuity without wiping the slate, so context anxiety (the model prematurely wrapping up as its window fills — observed in Sonnet 4.5) can persist. A reset wipes the slate, but only works if the handoff artifact carries enough state for the next agent to continue cleanly. Opus 4.6 largely removed this behavior natively; the source dropped context resets entirely on it.
  1. For subjective design work, weight the weak criteria and watch your wording. The four design criteria: design quality (coherent whole, not a collection of parts), originality (deliberate custom decisions vs. template defaults — purple-gradients-over-white-cards fails), craft (typographic hierarchy, spacing consistency, color harmony, contrast), functionality (usable without guessing). Weight design quality and originality more heavily, since craft and functionality are strong by default. Beware that criterion wording steers the generator: phrasing like "the best designs are museum quality" pushed outputs toward one visual convergence. Run 5–15 evaluation iterations per generation (full runs up to four hours), and after each evaluation let the generator choose: refine the current direction, or pivot entirely if scores plateau.
  1. Match harness complexity to model capability. On Opus 4.6 the source removed the sprint construct (the model decomposed work natively and ran coherently for over two hours), moved the evaluator from per-sprint grading to a single pass at the end of the run, and kept the planner (prevents under-scoping) and the evaluator (for tasks at the capability edge). Treat the evaluator as a conditional component, not a fixed yes/no.
  1. Simplify methodically: remove one component at a time. The source's first V2 attempt cut the harness back radically and could not replicate performance. What worked was removing a single component at a time and reviewing the impact on the final result before proceeding.
  1. Re-examine the harness whenever a new model lands. Strip pieces no longer load-bearing; add pieces that unlock new capability. The space of useful harness combinations moves rather than shrinks.
  1. Budget realistically. V2 reference run — "Build a fully featured DAW in the browser using the Web Audio API": planner 4.7 min / $0.46; Build R1 2 h 7 min / $71.08; QA R1 8.8 min / $3.24; Build R2 1 h 2 min / $36.89; QA R2 6.8 min / $3.09; Build R3 10.9 min / $5.88; QA R3 9.6 min / $4.06; total 3 h 50 min / $124.70. QA rounds catch display-only stubs (clips that can't be dragged, record buttons that toggle but capture no mic input, numeric sliders where graphical EQ curves belong).

Checklist

Setup:

  • [ ] Decide harness vs. solo: is the task beyond what the current model ships reliably alone? If not, run solo.
  • [ ] Planner prompt: ambitious scope, product-level spec, no granular technical details, AI features woven in.
  • [ ] Generator prompt: one feature per sprint, git commits, self-check before QA handoff.
  • [ ] Evaluator prompt: skeptical persona, Playwright MCP access, per-criterion hard thresholds, file:line specificity required in failures.
  • [ ] Calibrate evaluator with few-shot scored examples; plan for several rounds of QA-prompt iteration on divergence cases.
  • [ ] File-based communication channels defined for spec, contracts, and feedback.

Per sprint (or per build/QA round on newer models):

  • [ ] Sprint contract negotiated and agreed before any code: what "done" means + how it will be verified.
  • [ ] Generator implements exactly the contracted feature.
  • [ ] Evaluator drives the running app — UI, API endpoints, DB state — not just code review.
  • [ ] Any criterion below its hard threshold → sprint fails, detailed feedback returned, generator fixes.
  • [ ] On score plateau (design work): generator explicitly chooses refine vs. pivot.

On model upgrade:

  • [ ] Remove one harness component at a time; measure impact on the final result.
  • [ ] Test whether sprints are still needed (can the model decompose and stay coherent for 2+ hours natively?).
  • [ ] Consider moving the evaluator from per-sprint to a single end-of-run pass.
  • [ ] Keep the planner unless proven redundant — it prevents under-scoping.

Anti-patterns

  • Letting the generator gate its own work. Agents confidently praise mediocre output; use the separate evaluator as the gate.
  • Trusting an untuned evaluator. It will find real bugs and then approve anyway, and it tests superficially. Calibrate before relying on it.
  • Trusting surface functionality. The solo retro-game-maker run looked functional; entities rendered on screen but nothing responded to input, and the wiring between entity definitions and the game runtime was broken with no surface indication of where. Verification means driving the app, not seeing it render.
  • Over-specifying in the plan. Granular technical detail in the spec cascades wrong decisions into every sprint.
  • Compaction as a substitute for a clean slate. It preserves continuity but not the reset that eliminates context anxiety on models that exhibit it.
  • Loaded aesthetic language in criteria. Phrases like "museum quality" homogenize outputs; criterion wording is a steering input, treat it as such.
  • Running the full harness on tasks inside baseline capability. Pure overhead — the split pays only past the model's solo-reliability edge.
  • Freezing the harness across model generations. Components encode assumptions about model weakness; stale assumptions cost money and time without adding quality.

Source

  • Harness design for long-running application development — https://www.anthropic.com/engineering/harness-design-long-running-apps — 2026-03-24

Distilled from the official document(s) above on 2026-08-12. If this skill and the source disagree, trust the source.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.