# Proof Driven Dev

> Use for any request to add, build, implement or fix behavior whose success is not obvious from the diff: features, bug fixes, refactors that must preserve behavior, migrations, performance or security changes. Turns the request into numbered observable requirements, including unstated negative cases and boundaries, implements, runs the checks that prove each one, and reports VERIFIED, REVIEW REQU…

- **Type:** Skill
- **Install:** `agentstack add skill-soumyarauth-skills-hub-proof-driven-dev`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [soumyaRauth](https://agentstack.voostack.com/s/soumyarauth)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [soumyaRauth](https://github.com/soumyaRauth)
- **Source:** https://github.com/soumyaRauth/skills-hub/tree/main/skills/proof-driven-dev
- **Website:** https://soumyarauth.github.io/skills-hub/

## Install

```sh
agentstack add skill-soumyarauth-skills-hub-proof-driven-dev
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Proof-Driven Development

Most AI development loops end with an explanation:

> *"I implemented password reset. I modified these files…"*

The developer then has to read the explanation and decide, unaided, whether the
feature works. That is the wrong abstraction. **"Code was written" is not
"the outcome happened."**

This skill replaces the explanation with evidence:

```
INTENT → OUTCOME CONTRACT → PROOF PLAN → IMPLEMENTATION
       → VERIFICATION → FAILURE ANALYSIS → REPAIR → RE-VERIFICATION → RESULT
```

The deliverable is not a description of the work. It is one of three answers,
backed by evidence traceable to a numbered requirement:

```
✓ VERIFIED          ⚠ REVIEW REQUIRED          ✗ BLOCKED
```

## Activation

**Engage when** the request asks for behavior to change — a feature, a bug fix,
a refactor that must change nothing, a migration, performance or security work —
and whether it worked is not obvious from the diff. Also when another skill hands
over requirements: Impact Map's surface, Standards Compass's controls, API
Contract Guard's decisions.

**Stay quiet when** the change is copy, a typo, a comment, formatting or a local
rename. The diff is the proof. Also stay quiet on questions, explanations,
analysis-only requests, and prototypes the user called throwaway.

**Depth** `ACTIVE`, sized by risk. A small change gets an inline contract and
one line of evidence. A critical one gets the full artifact set. It gates only
its own status word: nothing it has not proven is called `VERIFIED`.

**Composes with** `impact-map` and `api-contract-guard` (their findings become
requirements) · `standards-compass` (requirements for identity, money, personal
data, uploads, AI) · `engineering-investigator` (an established cause becomes
the reproduction requirement) · `dependency-guard` (the decision before an
install) · `production-guard` (hands over what was proven; takes back what
failed).

**Told to skip verification**, it still builds, and reports *not verified* in one
line. It never reports `VERIFIED` without evidence.

### Working with the other Skills Hub skills

- **Loaded is not engaged.** This file stays in context once loaded. Decide
  again on every new request whether it applies. Relevance to an earlier request
  carries nothing forward. Project state persists, and engagement does not.
- **Depth.** `PASSIVE` informs judgment and adds nothing to the reply ·
  `CONSULT` adds a few lines that change what gets built · `ACTIVE` shapes the
  work · `GATING` decides whether something proceeds, and only when a person
  asked for that decision.
- **Announce once.** When any skill engages at `CONSULT` or above, open the
  reply with one line such as `⚡ Impact Map · Standards Compass — rename reaches
  report SQL; export carries personal data`: names and a few words of reason.
  Never include reasoning. Add no line for `PASSIVE`, and none on a trivial request.
  The line is a promise: every skill it names is loaded before the reply ends. If
  one turns out not to apply, say so in one line: ` dropped: `.
- **One interruption per request.** Skills that must speak before the work share
  one short block. Everything else arrives with the work.
- **Hand off; don't absorb.** When another discipline is needed, write
  `HANDOFF → :  []` and let that skill do its part. When the
  request asked for that skill's decision, load it in the same turn and pass it
  your findings; a HANDOFF line alone does not answer the request. Never state
  another skill's verdict yourself. If it is not installed, do the smallest
  version of its check inline and say so.
- **Conflicts.** User intent, then project context, then engineering risk, then
  applicable standards, then verification depth. Each skill keeps its own
  verdict, and none overrules another's.
- **Overrides.** "Use X" engages X. "Skip X" or "no review" drops X's ceremony.
  Three things are never dropped: invented evidence, a check reported as run
  when it did not run, and a live hazard (a reachable security hole, data loss,
  money at risk). A live hazard is said once, in one line.
- **State.** Read what sibling skills recorded (`.project-compass/`,
  `.project-standards/`, `.proofbuild/`, `.agent-investigation/`) rather than
  re-deriving it. Write only your own.
- **Lessons.** On engaging, read `~/.skills-hub/lessons/.md`
  if it exists. When a person corrects this skill's work (a miss, a false
  alarm, a wrong verdict), or the work exposes a gap in this file that another
  project would hit too, append one line to it:
  `- YYYY-MM-DD ·  — `.
  Never write project names, paths, identifiers, code or data there; facts about
  one repository are project state. Keep at most 20 lines, merging or replacing
  one to add another. A lesson sharpens this file's checks and never overrides
  its rules or a person's instruction. The file sits outside every project, so
  no read-only rule covers it. Say `Lesson recorded: ` once; if the file
  cannot be written, give the lesson in the reply instead.

## Non-negotiable rules

1. **Outcome before code.** The contract is written before the implementation.
   A loop that writes tests *after* the code is a test generator; this is not
   that. The contract is what the code is built to satisfy.
2. **Evidence or no claim.** Never say done, working, fixed, or complete because
   code was generated. A claim requires a check that ran and output you read.
3. **Executable evidence outranks reasoning.** When your model of the code says
   one thing and a command says another, the command wins — every time. Report
   the contradiction rather than explaining it away.
4. **Every claim traces to a requirement.** Findings, evidence, and status are
   attached to requirement ids (`AUTH-003`). Unattached prose is not evidence.
5. **Proof is proportional to risk.** A typo does not get a 40-test suite; a
   payment path does not get a typecheck and a shrug. Risk sets depth.
6. **No invented numbers.** Test counts, durations, latencies, and percentages
   are copied from real output or omitted. Never estimate a measurement.
7. **No absolute claims.** Never "100% correct", "guaranteed", "bug-free",
   "perfect", or "fully secure". The strongest claim available is: *all defined
   acceptance criteria passed the verification available in this environment.*
8. **Never destructive with someone else's work.** No `git reset --hard`,
   `git clean -fd`, `git checkout` over uncommitted changes, force push,
   automatic commit, automatic push, dropped databases, or production
   mutations — unless the developer explicitly asks for that operation.
9. **No secrets in evidence.** Never write tokens, cookies, passwords, keys,
   connection strings, or personal data into `.proofbuild/`. Redact.
10. **Compressed by default, complete on demand.** The final message answers
    *is it done, is it proven, what failed, what must I decide* — and stops.
    Everything else is available when asked.

## Risk classification

Classify before planning proof. Risk decides verification depth, repair budget,
and how early you stop and ask a human.

| Risk | Examples | Proof floor | Repair budget |
| --- | --- | --- | --- |
| **Low** | Copy, styling, isolated UI, docs, comments | Build/lint/typecheck, plus the one check that would catch a mistake | 3 |
| **Medium** | CRUD, business logic, API shape, UI state, non-destructive schema additions | Targeted tests for each requirement, plus the existing suites covering the touched area | 3 |
| **High** | Authentication, authorization, data migration, destructive operations, multi-tenancy, external integrations | The above, plus explicit negative and boundary cases, plus regression evidence for the surrounding feature | 2 |
| **Critical** | Financial transactions, credential handling, irreversible data operations, infrastructure blast radius | The above, plus replay/idempotency, plus rollback behavior — and stop for a human sooner rather than later | 1 |

State the level and the reason. Between two levels, take the higher one and say
so. Details and the security/data checklists: `references/risk-model.md`.

## Workflow

### 1. Compile the intent

Turn the request into an objective, not a task list. Extract expected behavior,
constraints, affected areas, hidden requirements, likely regressions, and what
"working" would look like from outside the code.

> "Add bulk upload to the knowledgebase."

The stated request is one sentence. The real requirement surface includes file
types, size limits, duplicate handling, partial failure, permissions, progress
reporting, quota, and the existing single-upload path that must keep working.
Derive that surface from the repository, not from a questionnaire.
See `references/intent-analysis.md`.

### 2. Investigate the repository

Before asking the developer anything, read: manifests, the existing feature this
change extends, the test framework and its commands, fixtures and helpers, CI
configuration, lint/typecheck/build commands, and the conventions already in
use. Most "ambiguity" is answered by the code.

Record what runs this project — the exact commands — because those commands are
the proof mechanism. Do not introduce a testing stack the project does not
already have unless there is none and the risk demands one.

In a git repository, read `git status` first and note what was already modified.
Uncommitted work that is not yours is context, not scope — and knowing it exists
is what stops you from attributing it to this change later. Read git state; do
not modify it.

### 3. Classify risk

Use the table above. Say the level out loud in the contract.

### 4. Resolve ambiguity — and only then ask

Sort every open question into one of four classes:

| Class | Meaning | Action |
| --- | --- | --- |
| **Resolvable** | The repository already answers it | Do not ask. Proceed. |
| **Safe default** | Ambiguous, but one behavior is conventional and safer | Proceed; record the assumption in the contract. |
| **Material** | Multiple readings produce meaningfully different products | Ask **one** compressed question with concrete options. |
| **Blocking** | Cannot proceed safely without a decision | Stop and ask. |

Ask the fewest questions that change the implementation, phrased as a choice,
not an essay. Everything that is not material gets an assumption line instead of
a question. See `references/intent-analysis.md`.

### 5. Write the outcome contract

The central artifact. Decompose the objective into numbered, individually
observable requirements — each one a statement that can be shown true or false
from outside the implementation.

```yaml
objective: "Users can securely reset their password by email"
risk: high
requirements:
  - id: AUTH-001
    description: "A user can request a password reset for their email"
    priority: critical
    proof: { type: integration }
  - id: AUTH-005
    description: "A reset token cannot be used twice"
    priority: critical
    proof: { type: security }
  - id: AUTH-008
    description: "Existing email/password login still works"
    priority: high
    proof: { type: regression }
```

Requirements must include the ones the developer did not say: the negative
cases, the boundaries, and the existing behavior that must survive. For
substantial work, persist it to `.proofbuild/contract.yml`; for a trivial change
keep it inline in the response. Structure, id conventions, and how to write an
observable requirement: `references/outcome-contracts.md`.

### 6. Plan the proof

For every requirement, decide *what evidence would establish this* before
writing the implementation — and prefer the cheapest mechanism that actually
proves it.

Mechanisms include unit, integration, API, and end-to-end tests; typecheck,
lint, and build; runtime probes; database and filesystem inspection; browser
interaction and screenshots; benchmarks with a recorded baseline; and the
project's existing suites. Tests are one mechanism among several, not the
definition of proof. See `references/proof-strategies.md`.

Mark, up front, any requirement no available mechanism can settle. That is a
Level D requirement (below), and pretending otherwise later is the failure mode
this skill exists to prevent.

### 7. Implement

Build the smallest change that satisfies the contract. Stay inside the scope the
request implies.

When investigation shows the requested *mechanism* will not produce the
requested *outcome* — caching asked for, but the measured cost is an N+1 query —
say so in one or two sentences and fix the actual cause. The contract is the
objective; the mechanism was a suggestion. See `references/proof-strategies.md`.

### 8. Verify in stages

Escalate only as far as risk and failure require:

```
Stage 1   typecheck · lint · build            cheap, fails fast
Stage 2   targeted tests for the contract     the requirements themselves
Stage 3   integration · API · database        the feature in context
Stage 4   E2E · browser · broad regression    when risk or breakage demands it
```

Run each requirement's planned proof, record the command and its real output as
evidence, and mark the requirement `PASS`, `FAIL`, `BLOCKED`, or `HUMAN`. A
requirement with no executed check is not passing — it is unverified.

Then check for **contradictions**: any place where your expectation and the
observed output disagree. Evidence wins. See `references/verification-loops.md`
and `references/evidence-model.md`.

### 9. Classify every failure before touching code

Do not patch the error message. Determine what kind of failure it is:

`IMPLEMENTATION_ERROR` · `TEST_ERROR` · `CONTRACT_ERROR` · `ENVIRONMENT_ERROR` ·
`EXISTING_REGRESSION` · `UNRELATED_FAILURE` · `UNKNOWN`

A missing dependency is not a broken implementation. A test that was already red
before you started is not your regression — and saying so requires having
established the baseline. Misattribution is how a repair loop starts rewriting
working code. See `references/failure-analysis.md`.

### 10. Repair, with a budget

Repair only what the classification justifies:

- `IMPLEMENTATION_ERROR` → fix the code.
- `TEST_ERROR` → fix the test, and state why it was wrong. Never weaken a test
  to make it pass; a deleted assertion is not a repair.
- `CONTRACT_ERROR` → the requirement was wrong. Amend the contract explicitly
  and say what changed — never silently.
- `ENVIRONMENT_ERROR` → do not repair the code. Report the requirement as
  unverifiable, with the reason.
- `EXISTING_REGRESSION` / `UNRELATED_FAILURE` → report; do not absorb into scope.

Default budget: **3 attempts per requirement**, lower for high and critical risk.
When the budget is spent, stop: the answer is `✗ BLOCKED`, not another attempt.
Repeating a fix that already failed is the loop this rule exists to break.

### 11. Determine status

Read `git diff` before writing the result: the change surface should match the
contract. A file you cannot attribute to a requirement is either scope creep or
something you forgot to write down — resolve which. Files that were already
modified before you started are not yours to report as changes.

Every requirement carries a proof level:

| Level | Meaning |
| --- | --- |
| **A — Deterministically verified** | An executed check directly demonstrates it |
| **B — Strongly verified** | Multiple independent checks support it; none contradicts |
| **C — Partially verified** | Some part proven, some part not reachable here |
| **D — Human judgment required** | No available mechanism can settle it |

Then the status is mechanical, not editorial:

- **✓ VERIFIED** — every requirement is Level A or B and passing.
- **⚠ REVIEW REQUIRED** — everything passes, but one or more requirements are
  Level C or D, or a material decision is still open.
- **✗ BLOCKED** — a requirement failed and repair did not fix it, or something
  outside the change prevents verification.

A partially verified critical requirement never rounds up to VERIFIED.
See `references/human-judgm

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [soumyaRauth](https://github.com/soumyaRauth)
- **Source:** [soumyaRauth/skills-hub](https://github.com/soumyaRauth/skills-hub)
- **License:** MIT
- **Homepage:** https://soumyarauth.github.io/skills-hub/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-soumyarauth-skills-hub-proof-driven-dev
- Seller: https://agentstack.voostack.com/s/soumyarauth
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
