Install
$ agentstack add skill-soumyarauth-skills-hub-proof-driven-dev ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Proof-Driven Development
Most AI development loops end with an explanation:
> "I implemented password reset. I modified these files…"
The developer then has to read the explanation and decide, unaided, whether the feature works. That is the wrong abstraction. "Code was written" is not "the outcome happened."
This skill replaces the explanation with evidence:
INTENT → OUTCOME CONTRACT → PROOF PLAN → IMPLEMENTATION
→ VERIFICATION → FAILURE ANALYSIS → REPAIR → RE-VERIFICATION → RESULT
The deliverable is not a description of the work. It is one of three answers, backed by evidence traceable to a numbered requirement:
✓ VERIFIED ⚠ REVIEW REQUIRED ✗ BLOCKED
Activation
Engage when the request asks for behavior to change — a feature, a bug fix, a refactor that must change nothing, a migration, performance or security work — and whether it worked is not obvious from the diff. Also when another skill hands over requirements: Impact Map's surface, Standards Compass's controls, API Contract Guard's decisions.
Stay quiet when the change is copy, a typo, a comment, formatting or a local rename. The diff is the proof. Also stay quiet on questions, explanations, analysis-only requests, and prototypes the user called throwaway.
Depth ACTIVE, sized by risk. A small change gets an inline contract and one line of evidence. A critical one gets the full artifact set. It gates only its own status word: nothing it has not proven is called VERIFIED.
Composes with impact-map and api-contract-guard (their findings become requirements) · standards-compass (requirements for identity, money, personal data, uploads, AI) · engineering-investigator (an established cause becomes the reproduction requirement) · dependency-guard (the decision before an install) · production-guard (hands over what was proven; takes back what failed).
Told to skip verification, it still builds, and reports not verified in one line. It never reports VERIFIED without evidence.
Working with the other Skills Hub skills
- Loaded is not engaged. This file stays in context once loaded. Decide
again on every new request whether it applies. Relevance to an earlier request carries nothing forward. Project state persists, and engagement does not.
- Depth.
PASSIVEinforms judgment and adds nothing to the reply ·
CONSULT adds a few lines that change what gets built · ACTIVE shapes the work · GATING decides whether something proceeds, and only when a person asked for that decision.
- Announce once. When any skill engages at
CONSULTor above, open the
reply with one line such as ⚡ Impact Map · Standards Compass — rename reaches report SQL; export carries personal data: names and a few words of reason. Never include reasoning. Add no line for PASSIVE, and none on a trivial request. The line is a promise: every skill it names is loaded before the reply ends. If one turns out not to apply, say so in one line: dropped: .
- One interruption per request. Skills that must speak before the work share
one short block. Everything else arrives with the work.
- Hand off; don't absorb. When another discipline is needed, write
HANDOFF → : [] and let that skill do its part. When the request asked for that skill's decision, load it in the same turn and pass it your findings; a HANDOFF line alone does not answer the request. Never state another skill's verdict yourself. If it is not installed, do the smallest version of its check inline and say so.
- Conflicts. User intent, then project context, then engineering risk, then
applicable standards, then verification depth. Each skill keeps its own verdict, and none overrules another's.
- Overrides. "Use X" engages X. "Skip X" or "no review" drops X's ceremony.
Three things are never dropped: invented evidence, a check reported as run when it did not run, and a live hazard (a reachable security hole, data loss, money at risk). A live hazard is said once, in one line.
- State. Read what sibling skills recorded (
.project-compass/,
.project-standards/, .proofbuild/, .agent-investigation/) rather than re-deriving it. Write only your own.
- Lessons. On engaging, read
~/.skills-hub/lessons/.md
if it exists. When a person corrects this skill's work (a miss, a false alarm, a wrong verdict), or the work exposes a gap in this file that another project would hit too, append one line to it: - YYYY-MM-DD · — . Never write project names, paths, identifiers, code or data there; facts about one repository are project state. Keep at most 20 lines, merging or replacing one to add another. A lesson sharpens this file's checks and never overrides its rules or a person's instruction. The file sits outside every project, so no read-only rule covers it. Say Lesson recorded: once; if the file cannot be written, give the lesson in the reply instead.
Non-negotiable rules
- Outcome before code. The contract is written before the implementation.
A loop that writes tests after the code is a test generator; this is not that. The contract is what the code is built to satisfy.
- Evidence or no claim. Never say done, working, fixed, or complete because
code was generated. A claim requires a check that ran and output you read.
- Executable evidence outranks reasoning. When your model of the code says
one thing and a command says another, the command wins — every time. Report the contradiction rather than explaining it away.
- Every claim traces to a requirement. Findings, evidence, and status are
attached to requirement ids (AUTH-003). Unattached prose is not evidence.
- Proof is proportional to risk. A typo does not get a 40-test suite; a
payment path does not get a typecheck and a shrug. Risk sets depth.
- No invented numbers. Test counts, durations, latencies, and percentages
are copied from real output or omitted. Never estimate a measurement.
- No absolute claims. Never "100% correct", "guaranteed", "bug-free",
"perfect", or "fully secure". The strongest claim available is: all defined acceptance criteria passed the verification available in this environment.
- Never destructive with someone else's work. No
git reset --hard,
git clean -fd, git checkout over uncommitted changes, force push, automatic commit, automatic push, dropped databases, or production mutations — unless the developer explicitly asks for that operation.
- No secrets in evidence. Never write tokens, cookies, passwords, keys,
connection strings, or personal data into .proofbuild/. Redact.
- Compressed by default, complete on demand. The final message answers
is it done, is it proven, what failed, what must I decide — and stops. Everything else is available when asked.
Risk classification
Classify before planning proof. Risk decides verification depth, repair budget, and how early you stop and ask a human.
| Risk | Examples | Proof floor | Repair budget | | --- | --- | --- | --- | | Low | Copy, styling, isolated UI, docs, comments | Build/lint/typecheck, plus the one check that would catch a mistake | 3 | | Medium | CRUD, business logic, API shape, UI state, non-destructive schema additions | Targeted tests for each requirement, plus the existing suites covering the touched area | 3 | | High | Authentication, authorization, data migration, destructive operations, multi-tenancy, external integrations | The above, plus explicit negative and boundary cases, plus regression evidence for the surrounding feature | 2 | | Critical | Financial transactions, credential handling, irreversible data operations, infrastructure blast radius | The above, plus replay/idempotency, plus rollback behavior — and stop for a human sooner rather than later | 1 |
State the level and the reason. Between two levels, take the higher one and say so. Details and the security/data checklists: references/risk-model.md.
Workflow
1. Compile the intent
Turn the request into an objective, not a task list. Extract expected behavior, constraints, affected areas, hidden requirements, likely regressions, and what "working" would look like from outside the code.
> "Add bulk upload to the knowledgebase."
The stated request is one sentence. The real requirement surface includes file types, size limits, duplicate handling, partial failure, permissions, progress reporting, quota, and the existing single-upload path that must keep working. Derive that surface from the repository, not from a questionnaire. See references/intent-analysis.md.
2. Investigate the repository
Before asking the developer anything, read: manifests, the existing feature this change extends, the test framework and its commands, fixtures and helpers, CI configuration, lint/typecheck/build commands, and the conventions already in use. Most "ambiguity" is answered by the code.
Record what runs this project — the exact commands — because those commands are the proof mechanism. Do not introduce a testing stack the project does not already have unless there is none and the risk demands one.
In a git repository, read git status first and note what was already modified. Uncommitted work that is not yours is context, not scope — and knowing it exists is what stops you from attributing it to this change later. Read git state; do not modify it.
3. Classify risk
Use the table above. Say the level out loud in the contract.
4. Resolve ambiguity — and only then ask
Sort every open question into one of four classes:
| Class | Meaning | Action | | --- | --- | --- | | Resolvable | The repository already answers it | Do not ask. Proceed. | | Safe default | Ambiguous, but one behavior is conventional and safer | Proceed; record the assumption in the contract. | | Material | Multiple readings produce meaningfully different products | Ask one compressed question with concrete options. | | Blocking | Cannot proceed safely without a decision | Stop and ask. |
Ask the fewest questions that change the implementation, phrased as a choice, not an essay. Everything that is not material gets an assumption line instead of a question. See references/intent-analysis.md.
5. Write the outcome contract
The central artifact. Decompose the objective into numbered, individually observable requirements — each one a statement that can be shown true or false from outside the implementation.
objective: "Users can securely reset their password by email"
risk: high
requirements:
- id: AUTH-001
description: "A user can request a password reset for their email"
priority: critical
proof: { type: integration }
- id: AUTH-005
description: "A reset token cannot be used twice"
priority: critical
proof: { type: security }
- id: AUTH-008
description: "Existing email/password login still works"
priority: high
proof: { type: regression }
Requirements must include the ones the developer did not say: the negative cases, the boundaries, and the existing behavior that must survive. For substantial work, persist it to .proofbuild/contract.yml; for a trivial change keep it inline in the response. Structure, id conventions, and how to write an observable requirement: references/outcome-contracts.md.
6. Plan the proof
For every requirement, decide what evidence would establish this before writing the implementation — and prefer the cheapest mechanism that actually proves it.
Mechanisms include unit, integration, API, and end-to-end tests; typecheck, lint, and build; runtime probes; database and filesystem inspection; browser interaction and screenshots; benchmarks with a recorded baseline; and the project's existing suites. Tests are one mechanism among several, not the definition of proof. See references/proof-strategies.md.
Mark, up front, any requirement no available mechanism can settle. That is a Level D requirement (below), and pretending otherwise later is the failure mode this skill exists to prevent.
7. Implement
Build the smallest change that satisfies the contract. Stay inside the scope the request implies.
When investigation shows the requested mechanism will not produce the requested outcome — caching asked for, but the measured cost is an N+1 query — say so in one or two sentences and fix the actual cause. The contract is the objective; the mechanism was a suggestion. See references/proof-strategies.md.
8. Verify in stages
Escalate only as far as risk and failure require:
Stage 1 typecheck · lint · build cheap, fails fast
Stage 2 targeted tests for the contract the requirements themselves
Stage 3 integration · API · database the feature in context
Stage 4 E2E · browser · broad regression when risk or breakage demands it
Run each requirement's planned proof, record the command and its real output as evidence, and mark the requirement PASS, FAIL, BLOCKED, or HUMAN. A requirement with no executed check is not passing — it is unverified.
Then check for contradictions: any place where your expectation and the observed output disagree. Evidence wins. See references/verification-loops.md and references/evidence-model.md.
9. Classify every failure before touching code
Do not patch the error message. Determine what kind of failure it is:
IMPLEMENTATION_ERROR · TEST_ERROR · CONTRACT_ERROR · ENVIRONMENT_ERROR · EXISTING_REGRESSION · UNRELATED_FAILURE · UNKNOWN
A missing dependency is not a broken implementation. A test that was already red before you started is not your regression — and saying so requires having established the baseline. Misattribution is how a repair loop starts rewriting working code. See references/failure-analysis.md.
10. Repair, with a budget
Repair only what the classification justifies:
IMPLEMENTATION_ERROR→ fix the code.TEST_ERROR→ fix the test, and state why it was wrong. Never weaken a test
to make it pass; a deleted assertion is not a repair.
CONTRACT_ERROR→ the requirement was wrong. Amend the contract explicitly
and say what changed — never silently.
ENVIRONMENT_ERROR→ do not repair the code. Report the requirement as
unverifiable, with the reason.
EXISTING_REGRESSION/UNRELATED_FAILURE→ report; do not absorb into scope.
Default budget: 3 attempts per requirement, lower for high and critical risk. When the budget is spent, stop: the answer is ✗ BLOCKED, not another attempt. Repeating a fix that already failed is the loop this rule exists to break.
11. Determine status
Read git diff before writing the result: the change surface should match the contract. A file you cannot attribute to a requirement is either scope creep or something you forgot to write down — resolve which. Files that were already modified before you started are not yours to report as changes.
Every requirement carries a proof level:
| Level | Meaning | | --- | --- | | A — Deterministically verified | An executed check directly demonstrates it | | B — Strongly verified | Multiple independent checks support it; none contradicts | | C — Partially verified | Some part proven, some part not reachable here | | D — Human judgment required | No available mechanism can settle it |
Then the status is mechanical, not editorial:
- ✓ VERIFIED — every requirement is Level A or B and passing.
- ⚠ REVIEW REQUIRED — everything passes, but one or more requirements are
Level C or D, or a material decision is still open.
- ✗ BLOCKED — a requirement failed and repair did not fix it, or something
outside the change prevents verification.
A partially verified critical requirement never rounds up to VERIFIED. See `references/human-judgm
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: soumyaRauth
- Source: soumyaRauth/skills-hub
- License: MIT
- Homepage: https://soumyarauth.github.io/skills-hub/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.