Install
$ agentstack add skill-waitdeadai-minmaxing-verify ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
/verify
THE VERIFIER — mandatory output validation against SPEC.md before accepting. This is an Independent verification pass with evidence for each criterion.
Isolation rule: Do not claim a different model, agent, process, or workspace performed verification unless the workflow artifact records metadata that proves it. When isolation metadata is unavailable, describe this as an independent verification pass against the spec, not as guaranteed separate execution.
MAXPARALLELAGENTS — ceiling for verification lanes. Split verification only when criteria or surfaces are independent enough to check separately.
Use when: After every implementation task, code changes, documentation changes, config changes, before shipping, "swarm verify", or whenever you need to validate output.
Swarm: "swarm verify" → /verify with an efficacy-first verification wave up to MAX_PARALLEL_AGENTS.
Never skip verification. If you can't prove it passes, it fails.
Purpose
Verify implementation output against SPEC.md. This is the critical quality gate that prevents spec drift and catches implementation errors early.
The Problem: Implementation verifies itself too casually → confirmation bias → bugs ship.
The Solution: Run a deliberately adversarial pass against SPEC.md. Find the flaws before the user does.
Execution Protocol
Step 1: Locate SPEC.md
- Find the relevant SPEC.md for this task
- If no SPEC.md exists → FAIL immediately:
`` VERIFICATION FAILED: No SPEC.md found. Cannot verify. Create SPEC.md first via /autoplan. ``
Step 2: Read Success Criteria
Extract all criteria from SPEC.md:
## Success Criteria
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
For non-trivial planning work, also read ## Agent-Native Estimate. If it is missing, human-equivalent only, or omits confidence, verification/review time, or blockers, mark the planning contract incomplete before accepting closeout.
For non-trivial file-changing work, also read ## Spec QA in the workflow artifact or the matching .taste/specqa/{run_id}/spec-qa.json. If /specqa is missing, if it did not run after SPEC.md and before implementation, if it blocked execution, or if it claims Opus 4.7 review without runtime identity proof, mark the planning contract incomplete before accepting closeout.
Step 3: Verify Each Criterion
For each criterion, perform verification:
For code changes:
- Read the changed files
- Run relevant tests
- Execute the code if applicable
- Inspect output against criterion
For documentation:
- Read the doc
- Verify structure matches spec
- Check all required sections exist
For config changes:
- Read the config file
- Verify values match spec
- Test that config takes effect
For /parallel runs:
- Read the parallel artifact when present
- If
.taste/parallel/{run_id}exists, run
scripts/parallel-aggregate.sh .taste/parallel/{run_id} and treat failure as verification failure
- Verify every worker packet returned the Worker Result Schema
- Check changed files against the ownership matrix
- Confirm sync barriers were honored before dependent work
- Confirm the effective budget did not exceed the hardware capacity profile,
MAX_PARALLEL_AGENTS, or Codexmax_threads - Treat worker summaries as claims until aggregate evidence proves them
- Reject worker success when the parent cannot cite command evidence or an
inspection finding that proves the aggregate result.
Step 4: Verification Results
## Verification Results: [Task Name]
### Against SPEC.md: /path/to/SPEC.md
### Verification Metadata
- Executor identity/model/workspace: [known value or unknown]
- Verifier identity/model/workspace: [known value or unknown]
- Isolation status: [proved separate / same session independent pass / unknown]
| Criterion | Status | Evidence |
|-----------|--------|----------|
| Criterion 1 | PASS | [test output, command result] |
| Criterion 2 | PASS | [inspection finding] |
| Criterion 3 | FAIL | Expected X, got Y |
### Verification Details
**[Criterion 1]**: PASS
- Expected: [what spec says]
- Actual: [what implementation does]
- Evidence: [test output, command, inspection]
**[Criterion 2]**: PASS
- Expected: [what spec says]
- Actual: [what implementation does]
- Evidence: [test output, command, inspection]
**[Criterion 3]**: FAIL
- Expected: [what spec says]
- Actual: [what implementation does]
- Evidence: [exact failure]
### Overall Result
- **ACCEPT** — All criteria pass
- **REJECT** — One or more criteria fail
- **CONDITIONAL** — Minor issues, human decision needed
Step 5: If REJECT
List specific failures and recommend fixes:
### Failures Detected
1. **[Criterion] failed**: Expected X, got Y
- Fix: [specific change needed]
2. **[Criterion] failed**: [reason]
- Fix: [specific change needed]
### Recommended Fixes
1. [specific fix for failure 1]
2. [specific fix for failure 2]
Step 5.5: Log Failures to Memory
On REJECT, log each failed criterion to error-solution tier:
# Log each failure as error-solution pair
bash scripts/memory.sh add error-solution "\"Failed criterion: [criterion]\"" "\"Fix: [recommended fix]\""
# Record failure in causal graph
python3 -c "
from memory.causal import record_outcome
factors = ['spec_drift', 'verify_failed', '[specific_failed_criterion]']
record_outcome(factors, 'failure')
" 2>/dev/null || echo "record_outcome: skipped"
Evidence Requirements
| Result | Evidence Required | |--------|-----------------| | PASS | Test output, command output, screenshot, inspection finding | | FAIL | Expected value, actual value, specific difference | | CONDITIONAL | What is uncertain, what would resolve it |
For /parallel, PASS also requires packet evidence, ownership-matrix checks, capacity checks, sync-barrier checks, and aggregate verification against SPEC.md. When run-level JSON sidecars exist, PASS also requires a passing scripts/parallel-aggregate.sh .taste/parallel/{run_id} result that reports effective lanes, bottleneck, worker result count, and critical path.
NOT acceptable as evidence:
- "Looks good"
- "Seems to work"
- "I think it's correct"
- No evidence at all
Quality Gates
- Must read SPEC.md before verifying
- Must verify ALL criteria, not just "most"
- Must provide EVIDENCE for each pass/fail
- Silent pass is not allowed — show evidence
- Silent fail is not allowed — show exactly what failed
- No SPEC.md = automatic FAIL
- Non-trivial planning work without a valid
Agent-Native Estimate= FAIL - Non-trivial file-changing work without valid
/specqaevidence = FAIL - Evidence-free closeout, failed-verification positive closeout, fake source
ledgers, and "tests passed" without command output or equivalent durable evidence = FAIL
- Machine-consumed estimate, verification, or worker result sidecars must pass
scripts/artifact-lint.sh when present.
- Harness-improvement claims should run
scripts/harness-eval.sh --jsonwhen
the static eval pack applies. Treat it as local gate coverage, not model-running benchmark coverage.
- Harness-health claims should cite
scripts/run-metrics.sh --jsonor
scripts/session-insights.sh --json; missing provider/cost/token data must remain insufficient_data.
- Parallel worker-output claims should cite
scripts/parallel-aggregate.sh
when .taste/parallel/{run_id} artifacts exist; failed aggregation means the parent has not verified the parallel result.
Anti-Patterns
- Verifying without SPEC.md → FAIL
- Skipping criteria → FAIL
- "Looks good" without evidence → FAIL
- Accepting output that doesn't match spec → FAIL
- Saying "mostly done" without specifics → FAIL
- Silent acceptance → FAIL
- Accepting human-equivalent-only estimates or hidden blockers → FAIL
Verifier Mindset
You are adversarial to the implementation. Your job is to find flaws.
- If you can't prove it passes → it FAILS
- Missing evidence → treat as FAIL
- Vague claims → demand specifics
- "Works fine" without test → FAIL
- Implementation self-verification without a separate evidence pass → BLOCK, use this skill
Chain Contract
This skill is a verification playbook. /workflow may reuse this guidance, but it should not depend on invoking /verify as a guaranteed nested continuation step.
When invoked directly by the user, return ACCEPT or REJECT with evidence.
When /workflow references this skill, the parent workflow decides whether to fix, re-verify, close out locally, or ship.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: waitdeadai
- Source: waitdeadai/minmaxing
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.