Install
$ agentstack add skill-michelkerkmeester-skilled-harness-spec-driven-agent-loops-sk-code-review ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ● Filesystem access Used
- ● Shell / process execution Used
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
code-review Mode - Stack-Agnostic Findings-First Review
Universal findings-first review baseline paired with sk-code surface standards evidence for the detected code surface.
1. WHEN TO USE
Activation Triggers
Use the code-review mode (of the sk-code family) when:
- A user asks for code review, PR review, quality gate, or merge readiness.
- A workflow dispatches
@reviewfor pre-commit or gate validation. - A user requests security/correctness risk analysis before merge.
- A user wants severity-ranked findings with file:line evidence.
Keyword Triggers
review, code review, pr review, audit, security review, quality gate, request changes, findings, blocking issues, merge readiness
Use Cases
- Review-only pass: findings-first output with no code edits.
- Gate validation: score + pass/fail recommendation for orchestrated workflows.
- Focused risk pass: security, concurrency, correctness, or removal-focused review.
When NOT to Use
- Feature implementation without review intent; use the surface skill (
code-webflow/code-opencode). - Pure documentation editing where code behavior is not being assessed.
- Git-only workflow tasks (branching, rebasing, commit hygiene) without code-quality evaluation intent.
- Applying review fixes after findings are accepted; use the surface skill (
code-webflow/code-opencode). - Author-side quality gates before review; use
code-quality. - Root-cause debugging; use the surface's
workflow-debug.mddoctrine. - Verification evidence collection; use the surface's
workflow-verify.mddoctrine.
2. SMART ROUTING
Primary Detection Signal
Review behavior follows a baseline+surface-evidence model:
- Baseline (always): the
code-reviewmode (of the sk-code family) findings-first doctrine. - Surface standards evidence (when available):
sk-codedetected surface resources. - Unknown surfaces: review against baseline security/correctness only and disclose uncertainty.
Phase Detection
TASK CONTEXT
|
+- STEP 0: Load the `code-review` mode baseline + `sk-code` surface evidence. The dispatcher / agent assembling the code-review prompt MUST prepend `CODE-REVIEW\n\n` as the first two lines of the rendered prompt before the reviewer LLM sees it. Reference resources stay unchanged.
+- STEP 1: Score intents (top-2 when ambiguity delta str:
return " ".join([
str(getattr(task, "text", "")),
str(getattr(task, "query", "")),
str(getattr(task, "description", "")),
" ".join(getattr(task, "keywords", []) or []),
]).lower()
def _guard_in_skill(relative_path: str) -> str:
resolved = (SKILL_ROOT / relative_path).resolve()
resolved.relative_to(SKILL_ROOT)
if resolved.suffix.lower() != ".md":
raise ValueError(f"Only markdown resources are routable: {relative_path}")
return resolved.relative_to(SKILL_ROOT).as_posix()
def discover_markdown_resources() -> set[str]:
docs = []
for base in RESOURCE_BASES:
if base.exists():
docs.extend(path for path in base.rglob("*.md") if path.is_file())
return {doc.relative_to(SKILL_ROOT).as_posix() for doc in docs}
def keyword_present(keyword: str, text: str) -> bool:
"""Boundary-aware match: bare substrings misroute ('pr' in 'improve prompt')."""
return re.search(rf"(? dict[str, float]:
text = _task_text(task)
scores = {intent: 0.0 for intent in INTENT_SIGNALS}
for intent, cfg in INTENT_SIGNALS.items():
for keyword in cfg["keywords"]:
if keyword_present(keyword, text):
scores[intent] += cfg["weight"]
return scores
def select_intents(scores: dict[str, float], ambiguity_delta: float = 1.0, max_intents: int = 2) -> list[str]:
ranked = sorted(scores.items(), key=lambda item: item[1], reverse=True)
if not ranked or ranked[0][1] 1 and ranked[1][1] > 0 and (ranked[0][1] - ranked[1][1]) str:
text = _task_text(task)
files = " ".join((workspace_files or []) + (changed_files or [])).lower()
if ".opencode/" in files or keyword_present("jsonc", text) or keyword_present("mcp", text):
return "sk-code:code-opencode"
if any(keyword_present(term, text) for term in ["frontend", "web", "css", "dom", "browser"]) or any(
marker in files for marker in ["next.config", "vite.config", "package.json", "src/"]
):
return "sk-code:code-webflow"
return "sk-code:unknown"
def route_review_resources(task, workspace_files=None, changed_files=None):
inventory = discover_markdown_resources()
text = _task_text(task)
scores = score_intents(task)
intents = select_intents(scores, ambiguity_delta=1.0)
loaded = []
seen = set()
def load_if_available(relative_path: str) -> None:
guarded = _guard_in_skill(relative_path)
if guarded in inventory and guarded not in seen:
load(guarded)
loaded.append(guarded)
seen.add(guarded)
for relative_path in DEFAULT_RESOURCES:
load_if_available(relative_path)
if sum(scores.values()) =4` or any other numeric threshold as a blocker.
#### Instance-Only Opt-Out
A finding may use the narrow fix path only when all are true:
- It is not P0/P1 security, path, auth/authz, sandboxing, env precedence, schema, persistence, or public-response behavior.
- `rg` proves no same-class producer or consumer.
- Verification is local and cheap: one focused test, one doc row, or one static audit command.
- The fix response includes the exact command evidence for the opt-out.
Otherwise, run the full fix completeness checklist.
### Phase 4: Output and Next Action
Required output contract:
```markdown
## Code Review Summary
**Files reviewed**: X files, Y lines changed
**Overall assessment**: [APPROVE / REQUEST_CHANGES / COMMENT]
**Baseline used**: [sk-code (`code-review`)]
**Surface evidence used**: [sk-code:code-webflow | sk-code:code-opencode | sk-code:unknown]
## Findings
### P0 - Critical
1. [path:line] Title
- Risk
- User impact
- Finding class: [instance-only | class-of-bug | cross-consumer | algorithmic | matrix/evidence | test-isolation]
- Scope proof: [grep/test evidence proving class coverage or instance-only status]
- affectedSurfaceHints: [optional string array of short producer/consumer surface names; recommended for actionable findings, required for cross-consumer findings]
- riskScore: [optional advisory number only; never gating]
- Recommended fix
### P1 - High
...
## Removal/Iteration Plan
## Next Steps
After reporting findings, request explicit next action before any implementation follow-up.
Final-line exact-string contract (MANDATORY)
Every review MUST end with exactly one of the following plain-text lines as the absolute final line of the output (no trailing whitespace, no variation):
Review status: APPROVED
Review status: REQUESTED_CHANGES
Review status: COMMENTED
Example output bottom:
...
## Next Steps
1. Fix the null-deref at src/foo.ts:42
2. Add input validation for the `/api/bar` endpoint
Review status: REQUESTED_CHANGES
Downstream automation parses this final line via exact string match — do not vary the format, add trailing punctuation, or wrap in Markdown formatting. The sole exception is the documented M-1 / M-2 skip output (§9): those lines begin with the exact Review status: COMMENTED and append a parenthetical reason, so a leading-verdict (grep / startsWith) parse still yields COMMENTED. A normal review must still end with one of the three exact lines above.
4. RULES
✅ ALWAYS
- Keep findings first; summaries follow findings.
- Enforce baseline security/correctness minimums regardless of surface.
- Include file:line evidence for actionable findings.
- State assumptions when evidence is incomplete.
- Identify
sk-codesurface evidence used for standards alignment.
⛔ NEVER
- Override surface-specific conventions with generic baseline style preferences.
- Approve code with unaddressed P0 security/correctness defects.
- Produce vague findings without concrete evidence.
- Mix unrelated cleanup into targeted fix recommendations.
- Do not implement fixes during review. Report findings only; implementation is a separate follow-up step.
⚠️ ESCALATE IF
- Surface detection is ambiguous and affects standards or verification commands.
- Baseline and surface guidance conflict in a non-deterministic way.
- Large diff size prevents reliable severity assignment without narrowed scope.
- Requested remediation exceeds review scope and becomes architecture redesign.
5. REFERENCES
Core References
- [review-core.md](./references/review-core.md) - Shared review doctrine: severity model, evidence rules, precedence, and finding schema.
- [review-ux-single-pass.md](./references/review-ux-single-pass.md) - Interactive single-pass review flow, presentation modes, and PR/pre-commit behavior.
- [quick-reference.md](./references/quick-reference.md) - Lightweight index for routing between shared doctrine and single-pass UX guidance.
- [pr-state-dedup.md](./references/pr-state-dedup.md) - Content-hash deduplication for unchanged pull-request reviews.
- [security-checklist.md](./assets/security-checklist.md) - Mandatory security and reliability checks.
- [code-quality-checklist.md](./assets/code-quality-checklist.md) - Correctness, performance, KISS, and DRY checks.
- [solid-checklist.md](./assets/solid-checklist.md) - SOLID (SRP/OCP/LSP/ISP/DIP) and architecture assessment prompts.
- [removal-plan.md](./assets/removal-plan.md) - Safe-now vs deferred removal planning template.
- [test-quality-checklist.md](./assets/test-quality-checklist.md) - Test quality, coverage, and anti-pattern detection.
Reference Loading Notes
- Load only the references needed for the selected intents.
- Keep Section 2 (
SMART ROUTING) as the authoritative routing source.
6. SUCCESS CRITERIA
- Review output is findings-first and severity-ordered.
code-reviewmode baseline +sk-codesurface evidence contract is explicit in report context.- Security/correctness minimums are always covered.
- Recommended fixes are actionable and scope-proportional.
7. INTEGRATION POINTS
- Primary review baseline for
@reviewagents in.opencode/agents/review.md. - Referenced by review-dispatch steps in
spec_kitandcreatecommand YAML workflows. - Complements, but does not replace, sibling ownership: the surface skills (
code-webflow/code-opencode) apply fixes and own the implement → debug → verify workflow doctrine, andcode-qualityowns author-side gates.
8. RELATED RESOURCES
Start with references/quick-reference.md, then load task-specific doctrine, assets, or scripts.
Manual Testing Playbook
Manual testing scenarios for the code-review mode (of the sk-code family) live in manual-testing-playbook/manual-testing-playbook.md (root index) plus per-feature sub-files under manual-testing-playbook//.md (both the category folder and the scenario file use bare descriptive slugs, no numeric prefix). Run scenarios via bash .opencode/skills/sk-doc/scripts/validate_document.py manual-testing-playbook/manual-testing-playbook.md for structural validation; execute scenarios in opencode/Claude/OpenCode sessions for behavioral verification.
9. PR-STATE EFFICIENCY GATES
9.1 M-1: PR-State Content-Hash Dedup
Prevents redundant re-reviews when a PR has not changed since the last review.
Signature computation:
diff_content_hash = sha256(git diff ...HEAD)
signature = sha256(commit_subject + "\u001f" + diff_content_hash)
Where commit_subject is the first line of git log ...HEAD --format=%s (latest commit subject).
Cache storage:
- Path:
.opencode/.code-review-cache/.jsonl - `
is computed assha256(git remote get-url origin).slice(0, 12)` - Each line is a JSON object:
{"signature": "", "timestamp": "", "prev_sha": ""} - Retention: keep last 100 entries per repo-ref, prune older entries on write
Skip behavior: When the current signature matches a prior cache entry, the review emits:
Review status: COMMENTED (no changes since last review at )
No full review analysis runs. Automation may treat COMMENTED as a pass (no new findings).
Cache write: After each full review completes, write the current signature + timestamp + HEAD SHA to the cache file.
9.2 M-2: Opt-In Minimum Evidence Gate
Skips full review for trivially small diffs to save compute, with a conservative taxonomy that never skips high-risk changes.
Enable gate:
export SK_CODE_REVIEW_MIN_CHANGED_LINES=50 # >0 enables; default 0 = disabled
Changed-line counting command:
git diff --numstat ...HEAD | awk '{added+=$1; removed+=$2} END {print added+removed}'
Conservative skip taxonomy — NEVER skip when diff touches:
| Risk Class | Path/File Patterns | Rationale | |---|---|---| | Security / Authentication / Authorization | auth*, *-auth-*, *permission*, *credential*, *token*, *secret*, *oauth*, *sso*, *login*, *session* | Compromised auth defeats everything | | Config files | *.config.*, *config*.json, *config*.yaml, *config*.toml, *.env*, *.ini, *.cfg | One-line config change can break production | | Persistence | *.sql, *migration*, *schema*, *db*.ts, *repository*, paths under /db/ or /migrations/ | Schema changes risk data loss | | Dependency manifests | package.json, package-lock.json, Cargo.toml, Cargo.lock, pyproject.toml, poetry.lock, requirements.txt, *.lock, Gemfile, Gemfile.lock | Transitive dependency changes are high-risk | | Sandboxing / Subprocess | *sandbox*, *subprocess*, *exec*, *spawn*, *eval* | Arbitrary code execution boundaries | | Public-facing responses | *.handler.ts, *-api*, *-route*, *-controller*, paths under /handlers/, /routes/, /api/ | User-visible behavior changes |
Skip behavior: When SK_CODE_REVIEW_MIN_CHANGED_LINES > 0, total changed lines config > default) — exactly like the §9.2 SK_CODE_REVIEW_MIN_CHANGED_LINES gate. Both are skill guidance the reviewer reads and applies in-loop, not a separate compiled dispatcher; this alias only NAMES and PERSISTS an already-existing routing behavior, adds no new tier, and relaxes no floor:
full(default / unset): the normal ALWAYS + CONDITIONAL + ON_DEMAND routing.ultra: bias intent selection toward the existingON_DEMANDreference set (the deep-dive tier) for the session, so a reviewer does not have to repeat "comprehensive / full review" each time.lite: maps to the existing M-2 conservative skip (§9.2) — it NEVER lowers the ALWAYS tier, the baseline security/correctness minimums, or the P0/P1/P2 contract. It cannot skip a review on a sensitive path (auth/config/persistence/deps/sandbox/public-response), exactly as M-2 already enforces.
The depth alias is advisory routing only; it must never be read as permission to relax a floor.
Related skills: sk-doc for skill authoring and packaging standards, sk-code for surface-aware standards, and system-spec-kit for packet-governed review workflows.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: MichelKerkmeester
- Source: MichelKerkmeester/skilled-harness_spec-driven-agent-loops
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.