AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Sk Code Review

skill-michelkerkmeester-skilled-harness-spec-driven-agent-loops-sk-code-review · by MichelKerkmeester

Stack-agnostic code-review for sk-code: findings-first severity, security/correctness minimums, and surface evidence.

No reviews yet
0 installs
19 views
0.0% view→install

Install

$ agentstack add skill-michelkerkmeester-skilled-harness-spec-driven-agent-loops-sk-code-review

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access Used
  • Shell / process execution Used
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-michelkerkmeester-skilled-harness-spec-driven-agent-loops-sk-code-review)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
24d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Sk Code Review? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

code-review Mode - Stack-Agnostic Findings-First Review

Universal findings-first review baseline paired with sk-code surface standards evidence for the detected code surface.

1. WHEN TO USE

Activation Triggers

Use the code-review mode (of the sk-code family) when:

  • A user asks for code review, PR review, quality gate, or merge readiness.
  • A workflow dispatches @review for pre-commit or gate validation.
  • A user requests security/correctness risk analysis before merge.
  • A user wants severity-ranked findings with file:line evidence.

Keyword Triggers

review, code review, pr review, audit, security review, quality gate, request changes, findings, blocking issues, merge readiness

Use Cases

  1. Review-only pass: findings-first output with no code edits.
  2. Gate validation: score + pass/fail recommendation for orchestrated workflows.
  3. Focused risk pass: security, concurrency, correctness, or removal-focused review.

When NOT to Use

  • Feature implementation without review intent; use the surface skill (code-webflow / code-opencode).
  • Pure documentation editing where code behavior is not being assessed.
  • Git-only workflow tasks (branching, rebasing, commit hygiene) without code-quality evaluation intent.
  • Applying review fixes after findings are accepted; use the surface skill (code-webflow / code-opencode).
  • Author-side quality gates before review; use code-quality.
  • Root-cause debugging; use the surface's workflow-debug.md doctrine.
  • Verification evidence collection; use the surface's workflow-verify.md doctrine.

2. SMART ROUTING

Primary Detection Signal

Review behavior follows a baseline+surface-evidence model:

  • Baseline (always): the code-review mode (of the sk-code family) findings-first doctrine.
  • Surface standards evidence (when available): sk-code detected surface resources.
  • Unknown surfaces: review against baseline security/correctness only and disclose uncertainty.

Phase Detection

TASK CONTEXT
    |
    +- STEP 0: Load the `code-review` mode baseline + `sk-code` surface evidence. The dispatcher / agent assembling the code-review prompt MUST prepend `CODE-REVIEW\n\n` as the first two lines of the rendered prompt before the reviewer LLM sees it. Reference resources stay unchanged.
    +- STEP 1: Score intents (top-2 when ambiguity delta  str:
    return " ".join([
        str(getattr(task, "text", "")),
        str(getattr(task, "query", "")),
        str(getattr(task, "description", "")),
        " ".join(getattr(task, "keywords", []) or []),
    ]).lower()

def _guard_in_skill(relative_path: str) -> str:
    resolved = (SKILL_ROOT / relative_path).resolve()
    resolved.relative_to(SKILL_ROOT)
    if resolved.suffix.lower() != ".md":
        raise ValueError(f"Only markdown resources are routable: {relative_path}")
    return resolved.relative_to(SKILL_ROOT).as_posix()

def discover_markdown_resources() -> set[str]:
    docs = []
    for base in RESOURCE_BASES:
        if base.exists():
            docs.extend(path for path in base.rglob("*.md") if path.is_file())
    return {doc.relative_to(SKILL_ROOT).as_posix() for doc in docs}

def keyword_present(keyword: str, text: str) -> bool:
    """Boundary-aware match: bare substrings misroute ('pr' in 'improve prompt')."""
    return re.search(rf"(? dict[str, float]:
    text = _task_text(task)
    scores = {intent: 0.0 for intent in INTENT_SIGNALS}
    for intent, cfg in INTENT_SIGNALS.items():
        for keyword in cfg["keywords"]:
            if keyword_present(keyword, text):
                scores[intent] += cfg["weight"]
    return scores

def select_intents(scores: dict[str, float], ambiguity_delta: float = 1.0, max_intents: int = 2) -> list[str]:
    ranked = sorted(scores.items(), key=lambda item: item[1], reverse=True)
    if not ranked or ranked[0][1]  1 and ranked[1][1] > 0 and (ranked[0][1] - ranked[1][1])  str:
    text = _task_text(task)
    files = " ".join((workspace_files or []) + (changed_files or [])).lower()

    if ".opencode/" in files or keyword_present("jsonc", text) or keyword_present("mcp", text):
        return "sk-code:code-opencode"
    if any(keyword_present(term, text) for term in ["frontend", "web", "css", "dom", "browser"]) or any(
        marker in files for marker in ["next.config", "vite.config", "package.json", "src/"]
    ):
        return "sk-code:code-webflow"
    return "sk-code:unknown"

def route_review_resources(task, workspace_files=None, changed_files=None):
    inventory = discover_markdown_resources()
    text = _task_text(task)
    scores = score_intents(task)
    intents = select_intents(scores, ambiguity_delta=1.0)

    loaded = []
    seen = set()

    def load_if_available(relative_path: str) -> None:
        guarded = _guard_in_skill(relative_path)
        if guarded in inventory and guarded not in seen:
            load(guarded)
            loaded.append(guarded)
            seen.add(guarded)

    for relative_path in DEFAULT_RESOURCES:
        load_if_available(relative_path)

    if sum(scores.values()) =4` or any other numeric threshold as a blocker.

#### Instance-Only Opt-Out

A finding may use the narrow fix path only when all are true:

- It is not P0/P1 security, path, auth/authz, sandboxing, env precedence, schema, persistence, or public-response behavior.
- `rg` proves no same-class producer or consumer.
- Verification is local and cheap: one focused test, one doc row, or one static audit command.
- The fix response includes the exact command evidence for the opt-out.

Otherwise, run the full fix completeness checklist.

### Phase 4: Output and Next Action

Required output contract:

```markdown
## Code Review Summary

**Files reviewed**: X files, Y lines changed
**Overall assessment**: [APPROVE / REQUEST_CHANGES / COMMENT]
**Baseline used**: [sk-code (`code-review`)]
**Surface evidence used**: [sk-code:code-webflow | sk-code:code-opencode | sk-code:unknown]

## Findings

### P0 - Critical
1. [path:line] Title
   - Risk
   - User impact
   - Finding class: [instance-only | class-of-bug | cross-consumer | algorithmic | matrix/evidence | test-isolation]
   - Scope proof: [grep/test evidence proving class coverage or instance-only status]
   - affectedSurfaceHints: [optional string array of short producer/consumer surface names; recommended for actionable findings, required for cross-consumer findings]
   - riskScore: [optional advisory number only; never gating]
   - Recommended fix

### P1 - High
...

## Removal/Iteration Plan

## Next Steps

After reporting findings, request explicit next action before any implementation follow-up.

Final-line exact-string contract (MANDATORY)

Every review MUST end with exactly one of the following plain-text lines as the absolute final line of the output (no trailing whitespace, no variation):

Review status: APPROVED
Review status: REQUESTED_CHANGES
Review status: COMMENTED

Example output bottom:

...
## Next Steps
1. Fix the null-deref at src/foo.ts:42
2. Add input validation for the `/api/bar` endpoint

Review status: REQUESTED_CHANGES

Downstream automation parses this final line via exact string match — do not vary the format, add trailing punctuation, or wrap in Markdown formatting. The sole exception is the documented M-1 / M-2 skip output (§9): those lines begin with the exact Review status: COMMENTED and append a parenthetical reason, so a leading-verdict (grep / startsWith) parse still yields COMMENTED. A normal review must still end with one of the three exact lines above.


4. RULES

✅ ALWAYS

  • Keep findings first; summaries follow findings.
  • Enforce baseline security/correctness minimums regardless of surface.
  • Include file:line evidence for actionable findings.
  • State assumptions when evidence is incomplete.
  • Identify sk-code surface evidence used for standards alignment.

⛔ NEVER

  • Override surface-specific conventions with generic baseline style preferences.
  • Approve code with unaddressed P0 security/correctness defects.
  • Produce vague findings without concrete evidence.
  • Mix unrelated cleanup into targeted fix recommendations.
  • Do not implement fixes during review. Report findings only; implementation is a separate follow-up step.

⚠️ ESCALATE IF

  • Surface detection is ambiguous and affects standards or verification commands.
  • Baseline and surface guidance conflict in a non-deterministic way.
  • Large diff size prevents reliable severity assignment without narrowed scope.
  • Requested remediation exceeds review scope and becomes architecture redesign.

5. REFERENCES

Core References

  • [review-core.md](./references/review-core.md) - Shared review doctrine: severity model, evidence rules, precedence, and finding schema.
  • [review-ux-single-pass.md](./references/review-ux-single-pass.md) - Interactive single-pass review flow, presentation modes, and PR/pre-commit behavior.
  • [quick-reference.md](./references/quick-reference.md) - Lightweight index for routing between shared doctrine and single-pass UX guidance.
  • [pr-state-dedup.md](./references/pr-state-dedup.md) - Content-hash deduplication for unchanged pull-request reviews.
  • [security-checklist.md](./assets/security-checklist.md) - Mandatory security and reliability checks.
  • [code-quality-checklist.md](./assets/code-quality-checklist.md) - Correctness, performance, KISS, and DRY checks.
  • [solid-checklist.md](./assets/solid-checklist.md) - SOLID (SRP/OCP/LSP/ISP/DIP) and architecture assessment prompts.
  • [removal-plan.md](./assets/removal-plan.md) - Safe-now vs deferred removal planning template.
  • [test-quality-checklist.md](./assets/test-quality-checklist.md) - Test quality, coverage, and anti-pattern detection.

Reference Loading Notes

  • Load only the references needed for the selected intents.
  • Keep Section 2 (SMART ROUTING) as the authoritative routing source.

6. SUCCESS CRITERIA

  • Review output is findings-first and severity-ordered.
  • code-review mode baseline + sk-code surface evidence contract is explicit in report context.
  • Security/correctness minimums are always covered.
  • Recommended fixes are actionable and scope-proportional.

7. INTEGRATION POINTS

  • Primary review baseline for @review agents in .opencode/agents/review.md.
  • Referenced by review-dispatch steps in spec_kit and create command YAML workflows.
  • Complements, but does not replace, sibling ownership: the surface skills (code-webflow / code-opencode) apply fixes and own the implement → debug → verify workflow doctrine, and code-quality owns author-side gates.

8. RELATED RESOURCES

Start with references/quick-reference.md, then load task-specific doctrine, assets, or scripts.

Manual Testing Playbook

Manual testing scenarios for the code-review mode (of the sk-code family) live in manual-testing-playbook/manual-testing-playbook.md (root index) plus per-feature sub-files under manual-testing-playbook//.md (both the category folder and the scenario file use bare descriptive slugs, no numeric prefix). Run scenarios via bash .opencode/skills/sk-doc/scripts/validate_document.py manual-testing-playbook/manual-testing-playbook.md for structural validation; execute scenarios in opencode/Claude/OpenCode sessions for behavioral verification.


9. PR-STATE EFFICIENCY GATES

9.1 M-1: PR-State Content-Hash Dedup

Prevents redundant re-reviews when a PR has not changed since the last review.

Signature computation:

diff_content_hash = sha256(git diff ...HEAD)
signature         = sha256(commit_subject + "\u001f" + diff_content_hash)

Where commit_subject is the first line of git log ...HEAD --format=%s (latest commit subject).

Cache storage:

  • Path: .opencode/.code-review-cache/.jsonl
  • ` is computed as sha256(git remote get-url origin).slice(0, 12)`
  • Each line is a JSON object: {"signature": "", "timestamp": "", "prev_sha": ""}
  • Retention: keep last 100 entries per repo-ref, prune older entries on write

Skip behavior: When the current signature matches a prior cache entry, the review emits:

Review status: COMMENTED (no changes since last review at )

No full review analysis runs. Automation may treat COMMENTED as a pass (no new findings).

Cache write: After each full review completes, write the current signature + timestamp + HEAD SHA to the cache file.

9.2 M-2: Opt-In Minimum Evidence Gate

Skips full review for trivially small diffs to save compute, with a conservative taxonomy that never skips high-risk changes.

Enable gate:

export SK_CODE_REVIEW_MIN_CHANGED_LINES=50  # >0 enables; default 0 = disabled

Changed-line counting command:

git diff --numstat ...HEAD | awk '{added+=$1; removed+=$2} END {print added+removed}'

Conservative skip taxonomy — NEVER skip when diff touches:

| Risk Class | Path/File Patterns | Rationale | |---|---|---| | Security / Authentication / Authorization | auth*, *-auth-*, *permission*, *credential*, *token*, *secret*, *oauth*, *sso*, *login*, *session* | Compromised auth defeats everything | | Config files | *.config.*, *config*.json, *config*.yaml, *config*.toml, *.env*, *.ini, *.cfg | One-line config change can break production | | Persistence | *.sql, *migration*, *schema*, *db*.ts, *repository*, paths under /db/ or /migrations/ | Schema changes risk data loss | | Dependency manifests | package.json, package-lock.json, Cargo.toml, Cargo.lock, pyproject.toml, poetry.lock, requirements.txt, *.lock, Gemfile, Gemfile.lock | Transitive dependency changes are high-risk | | Sandboxing / Subprocess | *sandbox*, *subprocess*, *exec*, *spawn*, *eval* | Arbitrary code execution boundaries | | Public-facing responses | *.handler.ts, *-api*, *-route*, *-controller*, paths under /handlers/, /routes/, /api/ | User-visible behavior changes |

Skip behavior: When SK_CODE_REVIEW_MIN_CHANGED_LINES > 0, total changed lines config > default) — exactly like the §9.2 SK_CODE_REVIEW_MIN_CHANGED_LINES gate. Both are skill guidance the reviewer reads and applies in-loop, not a separate compiled dispatcher; this alias only NAMES and PERSISTS an already-existing routing behavior, adds no new tier, and relaxes no floor:

  • full (default / unset): the normal ALWAYS + CONDITIONAL + ON_DEMAND routing.
  • ultra: bias intent selection toward the existing ON_DEMAND reference set (the deep-dive tier) for the session, so a reviewer does not have to repeat "comprehensive / full review" each time.
  • lite: maps to the existing M-2 conservative skip (§9.2) — it NEVER lowers the ALWAYS tier, the baseline security/correctness minimums, or the P0/P1/P2 contract. It cannot skip a review on a sensitive path (auth/config/persistence/deps/sandbox/public-response), exactly as M-2 already enforces.

The depth alias is advisory routing only; it must never be read as permission to relax a floor.


Related skills: sk-doc for skill authoring and packaging standards, sk-code for surface-aware standards, and system-spec-kit for packet-governed review workflows.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.