AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Code Review

skill-conorbronsdon-agent-skills-code-review · by conorbronsdon

Multi-agent deep review for code PRs in any repo. Use when asked to "deep review this PR," "multi-agent review," "review #N with subagents," "thorough review of this branch," or "is this PR ready to merge." Orchestrates Copilot + parallel subagents (adversarial, operational, reference-comparison) with scope-based escalation, stale-finding triage, and a hard iteration cap.

No reviews yet
0 installs
12 views
0.0% view→install

Install

$ agentstack add skill-conorbronsdon-agent-skills-code-review

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-conorbronsdon-agent-skills-code-review)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
8d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Code Review? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Code Review — Multi-Agent PR Review

User-invokable orchestrator for code-PR review across any repo: Copilot for line-level findings, parallel subagents for the architectural and operational misses line-level review can't see. (Validation evidence and origin story: [README](README.md).)

Fallbacks: no gh or not a GitHub repo → review the local diff (git diff ...HEAD) and skip the Copilot lane entirely; no subagent tool → run the review angles sequentially in the main thread. Detect what's available before assuming Copilot access.


Invocation: deliberately model-invocable — "review this PR" phrasing is the trigger. Pushes and issue-creation inside the flow are confirmation-gated.

Step 1: Assess PR Scope

The review target is $ARGUMENTS (a PR number, branch, or URL). If empty, infer it from the conversation or ask.

gh pr view N --json files,additions,deletions,title

Categorize:

| Tier | Examples | Treatment | |---|---|---| | Trivial | Typo, copy edit, single-line deps bump, README polish | Skip multi-agent. Copilot alone is enough. | | Standard | Component refactor, single-feature bug fix, contained logic change | Copilot + 1 subagent (pick the angle that matches the risk shape) | | High-stakes | Payment, auth, crypto/wallet, RPC, deployment configs, scale-relevant changes, new external SDK integrations | Copilot + 2 parallel subagents minimum (always include adversarial + operational) |

If trivial → stop reading this skill, just request Copilot review and ship.


Step 2: Spawn Subagents in Parallel

Three angles that empirically work. Pick by risk shape:

  • Adversarial — "Try to break it. What's the worst attack?" Finds security holes, abuse vectors, replay/race conditions, malformed-input handling.
  • Operational — "What fails in production at scale? Latency, cost, observability, timeouts, env config, retries, cold-start?" Finds the architectural P0s line-level review misses (platform execution limits are the canonical class).
  • Reference-comparison — "Does this match the upstream reference impl exactly? Where does it diverge and why?" Use for third-party SDK / spec integration (Phantom deep-link, Stripe, OAuth providers, RPC clients).

Spawn all chosen subagents in a single tool-call batch so they run in parallel. Prompt templates live in patterns/subagent-prompts.md.

Every subagent prompt must specify:

  1. Repo path (absolute) + branch + commit hash at HEAD
  2. Files to focus on (paste the changed-files list from Step 1)
  3. Confidence rating + critical findings first
  4. Cap report length at ~500 words

Step 3: Request Copilot Review (in parallel with subagents)

gh pr edit N --add-reviewer copilot-pull-request-reviewer

Don't wait for Copilot to start before spawning subagents. They run in parallel.


Step 4: Triage Findings

Three buckets:

Real findings — fix now. Push fix to the PR branch. Pushing changes external state: unless the user already asked you to fix or babysit this PR, confirm before the first push of the session.

Stale re-flags — Copilot will sometimes re-flag findings against unmoved diff lines after fixes ship at HEAD. Always read the file at HEAD before re-fixing. If the fix is already there, drop a one-line reply on the comment ("Fixed at ") and move on. Do not re-fix.

Deferred (filed) — tests, refactors, or work outside scope. File as a GitHub issue (confirm with the user before creating it), link the issue number in the merge commit message, move on. Don't expand the PR.


Step 5: Iteration Cap

Maximum 2 Copilot review rounds. Empirically, by round 3 the signal-to-noise ratio collapses (~50% stale re-flags). Ship after round 2. If subagents flag genuinely new issues after round 2, those are issue-tracker items, not PR-blockers.


Step 6: When NOT to Use This Skill

  • Already-merged PRs → use Claude Code's built-in /security-review if you suspect a problem.
  • Trivial-tier PRs (per Step 1) → just request Copilot, skip the multi-agent overhead.
  • PRs where a paid multi-agent service was already run → don't double-spend.
  • Skill / prompt / instruction-file repos (where the "code" is markdown) → multi-agent is overkill; a single review pass against the trigger-quality and instruction-quality dimensions is enough.

Empirical Evidence

Validation case (5 rounds on a high-stakes PR):

  • Copilot caught ~10 real line-level issues (null checks, error handling, type mismatches).
  • Subagents caught what Copilot missed across all 5 rounds:
  • P0: Vercel maxDuration not set — function would time out under realistic load. Architectural miss invisible at the line level.
  • Persistent-replay vector in the auth flow.
  • Transient poll tolerance (no retry budget on flaky upstream).
  • Env-configurable timeout (hardcoded value).
  • Structured logging (string-concat logs, unparseable in production).
  • bodyParser size limit (DoS surface).
  • Round 3+ pattern: Copilot re-flagged the same already-fixed lines. Confirmed the round-2 cap.

Subagents are not redundant with Copilot. They cover a different layer.


Cross-references

  • patterns/subagent-prompts.md — paste-ready prompt templates for the three angles.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.