Install
$ agentstack add skill-szarkans-multi-code-review ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Multi-model code review
!sh -c 'for p in "$CLAUDE_PLUGIN_ROOT/scripts" "$HOME/.claude/skills/multi/scripts" "./.claude/skills/multi/scripts"; do [ -x "$p/probe.sh" ] && { "$p/probe.sh"; echo "scripts-dir: $p"; exit 0; }; done; echo "probe: NOT FOUND — locate scripts/probe.sh in this plugin and run it yourself"'
Several models read the same code, and you decide what reaches the user. That is the whole idea: one model invents problems and walks past real ones, and you cannot tell which from a single report. Independent models disagree, and the disagreement is the signal.
You are the judge, not a dispatcher. The reviewers propose; they do not vote, and none of them gets the last word. When they conflict, open the code and decide.
This is normally reached at the end of a session — work is done, PR or push comes next. Act like it: the question is "is this ready", not "here are some observations".
$SCRIPTS below is whatever the probe printed as scripts-dir:.
The gate
From the probe lines above — they are already there, do not re-run it.
No Codex, or not logged in → stop. This is multi-model review; without a second model there is nothing here that a single-model review does not already do. Say so and point at codex login. Do not quietly deliver a one-model review wearing a three-model label.
Anything else missing is a note, not a stop: no OpenCode, no ponytail, not a git repo (fine — then the target is files, not a diff). Name what was missing in the report and carry on.
Decide what you are reviewing
Whatever the user named. A diff, a branch, two files, one function, a line range, a module nobody has touched in three years, the whole repository. There is no fixed vocabulary here and no menu — read what they wrote and work out what they mean, the way a colleague would.
When nothing was named, in this order:
- What this session was about. If you just wrote or changed something, that
is the target — you know which files, what the task was, where you were unsure, what you fixed blind. That is better than a diff, which in a dirty tree also holds debug leftovers and unrelated edits, and which misses the old code the change leans on.
- The branch, if it is ahead of main and the tree is clean.
- The uncommitted diff otherwise.
Ask only when the answer genuinely changes what gets reviewed and you cannot tell — dirty tree and they mentioned a PR, say. Never open with a questionnaire: this skill runs at the end of a session, sometimes inside an autonomous run, and three questions there are worse than a wrong guess you announced. Guess, say what you guessed, let them correct you.
Whatever you settle on, resolve it into concrete paths, or a concrete git range, before dispatching. Every reviewer must look at the same thing — otherwise "two of them agreed" means nothing, they just happened to read the same file. If the target is vague, resolve it and say what you resolved it to.
Say what you are about to do
Always, before launching anything. Short, then go — this is not a request for permission, and you do not wait for an answer:
Reviewing:
Running now: Codex · OpenCode · ponytail
If they wanted something else they will say so, and interrupting is cheaper than an interrogation.
Launch everything free, immediately
Codex, OpenCode and the ponytail lens cost nothing per run and take 30–70 seconds wall clock. There is never a reason to hold them back or make them conditional on a mode. Start the two external ones in the background, both at once — OpenCode spends about a minute just warming up — and do everything else while they run.
$SCRIPTS/collect-context.sh [--diff ] [--paths ""] > /tmp/multi-ctx.md
$SCRIPTS/review-codex.sh --target "" [--diff ] [--paths ""] \
--effort [--focus ""] \
--context /tmp/multi-ctx.md --out /tmp/multi-codex.txt
$SCRIPTS/review-opencode.sh --target "" [--diff ] [--paths ""] \
--model --fallback [--focus ""] \
--context /tmp/multi-ctx.md --out /tmp/multi-opencode.txt
--diff is what makes it a change review; leave it off and the reviewers read the actual code instead, with old code fully in scope. --paths narrows hard. --target is always required — it is the human sentence, and it is what keeps the reviewers pointed at the same thing.
collect-context.sh gathers the repo's CLAUDE.md/AGENTS.md and the .claude/rules/*.md matching the target, and every reviewer gets it. This is what separates this from three models guessing: an external reviewer that does not know the project's settled decisions spends its findings re-litigating them.
The ponytail lens — invoke the ponytail:ponytail-review skill on the same target whenever the probe found it. It hunts one thing, over-engineering, and that keeps the defect reviewers out of matters of taste entirely (see below). Its findings are a different kind of thing and never mix with defects: they get their own section and cannot corroborate or contradict a bug.
> If ponytail mode is active, its SubagentStart hook injects the YAGNI > ruleset into every sub-agent, including the ones hunting bugs. If findings > start reading like simplification advice, that is why; PONYTAIL_SUBAGENT_MATCHER > is the fix. Mention it once, move on.
Then decide how much Claude to spend
The external reviewers are free; your sub-agents are the user's money. So do not guess the depth up front — wait for the free results and decide on evidence. A two-line change both reviewers called clean does not need three sub-agents. A change where Codex reports two HIGH and OpenCode disagrees with one of them does.
Say the decision in one line when you make it: "going to normal — Codex found two HIGH, OpenCode contradicts one."
An explicitly named mode skips all of this. Obey it.
| | Claude sub-agents | when | |---|---|---| | lite | correctness only | small, low-risk, external reviewers agree and found little | | normal (default) | correctness · security · design | anything heading for a PR | | ultra | those three, plus execution, plus an adversarial second Codex pass (--adversarial), plus one verify per single-source finding | expensive to get wrong, or asked for |
Spawn them in parallel, in one message. Give each the target, the paths or range, the contents of /tmp/multi-ctx.md, and the user's own words if there were any.
Model: the argument if given, else MULTI_REVIEWER_MODEL from the probe, else the agent files' default (Sonnet). No mode raises it on its own — ultra buys depth through more angles and real verification, not a bigger model.
ultra is not a deeper code review — it reviews whether the task got done. Was there a plan and was it followed; is the thing actually finished or only finished-looking; what was silently skipped; what will detonate later. That needs the task context — the plan, the spec, the conversation. Hand execution what you actually know about the job. Without any of that, ultra degrades to normal plus an architecture angle, and you should say so rather than pretend.
Judge
Normalize everything to {file, line, severity, claim}. Codex reports - [P1] — :- with the explanation indented (P1/P2/P3 → high/medium/low) and absolute paths; make them repo-relative. OpenCode reports FILE:LINE | SEVERITY | reason. Sub-agents report FILE:LINE | SEVERITY | confidence NN with two lines under it.
Bucket by the underlying problem, not by wording — the same bug gets three different descriptions:
- Corroborated — two or more reviewers from different families (Claude /
Codex / OpenCode). Leads the report; independent agreement is the strongest evidence this pipeline produces.
- Single-source — one reviewer. Check each before the user sees it: open the
cited lines, confirm it is real and reachable. In ultra, spawn one verify per finding instead and take its verdict.
- Minor — a real defect that is simply small. Not verified — that costs more
than it is worth — and listed at the bottom, one line each.
- Dropped — only two things belong here: the code contradicts it, or (in a
diff review) it predates the change. "Too small" is never a reason.
Rules that make the report worth reading:
- Same scepticism for everyone. Codex being expensive does not make it
right; the cheap reviewer being cheap does not make it wrong. In measurement here the cheap one caught a real bug Codex missed.
- Disagreements are surfaced, not averaged. Read the code, decide, and put
the disagreement in the report — where good reviewers split is where the user should look.
- Nothing raised disappears silently. Every finding lands in a bucket, and a
dropped one carries its reason.
- Answer the user's own words first, if they gave any — even when the answer
is "no, that path is fine, here is why".
Report
Report only: no edits, no commits, no PR comments. This is the default shape, not a schema — drop empty sections, and match the surrounding conversation.
# 🔍 Multi-review — ·
Reviewers: Claude · Codex · OpenCode · ponytail
## ✅ Corroborated ()
1. **HIGH** `path/file.py:120` —
— Codex + correctness
## 🔸 Single-source, verified ()
- **MEDIUM** `path/file.py:88` — — [OpenCode] verified:
## ⚖️ Reviewers disagreed ()
- `path/file.py:44` — Codex calls it a race; correctness says the caller holds the lock.
## 🪒 Simplicity — ponytail ()
- `path/file.py:52-71` — delete: retry wrapper around an idempotent local call.
## 🔹 Minor ()
- `path/file.py:12` — [correctness] log line interpolates the wrong id; misleads during an incident.
## ⚪ Dropped ()
- [Codex] `path/x.py:12` — pre-existing, not introduced by this change
## Verdict
Then stop. Offer to fix the top findings or to re-run deeper — do not do either unprompted.
Loop mode (opt-in) — fix, re-review, repeat
Only when explicitly asked ("loop", "until it's clean"). It edits the working tree; say so before the first edit if they were not explicit.
Each round: re-run the reviewers → take what survived judging → fix what is new → go again. Stop when a round brings nothing new, when only Minor is left, or at the cap (default 3). Never fix a dropped finding — silencing a false positive is worse than the finding. Never commit. After each round list what changed as file:line one-liners. If the cap is hit with findings open, say so.
Edge cases
- Huge target — say up front it will be slow and shallow, and offer a
narrower one, rather than quietly reviewing four hundred files badly.
- Lockfiles, generated code, vendored trees — say so and skip the external
reviewers; there is nothing there for them.
- A reviewer dies or goes silent for 60s — one kill-and-restart, then treat
it as absent and name it in the report. A reviewer still writing is alive, however long it takes. Never block the whole review on one backend.
- OpenCode falls back to its free model — its output says so. Repeat it in
the reviewer line; the user is entitled to know which model actually ran.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: szarkans
- Source: szarkans/multi
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.