- Зачем:
- Round 1 adversarial self-review нашёл 6 находок (3 high, 3 medium).
Артефакт первой итерации содержал противоречия между SKILL и README,
неточности про upstream-механики (Codex continuation, Agent Teams
gating SendMessage), и оставлял undefined load-bearing branch
"Operator available?". Без фиксов скилл вводил пользователей без
Agent Teams в ступор при Round 2.
- Что:
- SKILL Step 2: уточнено, что SendMessage gated на Agent Teams; Task
spawn без Agent Teams one-shot; Codex поддерживает intra-session
continuation.
- SKILL Step 4: "REQUIRED SUB-SKILL ... if available" заменено на
"Use ... when available" с явным fallback на inlined принципы.
- SKILL Step 6: добавлена continuation-support table по хостам,
восстановлена секция "Detecting an operator" с heuristic
(interactive если direct human message + user-facing channel).
- README: переписан Agent Teams блок под operator-gated модель; Codex
section — корректное описание intra-session continuation через /agent.
- EXPERIMENT.md: добавлен banner о предшествии редизайну и устаревших
ссылках на adversarial-reviewer.md.
- Проверка:
- Прогнать /claude-team-review в новой сессии с Agent Teams enabled
(env var добавлен в ~/.claude/settings.json) — Round 2 для проверки
continuation через SendMessage и operator-gating паттерна.
19 KiB
name, description, user_invocable
| name | description | user_invocable |
|---|---|---|
| claude-team-review | Use when user says /claude-team-review, requests adversarial review of a plan or code change, wants peer review without external API dependencies, or needs to verify implementation against a plan before merging. | true |
Claude Team Review
Adversarial review of plans and code through a peer-reviewer subagent. The reviewer reads the project, runs tests and docs lookups, and delivers findings. The lead evaluates those findings (not blindly applies them), fixes what holds up, pushes back with reasoning on what doesn't, and asks for re-review. Up to 5 rounds.
Works on any host that supports subagents — Claude Code (via Task tool or Agent Teams) and Codex (native subagents) are both fine. The skill is platform-agnostic; the platform decides how to spawn.
When to invoke
/claude-team-review— auto-detect what to review/claude-team-review plan— force plan review/claude-team-review code— force code review/claude-team-review <file-path>— review a specific file (argument contains/or.)/claude-team-review xhigh— use max reasoning effort for the reviewer
Step 1: Determine review mode
Check in priority order:
1. Explicit argument (plan, code, file path) → use it. For plan,
skip all git checks and proceed to Step 2.
2. Plan Mode active (Claude Code) — if context contains the system
message "Plan mode is active" → mode = plan, skip git.
3. Auto-detect (no explicit argument, not in Plan Mode):
- Check for code changes (any non-empty output means changes exist):
git diff --name-only— unstagedgit diff --cached --name-only— staged
- Check if a plan exists in the current conversation context.
| Code changes? | Plan in context? | Mode |
|---|---|---|
| No | Yes | plan |
| Yes | Yes | code-vs-plan |
| Yes | No | code |
| No | No | Ask the user |
Step 2: Spawn the reviewer
Spawn a reviewer subagent using your host's standard mechanism, and
pass it the briefing assembled from reviewer-prompt.md.
- Claude Code: Task tool with
general-purposetype, OR — if Agent Teams is enabled — spawn a teammate. Note: only Agent Teams supports continuation between rounds (viaSendMessage, which is gated onCLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1). Without Agent Teams, the Task tool spawn is one-shot — Round 2+ has to go through the operator- gated fresh-spawn path in Step 6. - Codex: native subagent spawn (host orchestrates spawn/wait/consolidate).
Codex supports intra-session continuation — the lead can send follow-up
instructions to a running subagent thread, and
/agentswitches between active threads. - Other hosts: equivalent subagent mechanism. The skill works as long as the host can spawn a subagent and (optionally) route follow-up messages to it.
The reviewer is a full agent session — it explores the repo, runs git commands, and reads files on its own. Do not pre-collect diffs or file lists for it. Pass mode-specific context only:
- For
plan— path to plan file, or inline plan text - For
code— instruct it to use git status / git diff - For
code-vs-plan— pass the plan and instruct git lookup for changes
Fill the placeholders in reviewer-prompt.md and pass the result as the
subagent's prompt. The template includes operating stance, finding bar,
scope exclusions, and required output format.
Effort override: if the user passed xhigh, route the spawn through
a maximum-reasoning configuration if the host supports it.
Tool restrictions: the briefing explicitly forbids file modification.
If the host supports enforced read-only sandboxing (Claude Code
disallowedTools: Write, Edit, Codex sandbox_mode = "read-only"), apply
it on top — it's defense in depth, not the primary control.
Step 3: Show findings
When the reviewer responds, show the user the response verbatim — do not rephrase, summarize, or reorder:
## Team Review — Round N (mode: <plan|code|code-vs-plan>)
[Reviewer's response — verbatim]
Then check the verdict line:
| Verdict | Next step |
|---|---|
VERDICT: APPROVED |
Step 7 (final result) |
VERDICT: REVISE |
Step 4 (evaluate findings) |
| Unclear / missing | Send a message back asking for a clear verdict line |
| Round 5 reached | Step 7 with the "Max rounds reached" terminal state |
Step 4: Evaluate findings (do NOT apply yet)
External feedback = suggestions to evaluate, not orders to follow. This step exists because the reviewer may be technically wrong — and applying its findings blindly causes real damage (large structural edits based on cited issues that turn out to be feature requests, not bugs).
Use superpowers:receiving-code-review when available — it has the
full discipline for this step. The key principles are inlined below as
the always-available fallback (so the matrix step works on any host,
even where Superpowers is not installed).
Build the evaluation matrix
For each finding, fill out:
| # | Severity | Verified? | Type | Action |
|---|---|---|---|---|
| 1 | high | ✓ Context7 confirms behavior | arch | accept |
| 2 | critical | ✗ cited issue is feature request, not bug | tool-mechanic | reject with reasoning |
| 3 | medium | ✓ quick repro confirms | tool-mechanic | accept |
Action options are equal — accept, reject with reasoning, and
re-scope (apply a narrower fix). Reject and re-scope are not
exceptions; they are first-class outcomes.
Verification methods by finding type
| Finding type | What constitutes verification |
|---|---|
| Architectural / design | Reasoning + codebase grep, plus pattern check against existing code |
| Tool-mechanic (DSL syntax, config parser, API contract, library behavior) | Empirical test on the real system — reasoning is not enough |
| Style / convention | Match against actual codebase conventions |
| Security | Reasoning + concrete threat model |
Tool-mechanic findings are the most dangerous to accept on reasoning alone. Mental models of obscure tools are often wrong. If the reviewer cites an upstream issue or doc — open it. Do not trust the citation by number; issues get reclassified, closed, or turn out to describe a different case.
Receiving feedback — key principles
Inlined from superpowers:receiving-code-review for portability:
- Read all findings end-to-end before reacting
- Restate the technical claim in your own words (or ask)
- Verify against codebase / docs / a quick run before accepting
- Push back when wrong — with technical reasoning, not deference
- No performative agreement ("you're absolutely right" is a violation)
- Skip thanks. Just state the fix or the reasoning.
Show the matrix to the user
If an operator is present, show the matrix before applying. In headless or autonomous runs, proceed but be ready to explain each decision in the final summary.
Step 5: Apply or push back
For findings marked accept — fix them:
- Plan review: update the plan to address the finding
- Code review: edit files, run tests if applicable
- Code-vs-plan: either update the plan or the code, depending on which is wrong
Verify your own technical claims before publishing them. This is the reverse direction of receiving-code-review — not "don't accept someone else's unverified claim", but "don't publish your own".
When a fix or a reply to the reviewer makes a claim about tool mechanics (DSL syntax, config parser behavior, API contract, library behavior), verify it empirically:
- If a quick test is possible — run it (
docker run …, a real database container, a small repro script, whatever maps to the claim) - If a quick test is not possible — frame the claim as a hypothesis ("seems to", "needs verification") rather than as fact
Skip a fix that contradicts the user's explicit requirements — note this in the response to the reviewer.
Show the user a brief account:
### Round N fixes
- Applied: [#1 — what changed, 1 line]
- Re-scoped: [#3 — what changed, why narrower]
- Rejected: [#2 — short reason; full reasoning goes to the reviewer]
Step 6: Request re-review (Rounds 2–5)
Compose a structured response and send it to the reviewer.
Response format
I've evaluated the findings. Here's the state:
## Applied
- [#N]: [what was changed and why, 1–2 lines]
- ...
## Re-scoped
- [#N]: [scope/tone adjustment, with reasoning]
- ...
## Rejected with reasoning
- [#N]: [technical reason for not applying — not just "I disagree"]
- ...
## Specific asks for re-review
1. Are my rejections technically valid?
2. Any new issues introduced by the applied fixes?
3. [Any specific question about a high-risk fix]
[For plan mode with inline plans — append the full revised plan text:]
## Current revised plan
[Full text of the revised plan]
The three-section format gives the reviewer a chance to contest the rejections. A re-review that says "your rejection of #2 is valid; here's why" is just as useful as one that fixes new issues — both keep the loop honest.
Continuation vs fresh subagent
The cheap path is continuation — the same reviewer keeps context between rounds. Fresh-spawn is expensive: the new subagent must re-read the project from scratch. Because of that cost, fresh-spawn is always operator-gated, regardless of platform.
digraph re_review {
"Continuation supported by host\nand previous reviewer alive?" [shape=diamond];
"Continue same reviewer" [shape=box style=filled fillcolor=lightgreen];
"Operator available?" [shape=diamond];
"Ask operator:\nspawn fresh (full re-read),\nor conclude unverified?" [shape=box];
"Spawn fresh subagent\nwith PREVIOUS_FINDINGS block" [shape=box];
"Step 7 — NOT VERIFIED terminal state" [shape=box style=filled fillcolor=lightyellow];
"Continuation supported by host\nand previous reviewer alive?" -> "Continue same reviewer" [label="yes"];
"Continuation supported by host\nand previous reviewer alive?" -> "Operator available?" [label="no"];
"Operator available?" -> "Ask operator:\nspawn fresh (full re-read),\nor conclude unverified?" [label="yes"];
"Operator available?" -> "Step 7 — NOT VERIFIED terminal state" [label="no — headless"];
"Ask operator:\nspawn fresh (full re-read),\nor conclude unverified?" -> "Spawn fresh subagent\nwith PREVIOUS_FINDINGS block" [label="re-spawn"];
"Ask operator:\nspawn fresh (full re-read),\nor conclude unverified?" -> "Step 7 — NOT VERIFIED terminal state" [label="conclude"];
}
Continuation support, by host (verify against current docs before relying on this):
| Host | Continuation | Mechanism |
|---|---|---|
Claude Code + Agent Teams (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1) |
yes | SendMessage to the teammate inbox |
| Claude Code without Agent Teams | no — Task spawn is one-shot | — |
| Codex | yes — intra-session | follow-up message to a running thread; /agent to switch |
Two routes lead to the "ask operator" step:
- The host has no continuation mechanism for our use case (Claude Code without Agent Teams) — every round after Round 1 lands here
- The host has continuation but the previous reviewer is no longer reachable (process ended, thread closed, inbox dead)
Detecting a dead reviewer (continuation-supporting hosts). When the
send call returns only a routing acknowledgment (e.g. {"success": true, "message": "Message sent to reviewer's inbox"}) without review content —
the reviewer process has ended; the message went to a dead inbox. Do not
wait. Treat the reviewer as unreachable, fall through to the operator
question.
Detecting an operator. Treat the run as interactive (operator available) only if both conditions hold:
- The current conversation has a direct human message — not just an automated trigger, hook, scheduled run, or CI invocation
- The host exposes a user-facing channel to surface the question
When in doubt, default to presenting the choice — pausing once is cheaper than concluding unverified by mistake.
Asking the operator. Present the choice plainly:
The reviewer cannot be continued from the previous round
(reason: <no continuation on this host | previous reviewer ended>).
Fixes applied:
[List from Step 5]
Options:
(a) Spawn a new reviewer to verify fixes (expensive — full project re-read)
(b) Conclude the review — fixes applied, verification is on you
Fresh subagent re-review (if operator chose re-spawn). Fill the
{PREVIOUS_FINDINGS_BLOCK} placeholder in reviewer-prompt.md with:
- Verbatim previous findings
- The Applied / Re-scoped / Rejected-with-reasoning sections
- For plan mode with inline plans: the full revised plan text
The fresh reviewer has zero prior context — it must see the actual artifact, not a paraphrase.
When the reviewer responds — return to Step 3 with N+1.
Severity declination (soft signal)
Expect severity of findings to decline across rounds:
R1: 3 critical, 6 high, 5 medium (typical)
R2: 1 high, 1 medium, 3 low (good)
R3: 1 high (closing in)
R4: APPROVED (terminal)
If severity stays flat (e.g. high → high → high), something is structurally off — the lead may not understand the technology, the reviewer may be looping on the same misunderstanding, or the artifact has a deep problem that surface fixes can't reach. Pause and surface to the user. This is a soft signal, not a hard gate.
Step 7: Final result
Approved:
## Team Review — Summary (mode: <mode>)
**Status:** Approved after N round(s)
[Final review verbatim]
---
**Reviewed and approved by the reviewer. Awaiting your decision.**
Not re-verified (reviewer became unreachable mid-loop and the operator chose to conclude, or the run is headless):
## Team Review — Summary (mode: <mode>)
**Status:** NOT VERIFIED — fixes applied, reviewer did not re-verify
**Round N findings:**
[Verbatim findings from the last reviewer round]
**Applied / Re-scoped / Rejected:**
[The three sections from the last round]
---
**WARNING: This is NOT an approval. Fixes were applied but never verified
by the reviewer. Manual review of the fixes is required before merging.**
Maximum rounds reached:
## Team Review — Summary (mode: <mode>)
**Status:** Maximum reached (5 rounds) — not fully approved
**Remaining findings:**
[Unresolved issues]
---
**The reviewer still has findings. Please review them and decide how to proceed.**
Step 8: Cleanup
If the host provides a subagent / teammate cleanup mechanism, use it. Failures are non-blocking — subagents are reclaimed when the session ends.
Do NOT delete plan files that existed before the review.
Rules
- Lead actively evaluates and acts on findings — this is NOT message forwarding, and it is NOT blind acceptance
- Reviewer findings shown verbatim — do not rephrase or shorten
- Auto-detect mode from context; user arguments take priority
- The reviewer never writes files — enforced by the briefing rule, and by the host's sandbox if available
- The reviewer can run commands (tests, linters, git) and use MCP (Context7, web search) to verify findings
- Maximum 5 rounds to protect against infinite loops
- Show the user findings and the applied/rejected breakdown for each round
- Avoid creating auxiliary files (memory files, state files, logs, temporary markdown) — prefer working within the conversation context
- If a fix contradicts user requirements — skip it and explain why
- For re-review, prefer continuation (cheap). Fresh-spawn is the expensive path — always confirm with the operator before doing it. In headless runs without operator access, conclude unverified rather than auto-respawning.
- The ultimate goal is higher quality of plans, code, and other artifacts. Token economy is a means, not an end — never skip a verification step or cut a round short just to save tokens.
Red Flags — STOP and reconsider
When you catch yourself thinking any of these, you are about to violate the skill. Stop, re-read Steps 4–5.
| Thought | Reality |
|---|---|
| "Reviewer flagged this as critical — apply right away" | Build the matrix first. Verify before apply. |
| "The reviewer cites issue #N, I'll trust the number" | Open the issue. Citations age; issues get reclassified. |
"I know how <tool> works, no need to test the fix" |
Tool-mechanic claims need empirical verification, not reasoning. Run it or hedge it. |
| "I disagree with #2 but I'll just stay silent about it" | Reject with reasoning. The reviewer needs the chance to contest. |
| "You're absolutely right, applying now" | Performative agreement. Restate the requirement, then act. |
| "5 rounds is a lot, let me wrap this up at round 3" | Don't compress the loop to save tokens. Run it until terminal. |
| "Just paraphrase the findings to save space" | Verbatim. Always verbatim. Paraphrase loses signal. |
| "Severity hasn't dropped in 3 rounds, but I'll push through" | Pause. Surface to operator. Something is structurally off. |
| "Continuation isn't available — let me just spawn a fresh reviewer" | Ask the operator first. Fresh-spawn is expensive; the operator may prefer to conclude unverified. |
Comparison with adversarial-review
adversarial-review is a related but distinct skill (different repo):
Claude writes, Codex reviews — cross-model coverage.
| Aspect | adversarial-review (cross-model) | claude-team-review (same-host) |
|---|---|---|
| Reviewer model | External (GPT via Codex CLI from Claude) | Whatever the host runs |
| Cross-model blind spots | Yes — different model biases | No — same model family |
| Session persistence | Via codex exec resume |
Continuation if host supports it; otherwise operator-gated fresh-spawn |
| External dependencies | Codex CLI + OpenAI API key (from Claude) | None — uses host's native subagent mechanism |
| Reviewer capabilities | Read-only sandbox | Read + execute + MCP + web |
| Host compatibility | Designed for Claude Code as lead | Claude Code AND Codex as lead |
Use adversarial-review for cross-model diversity. Use claude-team-review
when you want zero external dependencies and a richer reviewer (tests, docs,
web), regardless of which host is the lead.