Files
adversarial-review/docs/superpowers/specs/2026-05-20-adversarial-review-refresh-design.md
ddadminandClaude Opus 4.7 1237e168c3 feat(skill): refresh adversarial-review под Codex 0.132
- Зачем:
  - актуализация под новый Codex CLI (0.132 убрал флаг -a,
    ввёл -c approval_policy / -c approvals_reviewer);
  - переход на безопасные операторские дефолты
    (sandbox=workspace-write для эмпирической верификации,
    явные overrides sandbox:* и approvals:*);
  - наведение порядка в принятии решений по раунду ревью:
    review_quality + evaluation matrix + structural operator gate
    вместо «оператор сам разбирается».
- Что:
  - SKILL.md: override-граммар (sandbox:*, approvals:*, model:*,
    low/medium/high/xhigh), детект OPERATOR_LANGUAGE, runtime hint,
    conditional dirty-file warning в Step 2 (после захвата REPO_ROOT),
    dual-layer mutation snapshots (git status + sha256sum input-файлов),
    operation-aware dispatch table, evaluation matrix + structural
    operator gate (batch-pause), structured resume body
    (Applied / Re-scoped / Rejected / Specific asks), final operator
    summary на OPERATOR_LANGUAGE.
  - references/runner.md: переведён на Sonnet runner, добавлен
    Step R2.5 bwrap preflight, Step R4.5 review_quality + bounded
    triage, расширена 11-полевая JSON-схема результата.
  - docs/DESIGN.md: §4.14–§4.21 с обоснованиями новых решений,
    §7.8 refresh-era smoke checks, §8 version log с двумя раундами
    dogfood-а этого refresh-а (включая R2-корректировку
    обоснования residual gap для уже-грязных tracked-файлов).
  - README.md: таблица дефолтов, Safety considerations (честно
    задокументирован residual gap), Linux sandbox prerequisites
    (bwrap + AppArmor user namespaces), Operator language,
    Final operator summary, troubleshooting.
  - examples/review-output.md: модель в сэмпле обновлена на gpt-5.5.
  - docs/superpowers/specs/2026-05-20-...: сохранена спека дизайна
    с уточнениями после dogfood.
- Проверка:
  - dogfood: 2 раунда /adversarial-review code на этом refresh-е,
    R2 верифицировал R2#1 (ordering) и R2#2 (Option B, honest
    residual gap);
  - smoke: codex --version ≥ 0.132.0, bwrap preflight зелёный
    на reference WSL2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 09:56:08 +03:00

32 KiB

Adversarial Review Refresh Design

Overview

This design updates the adversarial-review skill while preserving its core shape: Claude remains the lead, Codex provides an external cross-model adversarial review, and the Sonnet runner isolates mechanical Codex CLI work from the expensive lead context.

The refresh makes the skill more useful to the operator, more robust on modern Codex CLI setups, and less prone to blindly applying reviewer feedback. It does not turn the skill into a larger framework, add mandatory custom agent definitions, or move final decisions from the lead into the runner.

Goals

  • Use gpt-5.5 as the default Codex reviewer model while preserving explicit model and reasoning overrides.
  • Let Codex reviews use the operator's language while keeping parsed literals stable.
  • Improve reviewer capability by avoiding overly restrictive default sandboxing.
  • Avoid nested human approval prompts that the operator cannot see from the Claude session.
  • Keep runner-side work bounded and disposable so the lead context stays clean.
  • Add compact Sonnet triage to reduce repeated review rounds without letting Sonnet make final decisions.
  • Require the lead to evaluate findings before applying fixes.
  • Support operator sign-off for structural fixes without blocking explicitly autonomous runs.
  • Provide a final operator-facing summary of what changed across the review.
  • Document Linux sandbox prerequisites clearly enough for a human or installer agent to act on them.

Non-Goals

  • Do not require a custom Claude agent type or extra installation artifact.
  • Do not use danger-full-access as an automatic fallback.
  • Do not make the runner apply fixes, start extra review rounds, or decide final accept / reject / re-scope outcomes.
  • Do not move Codex stdout, stderr, or rollout contents into the main Claude context.
  • Do not add a mandatory detailed implementation plan for small instructional changes.
  • Do not mix documentation languages in repository files.

Default Reviewer Configuration

The default Codex invocation should be optimized for non-interactive review inside a Claude-runner child process:

CODEX_MODEL              = gpt-5.5
CODEX_REASONING          = high
CODEX_SANDBOX            = workspace-write
CODEX_APPROVAL_POLICY    = on-request
CODEX_APPROVALS_REVIEWER = auto_review

workspace-write is the default for all review modes (plan, code, code-vs-plan). The reasoning is load-bearing for this skill, so it is captured here rather than in a comment.

The reviewer needs to actually run things to verify findings: run tests, build the project, query the web or upstream APIs to confirm current behavior, and exercise project CLIs end-to-end. These are write-class operations — test runners produce output and cache files, build commands write artifacts, and most non-trivial CLI execution touches local state. A read-only sandbox blocks all of that.

Read-only is not a total verification blocker. It still allows file inspection, rg / grep searches, MCP-backed doc lookups, and --help-style CLI introspection that does not write to the workspace. But the highest-value findings come from the path read-only blocks: "the reviewer ran the tests and X failed", "the reviewer built the project and Y broke", "the reviewer queried the live API and the assumed signature does not exist." Forfeiting those to gain protection against .gitignored-file side effects is a poor trade.

This applies to plan reviews as much as to code. A plan reviewer should be able to run a test the plan relies on, build the project to confirm a structural claim, or hit an external API to validate an assumption — not just rg for cited filenames. A previous design pass made read-only the default for plan mode and was reverted for this reason.

The safety argument for forcing read-only is also weaker than it appears. The broader obra:superpowers skill family demonstrates that careful prompt-level discipline is sufficient to govern complex agent behavior in security-relevant contexts without sandbox-level restrictions. A concrete instance is superpowers:receiving-code-review: it constrains how a lead processes adversarial feedback — read every finding before reacting, restate the technical claim in own words, verify empirically (especially tool-mechanic claims) before accepting, push back with technical reasoning when wrong, no performative agreement — and that discipline is enforced entirely through prompt instructions, not through any tool-level sandboxing of the lead. The same pattern applies to the reviewer side of this skill: a strict instruction contract (see §Reviewer Behavior), backed by pre/post git status --porcelain and the /tmp sha256 snapshot for after-the-fact detection.

Read-only's specific failure modes are covered without a default sandbox change:

  • Tracked-file mutation is caught by git status --porcelain in §Workspace Mutation Detection.
  • /tmp review-input mutation is caught by the sha256sum snapshot in the same section.
  • Mutation of .gitignored state inside REPO_ROOT is bounded (re-seedable dev state; tests read .env.local, they do not write it) and is mitigated through documentation plus the sandbox:read-only override for operators who knowingly accept the trade-off.

Operators with a concrete reason to opt into read-only (sensitive local state, untrusted reviewer prompt source, or any other case where sandbox-level write protection outweighs verification capability) can do so per-run via sandbox:read-only. The override stays available exactly because the trade-off is real — for those operators. Defaulting the rest of the user base to that mode would gut the skill's main value.

The intended initial codex exec shape:

codex exec --json \
  -m gpt-5.5 \
  -c model_reasoning_effort=high \
  -s workspace-write \
  -c approval_policy='"on-request"' \
  -c approvals_reviewer='"auto_review"' \
  -C "<REPO_ROOT>" \
  -o /tmp/codex-review-<REVIEW_ID>.md \
  -

Codex CLI 0.132+ removed the top-level --ask-for-approval / -a flag from codex exec; approval policy is now expressed only through -c approval_policy=<value>. The spec uses -c form for both approval_policy and approvals_reviewer so the command shape stays self-consistent and future-proof against further short-flag churn. The implementation should verify the supported approval-control surface against the actual installed Codex CLI version during smoke testing.

-o writes only the assistant's final message. The reviewer prompt must deliver the full structured review (findings, severity tags, VERDICT) inside that final message; intermediate tool calls and reasoning will not appear in the review file.

workspace-write gives the reviewer enough capability to run local checks without forcing every useful command through a read-only wall. approval_policy="on-request" keeps boundary crossings explicit. approvals_reviewer="auto_review" avoids human approval prompts inside the nested Codex process, which the operator cannot reliably see or answer from the parent Claude session.

Auto-review is not a security boundary. It reviews approval requests; it does not inspect actions already permitted by the selected sandbox. The skill still needs strong reviewer instructions and workspace mutation checks.

Existing overrides remain:

  • model:<name>
  • low, medium, high, xhigh

New minimal overrides:

  • sandbox:read-only | workspace-write | danger-full-access | inherit
  • approvals:user | auto_review | never

Rules:

  • Default -s is workspace-write for every review mode. Read-only is a deliberate operator opt-in via sandbox:read-only, never an automatic per-mode default.
  • sandbox:inherit omits -s and relies on the user's Codex config. Because the effective sandbox is unknown until Codex actually launches, the runner skips bwrap preflight under inherit. If the inherited config selects a bwrap-backed mode on a host where bwrap is misconfigured, the failure surfaces as a §Runner Result Schema degraded_environmental result and is treated as terminal infrastructure failure on the initial dispatch per the dispatch table.
  • approvals:auto_review (the default) passes -c approval_policy='"on-request"' plus -c approvals_reviewer='"auto_review"'.
  • approvals:user passes -c approval_policy='"on-request"' only, omitting the approvals_reviewer override so Codex falls back to its default user reviewer. Allowed only by explicit override because nested human approvals can hang the run.
  • approvals:never passes -c approval_policy='"never"'; boundary crossings fail instead of asking.
  • sandbox:danger-full-access is explicit-only and should surface a warning.
  • Codex's untrusted approval policy is intentionally not exposed as an override; the skill needs predictable boundary semantics, not per-command trust prompts.
  • No silent fallback may change sandbox or approval semantics.

Resume commands should respect Codex CLI support for resume: sandbox and approval mode are properties of the initial session unless the current CLI explicitly supports changing them on resume. Do not pass unsupported -s flags or approval-related -c overrides to codex exec resume.

Runner Responsibilities

The runner remains a Sonnet subagent responsible for Codex CLI mechanics:

  • parse the input block;
  • write the prompt with the attempt-scoped session marker;
  • launch exactly one Codex operation (initial, resume, or fresh-exec);
  • own the one internal retry budget;
  • validate exit code, stderr, review file, and session id;
  • archive failed-resume diagnostics;
  • write the authoritative result JSON;
  • return the RUNNER_RESULT_AT: <path> line.

The runner gains bounded analysis:

  • run sandbox preflight when the selected mode requires a bwrap-backed sandbox;
  • detect obvious degraded reviews, such as sandbox or environment failures disguised as reviewer output;
  • extract finding count, maximum severity, and review quality;
  • produce compact triage metadata.

Triage rules:

  • Cover all critical and high findings.
  • Cover up to 10 medium findings.
  • If more medium findings remain, summarize the remainder and set an explicit truncation flag.
  • Use only cheap checks: file existence, section existence, simple rg, and small read-only snippets.
  • Do not run long test suites, web searches, or broad documentation lookups during triage.
  • Do not produce final accept, reject, or re-scope decisions.
  • If uncertain, mark that lead judgment is needed.

If triage fails but the Codex review is valid, the review should continue with a warning. If the review itself is degraded by infrastructure failure, the runner should not count it as a normal round.

Runner instructions must include negative examples:

  • Do not edit project files.
  • Do not apply fixes.
  • Do not run multiple Codex review rounds in one dispatch.
  • Do not delete /tmp/codex-* files.
  • Do not decide which findings the lead must accept.
  • If a command unexpectedly changes project files, stop and report it.

Runner Result Schema

The runner writes a single JSON file at RESULT_PATH. The schema below is the authoritative contract between runner and lead. New fields added by this refresh are marked (new); everything else preserves the pre-refresh shape.

{
  "result":              "success | timeout | launch_failure | infra_error | input_error",
  "verdict":             "APPROVED | REVISE | null",
  "review_file":         "<absolute path or null>",
  "codex_session_id":    "<uuid or null>",
  "attempt_id":          "<string>",
  "errors":              "<string or null>",
  "archived_stdout":     "<path or null>",
  "archived_stderr":     "<path or null>",
  "user_warning":        "<string or null>",
  "review_quality":      "valid | degraded_environmental | degraded_content | unknown",  // (new)
  "triage": {                                                                            // (new)
    "status":              "ok | skipped | failed",
    "finding_count":       "<int>",
    "max_severity":        "critical | high | medium | none",
    "covered_critical":    "<int>",
    "covered_high":        "<int>",
    "covered_medium":      "<int>",
    "truncated":           "<bool>",
    "needs_lead_judgment": "<bool>"
  }
}

Field semantics:

  • review_quality=valid — review file passes R4 checks and content matches a normal review shape.
  • review_quality=degraded_environmental — Codex returned an exit-0 pseudo-review caused by sandbox or environment failure (bwrap-EPERM, trust prompt, rate-limit stub, etc.). The text may contain VERDICT: and severity tags but describes an inability to perform the review.
  • review_quality=degraded_content — review parses cleanly but the runner's cheap heuristics suggest the body is not actionable (e.g. only a sandbox self-report, no concrete findings backing severity tags). Conservative catch — when in doubt, mark valid and let the lead decide.
  • review_quality=unknown — triage could not classify (e.g. triage step itself crashed). The lead should treat this as valid plus a warning.
  • triage.status=ok — triage ran, fields populated.
  • triage.status=skipped — triage skipped because Codex itself failed (no review to triage).
  • triage.status=failed — triage crashed; counts and severity may be missing. Review remains usable if review_quality=valid.

Lead-side dispatch is operation-aware because degraded_environmental on the first dispatch has no prior valid round to fall back on. The full table:

result OPERATION review_quality Lead action
success any valid Proceed to Step 5: show review verbatim, run evaluation matrix, advance round.
success any degraded_content Emit user_warning, show review verbatim, ask operator whether to advance the round.
success any unknown Emit user_warning, treat as valid for advancement, note in final operator summary.
success initial degraded_environmental Emit user_warning, do NOT show review verbatim, treat as terminal infrastructure failure. No prior valid review exists to recover from.
success resume degraded_environmental Emit user_warning, do NOT show review verbatim, do NOT count as a round, route to fallback chain using prior round's severity.
success fresh-exec degraded_environmental Emit user_warning, do NOT show review verbatim, terminal not-verified. Fresh-exec already was the fallback; a second environmental failure means the env is reliably broken.
timeout any n/a Terminal at main per existing rules.
launch_failure any n/a Terminal at main per existing rules (resume launch_failure routes to fresh-exec fallback per existing §7.4).
infra_error any n/a Show errors, abort.
input_error any n/a Show errors, abort (orchestration bug).

Backward compatibility: a runner that omits review_quality or triage fields entirely is treated as review_quality=unknown with triage.status=skipped. Existing field semantics (result, verdict, review_file, codex_session_id, user_warning, etc.) are unchanged.

Reviewer Behavior

The Codex reviewer is an auditor, not a contributor. The prompt should allow useful verification while prohibiting project mutation:

You may run commands to verify findings when useful.
Do not create, edit, delete, commit, or apply fixes to project files.
Prefer commands that do not mutate the working tree.
Do not run commands likely to rewrite generated files, snapshots, migrations,
lockfiles, or configs.
If verification would require mutation, report that limitation instead.
If a command unexpectedly changes files, stop and report it.

This is a deliberate compromise. A reviewer locked into a strict read-only sandbox often cannot run the checks needed to validate tool behavior, tests, or documentation-dependent claims. The primary safeguard is the instruction contract plus post-dispatch mutation detection, not a hard read-only shell.

Lead Responsibilities

The lead owns all product decisions. After each valid Codex review:

  1. Show the Codex review to the operator verbatim before applying fixes.

  2. Use runner triage only as a hint.

  3. Build a compact evaluation matrix:

    finding | severity | evidence | action | reason
    
  4. Choose one of three first-class actions for each finding: accept, re-scope, or reject with reasoning.

  5. Apply only accepted or re-scoped fixes.

  6. Send the reviewer a structured re-review prompt with Applied, Re-scoped, and Rejected with reasoning sections.

Reviewer findings are suggestions to evaluate, not orders to follow. The lead should verify findings proportionally to risk. Critical and high findings need more scrutiny; medium findings can often be accepted, narrowed, or rejected based on local checks and reasoning.

For plan reviews, the lead and reviewer must judge the plan at its declared level of abstraction. Missing implementation details are findings only when their absence blocks feasibility, safety, rollback, verification, or a public contract.

Operator Language

The skill should detect the operator language from recent user messages and ask Codex to respond in that language. If detection is unclear, use English.

Prompt block:

<language>
Respond in the operator's language: <detected language>.
Keep these machine-readable literals unchanged in English:
- [severity: critical|high|medium]
- VERDICT: APPROVED
- VERDICT: REVISE
</language>

Runtime prose shown to the operator should use the operator's language when practical. Repository documentation and skill files remain English.

Prompt Changes

Plan review prompts should add abstraction-level calibration:

Judge the plan at its declared level of abstraction.
Do not demand implementation details unless their absence blocks feasibility,
safety, rollback, verification, or a public contract.
If a detail can reasonably be decided during implementation, do not count it
as a finding.

Re-review prompts should replace the current "I've revised based on your feedback" shape with:

I've evaluated the findings.

## Applied
- ...

## Re-scoped
- ...

## Rejected with reasoning
- ...

## Specific asks for re-review
1. Are my rejections technically valid?
2. Did the applied/re-scoped fixes resolve the original findings?
3. Did the fixes introduce new issues?

Sonnet triage metadata should not be passed to Codex automatically. Codex sees verbatim findings and the lead's decisions, not the runner's notes.

Structural Operator Gate

Before applying structural fixes, the lead should pause once and ask the operator to approve the batch unless the operator explicitly requested autonomous work.

Structural changes include:

  • invocation grammar or argument semantics;
  • output format or parsed literals;
  • workflow steps, fallback semantics, or terminal states;
  • sandbox, approval, or security guarantees;
  • public configuration semantics;
  • schema, migration, or data format changes;
  • broad architectural rewrites;
  • any fix whose scope the lead is uncertain about.

Non-structural fixes include wording, factual clarifications, examples, and local changes that do not alter external behavior.

If the operator explicitly requested autonomous mode, the lead may apply structural fixes without pausing, but the final summary must state that structural changes were applied without operator sign-off due to autonomous mode.

Workspace Mutation Detection

Mutation detection runs at two layers: the repo tree (git-tracked + new untracked files) and the skill's own /tmp review inputs. A third class of mutation — gitignored files already inside REPO_ROOT — is documented as a known, operator-mitigated risk rather than detected automatically.

Repo tree. The main thread should snapshot workspace state with git status --porcelain before and after each runner dispatch, before Step 5 applies any fixes.

If tracked files changed during the runner dispatch, the skill must hard stop before applying fixes and show an operator diagnostic. This catches reviewer or runner mutation without requiring a custom agent type.

If only untracked generated artifacts appeared, the skill should warn and gate continuation. Some tools leave local artifacts, but the skill must not silently fold them into fixes.

Review inputs in /tmp. The materialized plan file (/tmp/codex-plan-<REVIEW_ID>.md when present), the prompt body (/tmp/codex-body-<REVIEW_ID>.md), and the resume body (/tmp/codex-resume-body-<REVIEW_ID>.md) sit outside REPO_ROOT and are not covered by git status. The main thread must hash these files before dispatch and re-hash after the runner returns:

sha256sum /tmp/codex-{body,plan,resume-body}-<REVIEW_ID>.md 2>/dev/null \
  > /tmp/codex-inputs-pre-<REVIEW_ID>.sha
# ... runner dispatch ...
sha256sum /tmp/codex-{body,plan,resume-body}-<REVIEW_ID>.md 2>/dev/null \
  > /tmp/codex-inputs-post-<REVIEW_ID>.sha
diff -q /tmp/codex-inputs-pre-<REVIEW_ID>.sha \
        /tmp/codex-inputs-post-<REVIEW_ID>.sha

A non-empty diff is treated identically to a tracked-file mutation: hard stop before applying fixes and surface the operator diagnostic.

Out-of-scope: gitignored files inside REPO_ROOT. .gitignored files that already exist inside REPO_ROOT (local SQLite DBs, .env.local, service-state directories, build caches) are NOT detected by either layer. Full-tree snapshotting would be too expensive to run every round, and a pre-dispatch confirmation about "ignored files present" would trigger on essentially every repo (node_modules, target/, .next/) and produce confirmation fatigue — security theater rather than protection.

The realistic mutation vector for these files is the reviewer running a project test or build command that side-effects on the ignored file — for example pytest triggering an unintended migration on dev.sqlite because the test settings point at it. The damage is bounded (the state is re-seedable; tests generally read .env.local, they do not write to it), but it is real on default workspace-write reviews.

The skill mitigates this with three operator-facing affordances rather than architectural force:

  1. The sandbox:read-only override is always available — operators who know they have sensitive ignored state can opt out of test execution at the cost of empirical verification (the reviewer loses the ability to run tests, linters, and most MCP-backed verification).
  2. The README must include a "Safety considerations" section with the concrete pytest → dev.sqlite vector, the opt-out guidance (with its explicit verification trade-off), and the worktree-isolation pattern (git worktree add /tmp/review-worktree <ref>) for sensitive repos.
  3. At the start of every review, the skill emits a one-line runtime hint: workspace-write in effect; pass sandbox:read-only if sensitive ignored state lives under REPO_ROOT. The hint appears exactly once per review (not per round) and is suppressed when the operator passed an explicit sandbox:* override.

This is a deliberate trade-off. Defaulting reviews to read-only or to an isolated worktree would gut the reviewer's empirical verification capability (tests, linters, builds, MCP doc lookups, web searches all require execute) — and empirical verification is precisely what makes adversarial review more valuable than a same-model self-check. The mitigation level is documentation and operator awareness rather than architectural force.

Linux Sandbox Prerequisites

The README should document sandbox prerequisites for Linux users and installer agents. With the default workspace-write mode, Codex may rely on bubblewrap and unprivileged user namespaces.

Diagnostic probe:

bwrap --dev-bind / / --unshare-net /bin/echo ok

If the probe fails on Ubuntu 24.04+ due to AppArmor user namespace restrictions, recommend the official bwrap AppArmor profile:

sudo apt install -y apparmor-profiles apparmor-utils
sudo install -m 0644 /usr/share/apparmor/extra-profiles/bwrap-userns-restrict \
  /etc/apparmor.d/bwrap-userns-restrict
sudo apparmor_parser -r /etc/apparmor.d/bwrap-userns-restrict

Then verify:

bwrap --dev-bind / / --unshare-net /bin/echo ok

The README should link to the official OpenAI Codex sandboxing documentation: https://developers.openai.com/codex/concepts/sandboxing.

Installer agents should not change AppArmor policy silently. They should probe, show the exact commands, request explicit permission, apply the profile only after approval, and re-run the probe.

Final Operator Summary

At every terminal state, the skill must provide an operator-facing summary in the operator's language. This summary comes after the final verbatim reviewer response and does not replace it.

Include:

  • final status: approved, maximum rounds reached, not verified, or aborted;
  • what changed across all review rounds;
  • findings applied, re-scoped, and rejected;
  • structural changes and whether operator sign-off was obtained;
  • verification performed and verification not performed;
  • remaining findings or risks;
  • a short explanation of what the status means for non-approved terminal states.

Constraints:

  • Do not include full diffs.
  • Do not repeat full reviewer findings unless unresolved findings matter.
  • Keep it concise and useful to the operator.
  • Build it from per-round decision summaries already shown in conversation.
  • Do not read Codex stdout, stderr, or rollout files from the main thread.
  • If context compaction makes history incomplete, state that limitation explicitly instead of inventing details.

Documentation Language Rule

Repository documentation and skill files stay in English:

  • SKILL.md
  • references/runner.md
  • README.md
  • docs/DESIGN.md
  • design specs under docs/superpowers/specs/

Runtime output may use the operator's language. Do not mix Russian and English inside repository documentation paragraphs except for machine-readable literals, CLI flags, config keys, JSON keys, or quoted runtime examples.

Files To Update

Expected implementation touches:

  • SKILL.md

    • default model, sandbox, approval parsing (note: approval policy is expressed via -c approval_policy=..., not the removed -a flag);
    • language detection and language block;
    • runner input schema extensions;
    • runner result schema consumption (review_quality, triage.*) and the operation-aware lead-side dispatch table from §Runner Result Schema;
    • evaluation matrix;
    • structural operator gate;
    • structured resume prompt;
    • workspace mutation snapshots at both layers (git status --porcelain AND sha256sum of /tmp/codex-{body,plan,resume-body}-*);
    • one-line runtime hint at the start of every review about the workspace-write default and the sandbox:read-only opt-out (suppressed when an explicit sandbox:* override is passed);
    • final operator summary.
  • references/runner.md

    • default Codex command flags (config-based approval policy via -c);
    • bwrap preflight before every dispatch that selects a bwrap-backed sandbox mode;
    • degraded-review detection;
    • bounded triage metadata;
    • emit the full runner result schema (§Runner Result Schema), including review_quality and the triage object;
    • stronger mandate and negative examples.
  • docs/DESIGN.md

    • rationale for reviewer permissions;
    • default model update;
    • nested approval reasoning;
    • Sonnet triage compromise;
    • operator-language behavior;
    • final summary rationale;
    • version and verification log update after smoke testing.
  • README.md

    • updated defaults (including config-based approval policy via -c, not -a);
    • sandbox/approval behavior;
    • Safety considerations section: the gitignored-file mutation vector (pytest → dev.sqlite migration), residual risk acknowledgment, the sandbox:read-only opt-out (with the explicit caveat that read-only blocks empirical verification — operators trade verification for write-protection), and the git worktree-based isolation pattern for sensitive repos;
    • Linux bwrap/AppArmor setup;
    • operator language behavior;
    • final summary behavior.

Do not renumber existing docs/DESIGN.md sections.

Verification

Use the existing smoke protocol in docs/DESIGN.md §7, updated for the new model and flags where needed.

Additional manual smoke checks:

  • codex exec launches with -s workspace-write, -c approval_policy='"on-request"', and -c approvals_reviewer='"auto_review"' on the target Codex CLI version (verify against the actual installed version — codex exec --help must list -s and -c; the absence of -a/--ask-for-approval is expected on 0.132+).
  • codex exec resume does not receive unsupported sandbox or approval flags.
  • bwrap-backed sandbox failure for workspace-write produces an early diagnostic before Codex launches, instead of a fake review round.
  • sandbox:read-only override still works and the runner skips approval prompts inside the nested process (the operator can verify this by running a review against a known-bwrap-failing host with the override).
  • language block preserves parseable [severity:] and VERDICT literals.
  • runner result JSON contains review_quality and triage fields when triage runs; legacy result JSON (no review_quality, no triage) is accepted by the lead as review_quality=unknown / triage.status=skipped.
  • triage failure does not fail an otherwise valid review.
  • a synthetic success + degraded_environmental on OPERATION=initial is treated as terminal infrastructure failure (no prior round to recover from); the verbatim review is NOT shown.
  • a synthetic success + degraded_environmental on OPERATION=resume emits the warning, does not consume a round, and routes to the existing fallback chain using the prior round's severity.
  • a synthetic success + degraded_environmental on OPERATION=fresh-exec is treated as terminal not-verified.
  • a synthetic success + degraded_content (on any operation) emits the warning, shows Step 5 verbatim, and prompts the operator for explicit advancement.
  • pre/post git status --porcelain catches workspace-tree mutation.
  • pre/post sha256sum of /tmp/codex-{body,plan,resume-body}-<REVIEW_ID>.md catches reviewer mutation of /tmp review inputs.
  • the runtime hint about workspace-write and the sandbox:read-only opt-out appears exactly once per review (regardless of mode) and is suppressed when the operator passed an explicit sandbox:* override.
  • final operator summary appears at approved, max-rounds, not-verified, and aborted terminal states.

Dogfood the result with:

  • plan review of this design;
  • code review after implementation.

No automated CI is required for this refresh.