feat(skill): refresh adversarial-review под Codex 0.132
- Зачем:
- актуализация под новый Codex CLI (0.132 убрал флаг -a,
ввёл -c approval_policy / -c approvals_reviewer);
- переход на безопасные операторские дефолты
(sandbox=workspace-write для эмпирической верификации,
явные overrides sandbox:* и approvals:*);
- наведение порядка в принятии решений по раунду ревью:
review_quality + evaluation matrix + structural operator gate
вместо «оператор сам разбирается».
- Что:
- SKILL.md: override-граммар (sandbox:*, approvals:*, model:*,
low/medium/high/xhigh), детект OPERATOR_LANGUAGE, runtime hint,
conditional dirty-file warning в Step 2 (после захвата REPO_ROOT),
dual-layer mutation snapshots (git status + sha256sum input-файлов),
operation-aware dispatch table, evaluation matrix + structural
operator gate (batch-pause), structured resume body
(Applied / Re-scoped / Rejected / Specific asks), final operator
summary на OPERATOR_LANGUAGE.
- references/runner.md: переведён на Sonnet runner, добавлен
Step R2.5 bwrap preflight, Step R4.5 review_quality + bounded
triage, расширена 11-полевая JSON-схема результата.
- docs/DESIGN.md: §4.14–§4.21 с обоснованиями новых решений,
§7.8 refresh-era smoke checks, §8 version log с двумя раундами
dogfood-а этого refresh-а (включая R2-корректировку
обоснования residual gap для уже-грязных tracked-файлов).
- README.md: таблица дефолтов, Safety considerations (честно
задокументирован residual gap), Linux sandbox prerequisites
(bwrap + AppArmor user namespaces), Operator language,
Final operator summary, troubleshooting.
- examples/review-output.md: модель в сэмпле обновлена на gpt-5.5.
- docs/superpowers/specs/2026-05-20-...: сохранена спека дизайна
с уточнениями после dogfood.
- Проверка:
- dogfood: 2 раунда /adversarial-review code на этом refresh-е,
R2 верифицировал R2#1 (ordering) и R2#2 (Option B, honest
residual gap);
- smoke: codex --version ≥ 0.132.0, bwrap preflight зелёный
на reference WSL2.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,718 @@
|
||||
# Adversarial Review Refresh Design
|
||||
|
||||
## Overview
|
||||
|
||||
This design updates the `adversarial-review` skill while preserving its core
|
||||
shape: Claude remains the lead, Codex provides an external cross-model
|
||||
adversarial review, and the Sonnet runner isolates mechanical Codex CLI work
|
||||
from the expensive lead context.
|
||||
|
||||
The refresh makes the skill more useful to the operator, more robust on modern
|
||||
Codex CLI setups, and less prone to blindly applying reviewer feedback. It does
|
||||
not turn the skill into a larger framework, add mandatory custom agent
|
||||
definitions, or move final decisions from the lead into the runner.
|
||||
|
||||
## Goals
|
||||
|
||||
- Use `gpt-5.5` as the default Codex reviewer model while preserving explicit
|
||||
model and reasoning overrides.
|
||||
- Let Codex reviews use the operator's language while keeping parsed literals
|
||||
stable.
|
||||
- Improve reviewer capability by avoiding overly restrictive default sandboxing.
|
||||
- Avoid nested human approval prompts that the operator cannot see from the
|
||||
Claude session.
|
||||
- Keep runner-side work bounded and disposable so the lead context stays clean.
|
||||
- Add compact Sonnet triage to reduce repeated review rounds without letting
|
||||
Sonnet make final decisions.
|
||||
- Require the lead to evaluate findings before applying fixes.
|
||||
- Support operator sign-off for structural fixes without blocking explicitly
|
||||
autonomous runs.
|
||||
- Provide a final operator-facing summary of what changed across the review.
|
||||
- Document Linux sandbox prerequisites clearly enough for a human or installer
|
||||
agent to act on them.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Do not require a custom Claude agent type or extra installation artifact.
|
||||
- Do not use `danger-full-access` as an automatic fallback.
|
||||
- Do not make the runner apply fixes, start extra review rounds, or decide
|
||||
final `accept` / `reject` / `re-scope` outcomes.
|
||||
- Do not move Codex stdout, stderr, or rollout contents into the main Claude
|
||||
context.
|
||||
- Do not add a mandatory detailed implementation plan for small instructional
|
||||
changes.
|
||||
- Do not mix documentation languages in repository files.
|
||||
|
||||
## Default Reviewer Configuration
|
||||
|
||||
The default Codex invocation should be optimized for non-interactive review
|
||||
inside a Claude-runner child process:
|
||||
|
||||
```text
|
||||
CODEX_MODEL = gpt-5.5
|
||||
CODEX_REASONING = high
|
||||
CODEX_SANDBOX = workspace-write
|
||||
CODEX_APPROVAL_POLICY = on-request
|
||||
CODEX_APPROVALS_REVIEWER = auto_review
|
||||
```
|
||||
|
||||
`workspace-write` is the default for all review modes (plan, code,
|
||||
code-vs-plan). The reasoning is load-bearing for this skill, so it is
|
||||
captured here rather than in a comment.
|
||||
|
||||
The reviewer needs to actually run things to verify findings: run
|
||||
tests, build the project, query the web or upstream APIs to confirm
|
||||
current behavior, and exercise project CLIs end-to-end. These are
|
||||
write-class operations — test runners produce output and cache files,
|
||||
build commands write artifacts, and most non-trivial CLI execution
|
||||
touches local state. A read-only sandbox blocks all of that.
|
||||
|
||||
Read-only is not a *total* verification blocker. It still allows file
|
||||
inspection, `rg` / `grep` searches, MCP-backed doc lookups, and
|
||||
`--help`-style CLI introspection that does not write to the workspace.
|
||||
But the highest-value findings come from the path read-only blocks:
|
||||
"the reviewer ran the tests and X failed", "the reviewer built the
|
||||
project and Y broke", "the reviewer queried the live API and the
|
||||
assumed signature does not exist." Forfeiting those to gain protection
|
||||
against `.gitignored`-file side effects is a poor trade.
|
||||
|
||||
This applies to plan reviews as much as to code. A plan reviewer
|
||||
should be able to run a test the plan relies on, build the project to
|
||||
confirm a structural claim, or hit an external API to validate an
|
||||
assumption — not just `rg` for cited filenames. A previous design
|
||||
pass made `read-only` the default for plan mode and was reverted for
|
||||
this reason.
|
||||
|
||||
The safety argument for forcing read-only is also weaker than it
|
||||
appears. The broader `obra:superpowers` skill family demonstrates that
|
||||
careful prompt-level discipline is sufficient to govern complex agent
|
||||
behavior in security-relevant contexts without sandbox-level
|
||||
restrictions. A concrete instance is `superpowers:receiving-code-review`:
|
||||
it constrains how a lead processes adversarial feedback — read every
|
||||
finding before reacting, restate the technical claim in own words,
|
||||
verify empirically (especially tool-mechanic claims) before accepting,
|
||||
push back with technical reasoning when wrong, no performative
|
||||
agreement — and that discipline is enforced **entirely through prompt
|
||||
instructions**, not through any tool-level sandboxing of the lead.
|
||||
The same pattern applies to the reviewer side of this skill: a strict
|
||||
instruction contract (see §Reviewer Behavior), backed by pre/post
|
||||
`git status --porcelain` and the `/tmp` sha256 snapshot for
|
||||
after-the-fact detection.
|
||||
|
||||
Read-only's specific failure modes are covered without a default
|
||||
sandbox change:
|
||||
|
||||
- Tracked-file mutation is caught by `git status --porcelain` in
|
||||
§Workspace Mutation Detection.
|
||||
- `/tmp` review-input mutation is caught by the `sha256sum` snapshot
|
||||
in the same section.
|
||||
- Mutation of `.gitignored` state inside `REPO_ROOT` is bounded
|
||||
(re-seedable dev state; tests read `.env.local`, they do not write
|
||||
it) and is mitigated through documentation plus the
|
||||
`sandbox:read-only` override for operators who knowingly accept the
|
||||
trade-off.
|
||||
|
||||
Operators with a concrete reason to opt into read-only (sensitive
|
||||
local state, untrusted reviewer prompt source, or any other case where
|
||||
sandbox-level write protection outweighs verification capability) can
|
||||
do so per-run via `sandbox:read-only`. The override stays available
|
||||
exactly because the trade-off is real — for those operators.
|
||||
Defaulting the rest of the user base to that mode would gut the
|
||||
skill's main value.
|
||||
|
||||
The intended initial `codex exec` shape:
|
||||
|
||||
```bash
|
||||
codex exec --json \
|
||||
-m gpt-5.5 \
|
||||
-c model_reasoning_effort=high \
|
||||
-s workspace-write \
|
||||
-c approval_policy='"on-request"' \
|
||||
-c approvals_reviewer='"auto_review"' \
|
||||
-C "<REPO_ROOT>" \
|
||||
-o /tmp/codex-review-<REVIEW_ID>.md \
|
||||
-
|
||||
```
|
||||
|
||||
Codex CLI 0.132+ removed the top-level `--ask-for-approval` / `-a` flag
|
||||
from `codex exec`; approval policy is now expressed only through
|
||||
`-c approval_policy=<value>`. The spec uses `-c` form for both
|
||||
`approval_policy` and `approvals_reviewer` so the command shape stays
|
||||
self-consistent and future-proof against further short-flag churn. The
|
||||
implementation should verify the supported approval-control surface
|
||||
against the actual installed Codex CLI version during smoke testing.
|
||||
|
||||
`-o` writes only the assistant's final message. The reviewer prompt must
|
||||
deliver the full structured review (findings, severity tags, VERDICT)
|
||||
inside that final message; intermediate tool calls and reasoning will
|
||||
not appear in the review file.
|
||||
|
||||
`workspace-write` gives the reviewer enough capability to run local
|
||||
checks without forcing every useful command through a read-only wall.
|
||||
`approval_policy="on-request"` keeps boundary crossings explicit.
|
||||
`approvals_reviewer="auto_review"` avoids human approval prompts inside
|
||||
the nested Codex process, which the operator cannot reliably see or
|
||||
answer from the parent Claude session.
|
||||
|
||||
Auto-review is not a security boundary. It reviews approval requests; it does
|
||||
not inspect actions already permitted by the selected sandbox. The skill still
|
||||
needs strong reviewer instructions and workspace mutation checks.
|
||||
|
||||
Existing overrides remain:
|
||||
|
||||
- `model:<name>`
|
||||
- `low`, `medium`, `high`, `xhigh`
|
||||
|
||||
New minimal overrides:
|
||||
|
||||
- `sandbox:read-only | workspace-write | danger-full-access | inherit`
|
||||
- `approvals:user | auto_review | never`
|
||||
|
||||
Rules:
|
||||
|
||||
- Default `-s` is `workspace-write` for every review mode. Read-only is a
|
||||
deliberate operator opt-in via `sandbox:read-only`, never an automatic
|
||||
per-mode default.
|
||||
- `sandbox:inherit` omits `-s` and relies on the user's Codex config.
|
||||
Because the effective sandbox is unknown until Codex actually
|
||||
launches, the runner skips bwrap preflight under `inherit`. If the
|
||||
inherited config selects a bwrap-backed mode on a host where bwrap
|
||||
is misconfigured, the failure surfaces as a §Runner Result Schema
|
||||
`degraded_environmental` result and is treated as terminal
|
||||
infrastructure failure on the initial dispatch per the dispatch table.
|
||||
- `approvals:auto_review` (the default) passes
|
||||
`-c approval_policy='"on-request"'` plus
|
||||
`-c approvals_reviewer='"auto_review"'`.
|
||||
- `approvals:user` passes `-c approval_policy='"on-request"'` only,
|
||||
omitting the `approvals_reviewer` override so Codex falls back to its
|
||||
default `user` reviewer. Allowed only by explicit override because
|
||||
nested human approvals can hang the run.
|
||||
- `approvals:never` passes `-c approval_policy='"never"'`; boundary
|
||||
crossings fail instead of asking.
|
||||
- `sandbox:danger-full-access` is explicit-only and should surface a warning.
|
||||
- Codex's `untrusted` approval policy is intentionally not exposed as an
|
||||
override; the skill needs predictable boundary semantics, not per-command
|
||||
trust prompts.
|
||||
- No silent fallback may change sandbox or approval semantics.
|
||||
|
||||
Resume commands should respect Codex CLI support for `resume`: sandbox and
|
||||
approval mode are properties of the initial session unless the current CLI
|
||||
explicitly supports changing them on resume. Do not pass unsupported `-s`
|
||||
flags or approval-related `-c` overrides to `codex exec resume`.
|
||||
|
||||
## Runner Responsibilities
|
||||
|
||||
The runner remains a Sonnet subagent responsible for Codex CLI mechanics:
|
||||
|
||||
- parse the input block;
|
||||
- write the prompt with the attempt-scoped session marker;
|
||||
- launch exactly one Codex operation (`initial`, `resume`, or `fresh-exec`);
|
||||
- own the one internal retry budget;
|
||||
- validate exit code, stderr, review file, and session id;
|
||||
- archive failed-resume diagnostics;
|
||||
- write the authoritative result JSON;
|
||||
- return the `RUNNER_RESULT_AT: <path>` line.
|
||||
|
||||
The runner gains bounded analysis:
|
||||
|
||||
- run sandbox preflight when the selected mode requires a bwrap-backed sandbox;
|
||||
- detect obvious degraded reviews, such as sandbox or environment failures
|
||||
disguised as reviewer output;
|
||||
- extract finding count, maximum severity, and review quality;
|
||||
- produce compact triage metadata.
|
||||
|
||||
Triage rules:
|
||||
|
||||
- Cover all critical and high findings.
|
||||
- Cover up to 10 medium findings.
|
||||
- If more medium findings remain, summarize the remainder and set an explicit
|
||||
truncation flag.
|
||||
- Use only cheap checks: file existence, section existence, simple `rg`, and
|
||||
small read-only snippets.
|
||||
- Do not run long test suites, web searches, or broad documentation lookups
|
||||
during triage.
|
||||
- Do not produce final `accept`, `reject`, or `re-scope` decisions.
|
||||
- If uncertain, mark that lead judgment is needed.
|
||||
|
||||
If triage fails but the Codex review is valid, the review should continue with
|
||||
a warning. If the review itself is degraded by infrastructure failure, the
|
||||
runner should not count it as a normal round.
|
||||
|
||||
Runner instructions must include negative examples:
|
||||
|
||||
- Do not edit project files.
|
||||
- Do not apply fixes.
|
||||
- Do not run multiple Codex review rounds in one dispatch.
|
||||
- Do not delete `/tmp/codex-*` files.
|
||||
- Do not decide which findings the lead must accept.
|
||||
- If a command unexpectedly changes project files, stop and report it.
|
||||
|
||||
## Runner Result Schema
|
||||
|
||||
The runner writes a single JSON file at `RESULT_PATH`. The schema below is
|
||||
the authoritative contract between runner and lead. New fields added by
|
||||
this refresh are marked `(new)`; everything else preserves the pre-refresh
|
||||
shape.
|
||||
|
||||
```json
|
||||
{
|
||||
"result": "success | timeout | launch_failure | infra_error | input_error",
|
||||
"verdict": "APPROVED | REVISE | null",
|
||||
"review_file": "<absolute path or null>",
|
||||
"codex_session_id": "<uuid or null>",
|
||||
"attempt_id": "<string>",
|
||||
"errors": "<string or null>",
|
||||
"archived_stdout": "<path or null>",
|
||||
"archived_stderr": "<path or null>",
|
||||
"user_warning": "<string or null>",
|
||||
"review_quality": "valid | degraded_environmental | degraded_content | unknown", // (new)
|
||||
"triage": { // (new)
|
||||
"status": "ok | skipped | failed",
|
||||
"finding_count": "<int>",
|
||||
"max_severity": "critical | high | medium | none",
|
||||
"covered_critical": "<int>",
|
||||
"covered_high": "<int>",
|
||||
"covered_medium": "<int>",
|
||||
"truncated": "<bool>",
|
||||
"needs_lead_judgment": "<bool>"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Field semantics:
|
||||
|
||||
- `review_quality=valid` — review file passes R4 checks and content matches
|
||||
a normal review shape.
|
||||
- `review_quality=degraded_environmental` — Codex returned an exit-0
|
||||
pseudo-review caused by sandbox or environment failure (bwrap-EPERM,
|
||||
trust prompt, rate-limit stub, etc.). The text may contain `VERDICT:`
|
||||
and severity tags but describes an inability to perform the review.
|
||||
- `review_quality=degraded_content` — review parses cleanly but the
|
||||
runner's cheap heuristics suggest the body is not actionable (e.g.
|
||||
only a sandbox self-report, no concrete findings backing severity
|
||||
tags). Conservative catch — when in doubt, mark `valid` and let the
|
||||
lead decide.
|
||||
- `review_quality=unknown` — triage could not classify (e.g. triage step
|
||||
itself crashed). The lead should treat this as `valid` plus a warning.
|
||||
- `triage.status=ok` — triage ran, fields populated.
|
||||
- `triage.status=skipped` — triage skipped because Codex itself failed
|
||||
(no review to triage).
|
||||
- `triage.status=failed` — triage crashed; counts and severity may be
|
||||
missing. Review remains usable if `review_quality=valid`.
|
||||
|
||||
Lead-side dispatch is **operation-aware** because `degraded_environmental`
|
||||
on the first dispatch has no prior valid round to fall back on. The full
|
||||
table:
|
||||
|
||||
| `result` | `OPERATION` | `review_quality` | Lead action |
|
||||
|------------------|---------------|--------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `success` | any | `valid` | Proceed to Step 5: show review verbatim, run evaluation matrix, advance round. |
|
||||
| `success` | any | `degraded_content` | Emit `user_warning`, show review verbatim, ask operator whether to advance the round. |
|
||||
| `success` | any | `unknown` | Emit `user_warning`, treat as `valid` for advancement, note in final operator summary. |
|
||||
| `success` | `initial` | `degraded_environmental` | Emit `user_warning`, do NOT show review verbatim, treat as terminal infrastructure failure. No prior valid review exists to recover from. |
|
||||
| `success` | `resume` | `degraded_environmental` | Emit `user_warning`, do NOT show review verbatim, do NOT count as a round, route to fallback chain using prior round's severity. |
|
||||
| `success` | `fresh-exec` | `degraded_environmental` | Emit `user_warning`, do NOT show review verbatim, terminal not-verified. Fresh-exec already was the fallback; a second environmental failure means the env is reliably broken. |
|
||||
| `timeout` | any | n/a | Terminal at main per existing rules. |
|
||||
| `launch_failure` | any | n/a | Terminal at main per existing rules (`resume` `launch_failure` routes to fresh-exec fallback per existing §7.4). |
|
||||
| `infra_error` | any | n/a | Show `errors`, abort. |
|
||||
| `input_error` | any | n/a | Show `errors`, abort (orchestration bug). |
|
||||
|
||||
Backward compatibility: a runner that omits `review_quality` or `triage`
|
||||
fields entirely is treated as `review_quality=unknown` with
|
||||
`triage.status=skipped`. Existing field semantics (`result`, `verdict`,
|
||||
`review_file`, `codex_session_id`, `user_warning`, etc.) are unchanged.
|
||||
|
||||
## Reviewer Behavior
|
||||
|
||||
The Codex reviewer is an auditor, not a contributor. The prompt should allow
|
||||
useful verification while prohibiting project mutation:
|
||||
|
||||
```text
|
||||
You may run commands to verify findings when useful.
|
||||
Do not create, edit, delete, commit, or apply fixes to project files.
|
||||
Prefer commands that do not mutate the working tree.
|
||||
Do not run commands likely to rewrite generated files, snapshots, migrations,
|
||||
lockfiles, or configs.
|
||||
If verification would require mutation, report that limitation instead.
|
||||
If a command unexpectedly changes files, stop and report it.
|
||||
```
|
||||
|
||||
This is a deliberate compromise. A reviewer locked into a strict read-only
|
||||
sandbox often cannot run the checks needed to validate tool behavior, tests, or
|
||||
documentation-dependent claims. The primary safeguard is the instruction
|
||||
contract plus post-dispatch mutation detection, not a hard read-only shell.
|
||||
|
||||
## Lead Responsibilities
|
||||
|
||||
The lead owns all product decisions. After each valid Codex review:
|
||||
|
||||
1. Show the Codex review to the operator verbatim before applying fixes.
|
||||
2. Use runner triage only as a hint.
|
||||
3. Build a compact evaluation matrix:
|
||||
|
||||
```text
|
||||
finding | severity | evidence | action | reason
|
||||
```
|
||||
|
||||
4. Choose one of three first-class actions for each finding:
|
||||
`accept`, `re-scope`, or `reject with reasoning`.
|
||||
5. Apply only accepted or re-scoped fixes.
|
||||
6. Send the reviewer a structured re-review prompt with `Applied`,
|
||||
`Re-scoped`, and `Rejected with reasoning` sections.
|
||||
|
||||
Reviewer findings are suggestions to evaluate, not orders to follow. The lead
|
||||
should verify findings proportionally to risk. Critical and high findings need
|
||||
more scrutiny; medium findings can often be accepted, narrowed, or rejected
|
||||
based on local checks and reasoning.
|
||||
|
||||
For plan reviews, the lead and reviewer must judge the plan at its declared
|
||||
level of abstraction. Missing implementation details are findings only when
|
||||
their absence blocks feasibility, safety, rollback, verification, or a public
|
||||
contract.
|
||||
|
||||
## Operator Language
|
||||
|
||||
The skill should detect the operator language from recent user messages and
|
||||
ask Codex to respond in that language. If detection is unclear, use English.
|
||||
|
||||
Prompt block:
|
||||
|
||||
```text
|
||||
<language>
|
||||
Respond in the operator's language: <detected language>.
|
||||
Keep these machine-readable literals unchanged in English:
|
||||
- [severity: critical|high|medium]
|
||||
- VERDICT: APPROVED
|
||||
- VERDICT: REVISE
|
||||
</language>
|
||||
```
|
||||
|
||||
Runtime prose shown to the operator should use the operator's language when
|
||||
practical. Repository documentation and skill files remain English.
|
||||
|
||||
## Prompt Changes
|
||||
|
||||
Plan review prompts should add abstraction-level calibration:
|
||||
|
||||
```text
|
||||
Judge the plan at its declared level of abstraction.
|
||||
Do not demand implementation details unless their absence blocks feasibility,
|
||||
safety, rollback, verification, or a public contract.
|
||||
If a detail can reasonably be decided during implementation, do not count it
|
||||
as a finding.
|
||||
```
|
||||
|
||||
Re-review prompts should replace the current "I've revised based on your
|
||||
feedback" shape with:
|
||||
|
||||
```text
|
||||
I've evaluated the findings.
|
||||
|
||||
## Applied
|
||||
- ...
|
||||
|
||||
## Re-scoped
|
||||
- ...
|
||||
|
||||
## Rejected with reasoning
|
||||
- ...
|
||||
|
||||
## Specific asks for re-review
|
||||
1. Are my rejections technically valid?
|
||||
2. Did the applied/re-scoped fixes resolve the original findings?
|
||||
3. Did the fixes introduce new issues?
|
||||
```
|
||||
|
||||
Sonnet triage metadata should not be passed to Codex automatically. Codex sees
|
||||
verbatim findings and the lead's decisions, not the runner's notes.
|
||||
|
||||
## Structural Operator Gate
|
||||
|
||||
Before applying structural fixes, the lead should pause once and ask the
|
||||
operator to approve the batch unless the operator explicitly requested
|
||||
autonomous work.
|
||||
|
||||
Structural changes include:
|
||||
|
||||
- invocation grammar or argument semantics;
|
||||
- output format or parsed literals;
|
||||
- workflow steps, fallback semantics, or terminal states;
|
||||
- sandbox, approval, or security guarantees;
|
||||
- public configuration semantics;
|
||||
- schema, migration, or data format changes;
|
||||
- broad architectural rewrites;
|
||||
- any fix whose scope the lead is uncertain about.
|
||||
|
||||
Non-structural fixes include wording, factual clarifications, examples, and
|
||||
local changes that do not alter external behavior.
|
||||
|
||||
If the operator explicitly requested autonomous mode, the lead may apply
|
||||
structural fixes without pausing, but the final summary must state that
|
||||
structural changes were applied without operator sign-off due to autonomous
|
||||
mode.
|
||||
|
||||
## Workspace Mutation Detection
|
||||
|
||||
Mutation detection runs at two layers: the repo tree (git-tracked + new
|
||||
untracked files) and the skill's own `/tmp` review inputs. A third class
|
||||
of mutation — gitignored files already inside `REPO_ROOT` — is documented
|
||||
as a known, operator-mitigated risk rather than detected automatically.
|
||||
|
||||
**Repo tree.** The main thread should snapshot workspace state with
|
||||
`git status --porcelain` before and after each runner dispatch, before
|
||||
Step 5 applies any fixes.
|
||||
|
||||
If tracked files changed during the runner dispatch, the skill must hard
|
||||
stop before applying fixes and show an operator diagnostic. This catches
|
||||
reviewer or runner mutation without requiring a custom agent type.
|
||||
|
||||
If only untracked generated artifacts appeared, the skill should warn and
|
||||
gate continuation. Some tools leave local artifacts, but the skill must
|
||||
not silently fold them into fixes.
|
||||
|
||||
**Review inputs in `/tmp`.** The materialized plan file
|
||||
(`/tmp/codex-plan-<REVIEW_ID>.md` when present), the prompt body
|
||||
(`/tmp/codex-body-<REVIEW_ID>.md`), and the resume body
|
||||
(`/tmp/codex-resume-body-<REVIEW_ID>.md`) sit outside `REPO_ROOT` and are
|
||||
not covered by `git status`. The main thread must hash these files before
|
||||
dispatch and re-hash after the runner returns:
|
||||
|
||||
```bash
|
||||
sha256sum /tmp/codex-{body,plan,resume-body}-<REVIEW_ID>.md 2>/dev/null \
|
||||
> /tmp/codex-inputs-pre-<REVIEW_ID>.sha
|
||||
# ... runner dispatch ...
|
||||
sha256sum /tmp/codex-{body,plan,resume-body}-<REVIEW_ID>.md 2>/dev/null \
|
||||
> /tmp/codex-inputs-post-<REVIEW_ID>.sha
|
||||
diff -q /tmp/codex-inputs-pre-<REVIEW_ID>.sha \
|
||||
/tmp/codex-inputs-post-<REVIEW_ID>.sha
|
||||
```
|
||||
|
||||
A non-empty diff is treated identically to a tracked-file mutation: hard
|
||||
stop before applying fixes and surface the operator diagnostic.
|
||||
|
||||
**Out-of-scope: gitignored files inside `REPO_ROOT`.** `.gitignored` files
|
||||
that already exist inside `REPO_ROOT` (local SQLite DBs, `.env.local`,
|
||||
service-state directories, build caches) are NOT detected by either
|
||||
layer. Full-tree snapshotting would be too expensive to run every round,
|
||||
and a pre-dispatch confirmation about "ignored files present" would
|
||||
trigger on essentially every repo (`node_modules`, `target/`, `.next/`)
|
||||
and produce confirmation fatigue — security theater rather than
|
||||
protection.
|
||||
|
||||
The realistic mutation vector for these files is the reviewer running a
|
||||
project test or build command that side-effects on the ignored file —
|
||||
for example `pytest` triggering an unintended migration on `dev.sqlite`
|
||||
because the test settings point at it. The damage is bounded (the state
|
||||
is re-seedable; tests generally read `.env.local`, they do not write to
|
||||
it), but it is real on default workspace-write reviews.
|
||||
|
||||
The skill mitigates this with three operator-facing affordances rather
|
||||
than architectural force:
|
||||
|
||||
1. The `sandbox:read-only` override is always available — operators who
|
||||
know they have sensitive ignored state can opt out of test execution
|
||||
at the cost of empirical verification (the reviewer loses the
|
||||
ability to run tests, linters, and most MCP-backed verification).
|
||||
2. The README must include a "Safety considerations" section with the
|
||||
concrete pytest → `dev.sqlite` vector, the opt-out guidance (with
|
||||
its explicit verification trade-off), and the worktree-isolation
|
||||
pattern (`git worktree add /tmp/review-worktree <ref>`) for sensitive
|
||||
repos.
|
||||
3. At the start of every review, the skill emits a one-line runtime
|
||||
hint: `workspace-write in effect; pass sandbox:read-only if
|
||||
sensitive ignored state lives under REPO_ROOT`. The hint appears
|
||||
exactly once per review (not per round) and is suppressed when the
|
||||
operator passed an explicit `sandbox:*` override.
|
||||
|
||||
This is a deliberate trade-off. Defaulting reviews to `read-only` or to
|
||||
an isolated worktree would gut the reviewer's empirical verification
|
||||
capability (tests, linters, builds, MCP doc lookups, web searches all
|
||||
require execute) — and empirical verification is precisely what makes
|
||||
adversarial review more valuable than a same-model self-check. The
|
||||
mitigation level is documentation and operator awareness rather than
|
||||
architectural force.
|
||||
|
||||
## Linux Sandbox Prerequisites
|
||||
|
||||
The README should document sandbox prerequisites for Linux users and installer
|
||||
agents. With the default `workspace-write` mode, Codex may rely on bubblewrap
|
||||
and unprivileged user namespaces.
|
||||
|
||||
Diagnostic probe:
|
||||
|
||||
```bash
|
||||
bwrap --dev-bind / / --unshare-net /bin/echo ok
|
||||
```
|
||||
|
||||
If the probe fails on Ubuntu 24.04+ due to AppArmor user namespace
|
||||
restrictions, recommend the official bwrap AppArmor profile:
|
||||
|
||||
```bash
|
||||
sudo apt install -y apparmor-profiles apparmor-utils
|
||||
sudo install -m 0644 /usr/share/apparmor/extra-profiles/bwrap-userns-restrict \
|
||||
/etc/apparmor.d/bwrap-userns-restrict
|
||||
sudo apparmor_parser -r /etc/apparmor.d/bwrap-userns-restrict
|
||||
```
|
||||
|
||||
Then verify:
|
||||
|
||||
```bash
|
||||
bwrap --dev-bind / / --unshare-net /bin/echo ok
|
||||
```
|
||||
|
||||
The README should link to the official OpenAI Codex sandboxing documentation:
|
||||
`https://developers.openai.com/codex/concepts/sandboxing`.
|
||||
|
||||
Installer agents should not change AppArmor policy silently. They should probe,
|
||||
show the exact commands, request explicit permission, apply the profile only
|
||||
after approval, and re-run the probe.
|
||||
|
||||
## Final Operator Summary
|
||||
|
||||
At every terminal state, the skill must provide an operator-facing summary in
|
||||
the operator's language. This summary comes after the final verbatim reviewer
|
||||
response and does not replace it.
|
||||
|
||||
Include:
|
||||
|
||||
- final status: approved, maximum rounds reached, not verified, or aborted;
|
||||
- what changed across all review rounds;
|
||||
- findings applied, re-scoped, and rejected;
|
||||
- structural changes and whether operator sign-off was obtained;
|
||||
- verification performed and verification not performed;
|
||||
- remaining findings or risks;
|
||||
- a short explanation of what the status means for non-approved terminal
|
||||
states.
|
||||
|
||||
Constraints:
|
||||
|
||||
- Do not include full diffs.
|
||||
- Do not repeat full reviewer findings unless unresolved findings matter.
|
||||
- Keep it concise and useful to the operator.
|
||||
- Build it from per-round decision summaries already shown in conversation.
|
||||
- Do not read Codex stdout, stderr, or rollout files from the main thread.
|
||||
- If context compaction makes history incomplete, state that limitation
|
||||
explicitly instead of inventing details.
|
||||
|
||||
## Documentation Language Rule
|
||||
|
||||
Repository documentation and skill files stay in English:
|
||||
|
||||
- `SKILL.md`
|
||||
- `references/runner.md`
|
||||
- `README.md`
|
||||
- `docs/DESIGN.md`
|
||||
- design specs under `docs/superpowers/specs/`
|
||||
|
||||
Runtime output may use the operator's language. Do not mix Russian and English
|
||||
inside repository documentation paragraphs except for machine-readable literals,
|
||||
CLI flags, config keys, JSON keys, or quoted runtime examples.
|
||||
|
||||
## Files To Update
|
||||
|
||||
Expected implementation touches:
|
||||
|
||||
- `SKILL.md`
|
||||
- default model, sandbox, approval parsing (note: approval policy is
|
||||
expressed via `-c approval_policy=...`, not the removed `-a` flag);
|
||||
- language detection and language block;
|
||||
- runner input schema extensions;
|
||||
- runner result schema consumption (`review_quality`, `triage.*`) and the
|
||||
operation-aware lead-side dispatch table from §Runner Result Schema;
|
||||
- evaluation matrix;
|
||||
- structural operator gate;
|
||||
- structured resume prompt;
|
||||
- workspace mutation snapshots at both layers (`git status --porcelain`
|
||||
AND `sha256sum` of `/tmp/codex-{body,plan,resume-body}-*`);
|
||||
- one-line runtime hint at the start of every review about the
|
||||
`workspace-write` default and the `sandbox:read-only` opt-out
|
||||
(suppressed when an explicit `sandbox:*` override is passed);
|
||||
- final operator summary.
|
||||
|
||||
- `references/runner.md`
|
||||
- default Codex command flags (config-based approval policy via `-c`);
|
||||
- bwrap preflight before every dispatch that selects a bwrap-backed
|
||||
sandbox mode;
|
||||
- degraded-review detection;
|
||||
- bounded triage metadata;
|
||||
- emit the full runner result schema (§Runner Result Schema), including
|
||||
`review_quality` and the `triage` object;
|
||||
- stronger mandate and negative examples.
|
||||
|
||||
- `docs/DESIGN.md`
|
||||
- rationale for reviewer permissions;
|
||||
- default model update;
|
||||
- nested approval reasoning;
|
||||
- Sonnet triage compromise;
|
||||
- operator-language behavior;
|
||||
- final summary rationale;
|
||||
- version and verification log update after smoke testing.
|
||||
|
||||
- `README.md`
|
||||
- updated defaults (including config-based approval policy via `-c`,
|
||||
not `-a`);
|
||||
- sandbox/approval behavior;
|
||||
- Safety considerations section: the gitignored-file mutation vector
|
||||
(pytest → `dev.sqlite` migration), residual risk acknowledgment, the
|
||||
`sandbox:read-only` opt-out (with the explicit caveat that read-only
|
||||
blocks empirical verification — operators trade verification for
|
||||
write-protection), and the `git worktree`-based isolation pattern for
|
||||
sensitive repos;
|
||||
- Linux bwrap/AppArmor setup;
|
||||
- operator language behavior;
|
||||
- final summary behavior.
|
||||
|
||||
Do not renumber existing `docs/DESIGN.md` sections.
|
||||
|
||||
## Verification
|
||||
|
||||
Use the existing smoke protocol in `docs/DESIGN.md §7`, updated for the new
|
||||
model and flags where needed.
|
||||
|
||||
Additional manual smoke checks:
|
||||
|
||||
- `codex exec` launches with `-s workspace-write`,
|
||||
`-c approval_policy='"on-request"'`, and
|
||||
`-c approvals_reviewer='"auto_review"'` on the target Codex CLI version
|
||||
(verify against the actual installed version — `codex exec --help`
|
||||
must list `-s` and `-c`; the absence of `-a/--ask-for-approval` is
|
||||
expected on 0.132+).
|
||||
- `codex exec resume` does not receive unsupported sandbox or approval
|
||||
flags.
|
||||
- bwrap-backed sandbox failure for `workspace-write` produces an early
|
||||
diagnostic before Codex launches, instead of a fake review round.
|
||||
- `sandbox:read-only` override still works and the runner skips approval
|
||||
prompts inside the nested process (the operator can verify this by
|
||||
running a review against a known-bwrap-failing host with the override).
|
||||
- language block preserves parseable `[severity:]` and `VERDICT` literals.
|
||||
- runner result JSON contains `review_quality` and `triage` fields when
|
||||
triage runs; legacy result JSON (no `review_quality`, no `triage`) is
|
||||
accepted by the lead as `review_quality=unknown` /
|
||||
`triage.status=skipped`.
|
||||
- triage failure does not fail an otherwise valid review.
|
||||
- a synthetic `success + degraded_environmental` on `OPERATION=initial` is
|
||||
treated as terminal infrastructure failure (no prior round to recover
|
||||
from); the verbatim review is NOT shown.
|
||||
- a synthetic `success + degraded_environmental` on `OPERATION=resume`
|
||||
emits the warning, does not consume a round, and routes to the existing
|
||||
fallback chain using the prior round's severity.
|
||||
- a synthetic `success + degraded_environmental` on `OPERATION=fresh-exec`
|
||||
is treated as terminal not-verified.
|
||||
- a synthetic `success + degraded_content` (on any operation) emits the
|
||||
warning, shows Step 5 verbatim, and prompts the operator for explicit
|
||||
advancement.
|
||||
- pre/post `git status --porcelain` catches workspace-tree mutation.
|
||||
- pre/post `sha256sum` of `/tmp/codex-{body,plan,resume-body}-<REVIEW_ID>.md`
|
||||
catches reviewer mutation of `/tmp` review inputs.
|
||||
- the runtime hint about `workspace-write` and the `sandbox:read-only`
|
||||
opt-out appears exactly once per review (regardless of mode) and is
|
||||
suppressed when the operator passed an explicit `sandbox:*` override.
|
||||
- final operator summary appears at approved, max-rounds, not-verified,
|
||||
and aborted terminal states.
|
||||
|
||||
Dogfood the result with:
|
||||
|
||||
- plan review of this design;
|
||||
- code review after implementation.
|
||||
|
||||
No automated CI is required for this refresh.
|
||||
Reference in New Issue
Block a user