1082-line design note covering:
- Empirical facts about Codex CLI 0.121.0 (invocation, streams,
resume semantics, known failure modes) with copy-pasteable
verification commands.
- Claude Code harness facts (Bash truncation, cwd drift, Opus
literal-interpretation tendencies).
- 12 design decisions in a uniform format: what, where in SKILL.md,
alternatives considered, why chosen, trade-offs accepted.
- Rejected ideas (marker files, per-round naming, $(pwd), etc.) with
reasons, so future contributors don't re-propose them.
- Prior diagnostic errors from a previous agent-auditor's dump that
turned out to be wrong when verified, kept as a methodological
lesson.
- Smoke-test protocol (§7) with concrete commands and expected
outputs so any maintainer can verify the Codex contract still holds
in minutes.
- Update protocol: when and how to revise this file, with a pointer
that future Opus generations interpret instructions more literally
and SKILL.md hardening must track that.
- Mermaid flow diagram of the round-trip.
Intended audiences: future Claude sessions resuming work on the skill,
human developers, and new contributors. The file is self-contained —
does not rely on conversation history that produced the current design.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Address 10+ findings from two rounds of adversarial review of the
previous Step 4/5/7 design. Major changes:
- Use `codex exec --json` so `thread_id` can be parsed deterministically
from the first JSONL line on stdout (bypasses the ~30KB Bash-tool
truncation that could drop stderr metadata in the old flow).
- Capture REPO_ROOT via `git rev-parse --show-toplevel` at Step 2 and
substitute the absolute path literally. Pin the initial exec with
`-C "${REPO_ROOT}"` and prefix every resume with `cd '${REPO_ROOT}' &&`
because `codex exec resume` has no `-C` flag and inherits cwd from
the invoking shell.
- Drop `resume --last` from the fallback chain (cwd filtering is not
enough to distinguish our session from unrelated parallel codex runs).
- Update CODEX_SESSION_ID only on full success (exit 0, no stderr error
line, review file contains VERDICT and findings on REVISE); rotate
to the resumed session's new thread_id each round.
- Harden the "show review" gate (Step 5 "YOUR NEXT MESSAGE" instruction
and Step 6 precondition check) now that --json stdout no longer leaks
review text into the Bash tool result.
- Add strict check order for launch and resume (exit → stderr → review
file) so we never commit a broken session-id on a half-failed run.
- Replace silent fresh-exec fallback with interactive ask / headless
severity-based decision. Fresh-exec prompt rebuilds prior rounds from
conversation history.
- Bare repo / submodule / shell-hostile paths abort at Step 2 with a
clear message rather than failing silently later.
- Conditional cleanup: keep temp files on abort paths for diagnostics.
- Expand REVIEW_ID random to 8 digits.
README: update permissions (add stdout JSONL read, resume-prompt write,
narrower `cd * && ... codex exec resume *` pattern) and troubleshooting
(NOT VERIFIED outcome, bare repo, submodule).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
When the skill is available in shared ~/.agents/skills/, Codex CLI
picks it up and tries to follow its instructions — launching itself
recursively. The blockquote explains the architectural constraint
and tells Codex to review directly instead.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Fix recommended permissions: add missing Write(/tmp/codex-prompt-*),
remove overbroad rm rule (cleanup is best-effort)
- Rename claude-plan-* → codex-plan-* so all temp files share codex-* prefix
- Extract session ID via Read tool instead of grep (no extra permission needed)
- Add UUID format spec for session ID validation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add Plan Mode /tmp write limitation to SKILL.md (Step 4) and README
- Document that `codex exec resume` inherits sandbox from original session
- Remove none/low reasoning effort options (minimum is now medium)
- Add .claude to .gitignore (plan files from testing)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Зачем:
- длинные XML-промпты (60+ строк) ломали shell quoting при inline-передаче в codex exec.
- session ID терялся из-за 2>/dev/null на stderr, делая resume невозможным.
- Что:
- промпт записывается в temp-файл, передаётся через stdin: `codex exec ... - < file`.
- stderr перенаправлен в temp-файл, session ID извлекается через grep.
- resume унифицирован: тот же stdin-механизм вместо inline-аргумента.
- fallback fresh exec явно обновляет CODEX_SESSION_ID.
- cleanup дополнен новыми temp-файлами (prompt, stderr).
- Проверка:
- smoke test: plan review → 2 раунда с resume через session ID — OK.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Why:
- English makes the skill accessible to a wider audience
- Permission prompts on every git/codex call hurt UX
- What:
- Translated all SKILL.md instructions and rules to English
- Added recommended permissions section to README
- Removed literal ## from output_format to avoid Claude Code
security warning about # in quoted arguments
- Removed overly broad Bash(codex *) permission rule
- Added explicit note about codex exec scope limitations
- Verify:
- /adversarial-review produces structured output with markdown headers
- No "Newline followed by #" security warning on codex exec
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fixes from adversarial code-vs-plan review (3 rounds):
- Verdict format in prompts now matches parser (bare tokens)
- Missing verdict treated as parse failure, not approval
- README: softened backend swappability to "designed for extensibility"
- Example: replaced incorrect FK scenario with valid transaction bug
- Example: aligned fixes and round-2 summary with round-1 finding
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Baseline copy of the working codex-review SKILL.md from dotfiles
before adversarial prompt rewrite and rebranding.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>