Source profileQuality 86/100

notque/vexjoy-agent/skills/meta/codex/SKILL.md

codex

Run benchmark-selected GPT-5.6 work through the Codex CLI.

Source repository stars
413
Declared platforms
1
Static risk flags
1
Last source update
2026-07-25
Source checked
2026-08-04

Decision brief

What it does—and where it fits

Run a benchmark-selected GPT-5.6 task through the Codex CLI (codex exec) and return the result. This is the OpenAI execution lane — the general-purpose lane for work the model-selection policy sends to GPT-5.6, and the canonical owner of general codex exec mechanics — when the C…

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexDeclaredSource recordInstall path and trigger
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/notque/vexjoy-agent --skill "skills/meta/codex"
    Safe inspection promptEditorial

    Inspect the Agent Skill "codex" from https://github.com/notque/vexjoy-agent/blob/b19dacd072f5befd29b525b25dbecc7a1cd86d92/skills/meta/codex/SKILL.md at commit b19dacd072f5befd29b525b25dbecc7a1cd86d92. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Phase 1: DECIDE — does this task belong on GPT-5.6?

      Policy mirror — canonical copy: /do SKILL.md, Model Selection (edit there first, then here). Rankings, higher = better; cost = avg USD per task, written as a plain number (slash-command templating corrupts dollar-digit sequences in injected skill bodies), what the owner actually…

      Policy mirror — canonical copy: /do SKILL.md, Model Selection (edit there first, then here). Rankings, higher = better; cost = avg USD per task, written as a plain number (slash-command templating corrupts dollar-digit…Run deterministic work as scripts, not through Codex. The /do model policy selects the lane and passes model plus effort. Legacy GPT-5.5, all Luna choices, and the other non-default GPT-5.6 settings are manual-only; do…These are defaults, not limits. Standing permission to escalate when output misses the bar applies within the policy; max still needs an explicit override. For anything that ships, intelligence taste cost; cost is a tie…
    2. 02

      Phase 2: WRAP — how GPT-5.6 runs from this harness

      Wrapper symmetry: the wrapper is needed for whichever model family is NOT the current harness.

      Under Claude Code (current default): GPT-5.6 runs through a wrapper — either the dispatched agent runs codex exec via Bash with a self-contained prompt, or a thin Claude wrapper agent (model: "sonnet", low effort) write…Under the Codex harness: Claude models require the wrapper instead.Claude models under Claude Code need no wrapper — just the Agent/Workflow model parameter.
    3. 03

      Phase 3: PROMPT — write a self-contained prompt

      Codex runs in its own process with no conversation history. The prompt must carry everything:

      Context — one short paragraph: what the repo/data is, what state matters.Task — the concrete operation, with file paths relative to the working directory. Let codex read files itself; embedding large content wastes tokens and loses formatting.Output format — the exact structure to return (table, JSON, diff), so the wrapper can consume it without a second pass.
    4. 04

      Phase 4: RUN

      Pass the policy-selected model and effort explicitly. Do not rely on a local default that can silently select a deprecated model.

      Pass the policy-selected model and effort explicitly. Do not rely on a local default that can silently select a deprecated model.Investigation / data analysis (default for anything that only reads):Set CODEXMODEL and CODEXEFFORT from the /do selection before invoking the CLI; do not substitute a local default.
    5. 05

      Error handling

      Cause: Codex CLI not installed on this host. Solution: fall back to the policy's Claude pick (model: "sonnet" for mechanical work) and report which lane ran; install via the owner's codex setup when authorized.

      Cause: Codex CLI not installed on this host. Solution: fall back to the policy's Claude pick (model: "sonnet" for mechanical work) and report which lane ran; install via the owner's codex setup when authorized.Cause: the bwrap sandbox fails in some containerized/VM environments. Solution: for read-only work, retry without -s read-only only if the environment already provides external sandboxing (Claude Code does); -s read-onl…Cause: prompt exceeded length limits or codex wrote to stdout only. Solution: shorten the prompt (point codex at files instead of embedding content); capture stdout as fallback.

    Permission review

    Static risk signals and limitations

    Reads files

    low · line 71

    The documentation asks the agent to read local files, directories, or repositories.

    codex exec -m "$CODEX_MODEL" -c "model_reasoning_effort=\"$CODEX_EFFORT\"" -s read-only --skip-git-repo-check -o "$TMPFILE" "$(cat <<'PROMPT'

    Reads files

    low · line 80

    The documentation asks the agent to read local files, directories, or repositories.

    *Write tasks (clear-spec implementation, migrations):** drop `-s read-only`; run from the target repo's working directory; review the diff (`git status --short`, `git diff`) before committing anything.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score86/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars413SourceRepository attention, not individual Skill quality
    Compatibility1 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    notque/vexjoy-agent
    Skill path
    skills/meta/codex/SKILL.md
    Commit
    b19dacd072f5befd29b525b25dbecc7a1cd86d92
    License
    MIT
    Collected
    2026-08-04
    Default branch
    main
    View the original SKILL.md

    Codex — the GPT-5.6 Execution Lane

    Run a benchmark-selected GPT-5.6 task through the Codex CLI (codex exec) and return the result. This is the OpenAI execution lane — the general-purpose lane for work the model-selection policy sends to GPT-5.6, and the canonical owner of general codex exec mechanics — when the CLI changes, update here first. GPT selections are reachable only through this CLI; the Agent tool's model parameter covers Claude models only.

    Under Claude Code, this skill runs only on explicit invocation or cross-provider escalation, never as the automatic default. The harness-native model lane under Claude Code is the Anthropic lane (Opus 5). This skill is a deliberate cross-provider tool — codex review as a second-opinion, codex exec for a GPT-specific constraint — not a routing default.

    Two flows keep their own specialized codex integration — route to them instead of re-implementing here:

    Existing flowOwnsWhere
    PR / code review via codexcodex exec review, finding triage, report synthesisskills/process/pr-workflow/references/codex-review.md
    Sprite/image generation backendcodex image backend selection and invocationskills/game/game-sprite-pipeline/references/backend-chain.md

    Phase 1: DECIDE — does this task belong on GPT-5.6?

    Policy mirror — canonical copy: /do SKILL.md, Model Selection (edit there first, then here). Rankings, higher = better; cost = avg USD per task, written as a plain number (slash-command templating corrupts dollar-digit sequences in injected skill bodies), what the owner actually pays.

    Task classModel / effortDeepSWE Pass@1 / cost / output tokens / steps
    Low-risk assistancegpt-5.6-terra / high54 / 1.13 / 22k / 34
    Standard implementationgpt-5.6-sol / high69 / 3.47 / 28k / 37
    High-risk implementation or reviewgpt-5.6-sol / xhigh71 / 4.70 / 41k / 44
    Exceptional explicit escalationgpt-5.6-sol / max73 / 8.39 / 60k / 61

    Run deterministic work as scripts, not through Codex. The /do model policy selects the lane and passes model plus effort. Legacy GPT-5.5, all Luna choices, and the other non-default GPT-5.6 settings are manual-only; do not substitute them automatically. Luna max, for example, saves 0.44 USD versus Sol high but consumes 45k more output tokens and 65 more steps for two fewer Pass@1 points. Consult the canonical table in /do SKILL.md.

    These are defaults, not limits. Standing permission to escalate when output misses the bar applies within the policy; max still needs an explicit override. For anything that ships, intelligence > taste > cost; cost is a tie-breaker only.

    Gate: task has a GPT-5.6 policy selection. Otherwise route to scripts or the policy's Claude pick and stop here.

    Phase 2: WRAP — how GPT-5.6 runs from this harness

    Wrapper symmetry: the wrapper is needed for whichever model family is NOT the current harness.

    • Under Claude Code (current default): GPT-5.6 runs through a wrapper — either the dispatched agent runs codex exec via Bash with a self-contained prompt, or a thin Claude wrapper agent (model: "sonnet", low effort) writes the self-contained codex prompt, runs it, and returns the result.
    • Under the Codex harness: Claude models require the wrapper instead.
    • Claude models under Claude Code need no wrapper — just the Agent/Workflow model parameter.

    Pick the direct-Bash form when the calling agent already holds the task context; pick the thin wrapper agent for fan-out (one wrapper per data source) so the orchestrator stays lean.

    Availability check first: command -v codex — when absent, fall back to the policy's Claude pick (model: "sonnet" for mechanical work) and tell the user in one line which lane ran.

    Phase 3: PROMPT — write a self-contained prompt

    Codex runs in its own process with no conversation history. The prompt must carry everything:

    1. Context — one short paragraph: what the repo/data is, what state matters.
    2. Task — the concrete operation, with file paths relative to the working directory. Let codex read files itself; embedding large content wastes tokens and loses formatting.
    3. Output format — the exact structure to return (table, JSON, diff), so the wrapper can consume it without a second pass.

    Prompt hygiene (hard rule): codex prompts leave the machine. Send only public content — secrets, credentials, and private component names (anything sourced from INDEX.local.json or other local-only inventories) stay out. Run the deterministic scan on the prompt text before executing:

    printf '%s' "$PROMPT" | rg -n "Bearer|Authorization|token|secret|api[_-]?key|password|PRIVATE KEY" && echo "HYGIENE VIOLATION"
    

    On a hit or a private component name: scrub the flagged content when the task survives without it; otherwise reroute the task to a Claude model. A bare refusal is not an outcome.

    Phase 4: RUN

    Pass the policy-selected model and effort explicitly. Do not rely on a local default that can silently select a deprecated model.

    Investigation / data analysis (default for anything that only reads):

    Set CODEX_MODEL and CODEX_EFFORT from the /do selection before invoking the CLI; do not substitute a local default.

    TMPFILE=$(mktemp)
    codex exec -m "$CODEX_MODEL" -c "model_reasoning_effort=\"$CODEX_EFFORT\"" -s read-only --skip-git-repo-check -o "$TMPFILE" "$(cat <<'PROMPT'
    [self-contained prompt]
    PROMPT
    )"
    cat "$TMPFILE"
    

    -s read-only sandboxes the run to reads — verified working on this host. Use it for every investigation or analysis prompt not covered by an existing codex flow, because a read-only task never needs write access and the sandbox makes that deterministic.

    Write tasks (clear-spec implementation, migrations): drop -s read-only; run from the target repo's working directory; review the diff (git status --short, git diff) before committing anything.

    Reviews: use codex exec review via the pr-workflow codex-review flow (table above), not a hand-rolled prompt.

    Gate: exit code 0 AND output matches the requested format. Non-zero exit: report stderr and stop — codex failures are auth/API/prompt-length issues that a blind retry won't fix. Verify the output against a deterministic check where one exists (counts, file lists, test runs) before passing it upstream — GPT-5.6 output is evidence, not verdict. State the model and effort in the result so the caller can apply the escalation rule.

    Error handling

    codex: command not found

    Cause: Codex CLI not installed on this host. Solution: fall back to the policy's Claude pick (model: "sonnet" for mechanical work) and report which lane ran; install via the owner's codex setup when authorized.

    Sandbox error mentioning bwrap / Failed RTM_NEWADDR

    Cause: the bwrap sandbox fails in some containerized/VM environments. Solution: for read-only work, retry without -s read-only only if the environment already provides external sandboxing (Claude Code does); -s read-only and --dangerously-bypass-approvals-and-sandbox are mutually exclusive — use one.

    Output missing or truncated in -o file

    Cause: prompt exceeded length limits or codex wrote to stdout only. Solution: shorten the prompt (point codex at files instead of embedding content); capture stdout as fallback.

    References

    • /do SKILL.md, Model Selection — canonical policy table and routing decision rules
    • skills/process/pr-workflow/references/codex-review.md — review-specific codex flow
    • skills/game/game-sprite-pipeline/references/backend-chain.md — codex image backend

    Alternatives

    Compare before choosing