Best for
- "How do I run Codex headless?"
- "Can I automate Codex workflows?"
- "How to use Codex in CI/CD?"
sunholo-data/ailang/.agents/skills/headless-runner/SKILL.md
Run Codex in headless/programmatic mode for automation, CI/CD, and agent workflows. Use when user asks about headless mode, programmatic execution, scripting Codex, or automating Codex workflows.
Decision brief
Run Codex programmatically from scripts, CI/CD pipelines, and autonomous agent workflows. Headless mode automatically loads all project configuration (.Codex/ directory), giving you full access to skills, agents, hooks, and commands.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Declared | Source record | Install path and trigger |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/sunholo-data/ailang --skill ".agents/skills/headless-runner"Inspect the Agent Skill "headless-runner" from https://github.com/sunholo-data/ailang/blob/9944e264e3b9043881978731dccd258f561082a3/.agents/skills/headless-runner/SKILL.md at commit 9944e264e3b9043881978731dccd258f561082a3. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Review the “Quick Start” section in the pinned source before continuing.
Use for: Simple, one-off tasks
Codex -p "Step 1" --output-format json step1.json ARTIFACT1=$(jq -r '.artifact' step1.json)
Codex -p "Step 2 using $ARTIFACT1" --output-format json step2.json ARTIFACT2=$(jq -r '.artifact' step2.json)
Codex -p "Step 3 using $ARTIFACT2" bash !/bin/bash
Permission review
The documentation asks the agent to run terminal commands or scripts.
Run headless command with automatic retry on failure.Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 93/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 33 | Source | Repository attention, not individual Skill quality |
| Compatibility | 1 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Run Codex programmatically from scripts, CI/CD pipelines, and autonomous agent workflows. Headless mode automatically loads all project configuration (.Codex/ directory), giving you full access to skills, agents, hooks, and commands.
Most common usage:
# Basic headless invocation (from project directory)
Codex -p "Your prompt here"
# With JSON output for programmatic parsing
Codex -p "Run eval baseline for v0.3.14" --output-format json
# Control tool access
Codex -p "Analyze failures" --allowedTools "Bash,Read,Grep"
# Multi-turn conversation
Codex -p "Start task" --output-format json > result.json
SESSION_ID=$(jq -r '.session_id' result.json)
Codex --resume $SESSION_ID "Continue with next step"
What gets loaded automatically:
.Codex/settings.json and .Codex/settings.local.json.Codex/agents/ (all project agents).Codex/skills/ (all project skills).Codex/commands/Invoke this skill when user asks about:
# Text output (default)
Codex -p "Prompt here"
# JSON output with metadata
Codex -p "Prompt here" --output-format json
# Returns: {session_id, result, cost, duration, ...}
# Streaming JSON (for long-running tasks)
Codex -p "Prompt here" --output-format stream-json
# Allow specific tools
Codex -p "Task" --allowedTools "Bash,Read,Write"
# Allow all tools (use with caution)
Codex -p "Task" --allowedTools "*"
# Permission mode for edits
Codex -p "Task" --permission-mode acceptEdits
# Resume specific session
Codex --resume SESSION_ID "Continue task"
# Continue most recent session
Codex --continue "Next instruction"
# Extract session ID from JSON output
SESSION_ID=$(Codex -p "Start" --output-format json | jq -r '.session_id')
Codex --resume $SESSION_ID "Continue"
# .github/workflows/eval-baseline.yml
- name: Run eval baseline
run: |
Codex -p "Use eval-orchestrator agent to run baseline for ${{ github.ref_name }}" \
--output-format json \
--allowedTools "Bash,Read,Write" \
> eval_result.json
- name: Check for failures
run: |
FAILURES=$(jq -r '.failures' eval_result.json)
if [ "$FAILURES" -gt 0 ]; then
echo "::error::Eval baseline has $FAILURES failures"
exit 1
fi
#!/bin/bash
# cron_daily_check.sh - Run via cron daily
cd /path/to/project
# Check agent inbox
Codex -p "Use agent-inbox skill to check for unread messages" \
--output-format json > inbox.json
# If messages exist, notify
UNREAD=$(jq -r '.unreadCount' inbox.json)
if [ "$UNREAD" -gt 0 ]; then
echo "Found $UNREAD unread agent messages"
# Send notification, create issue, etc.
fi
#!/bin/bash
# autonomous_sprint_cycle.sh
# Agent A: Create design doc
Codex -p "Use design-doc-creator to document feature X" \
--output-format json > design.json
DESIGN_DOC=$(jq -r '.artifactPath' design.json)
# Agent B: Plan sprint from design
Codex -p "Use sprint-planner to create plan from $DESIGN_DOC" \
--output-format json > plan.json
PLAN_FILE=$(jq -r '.planPath' plan.json)
# Agent C: Execute sprint
Codex -p "Use sprint-executor to execute $PLAN_FILE" \
--output-format json > execution.json
#!/bin/bash
# test_agent_quality.sh
# Run eval with specific model
Codex -p "Use eval-orchestrator: run suite with gpt5-mini only" \
--output-format json > results.json
# Parse results
SUCCESS_RATE=$(jq -r '.successRate' results.json)
# Assert quality threshold
if (( $(echo "$SUCCESS_RATE < 0.75" | bc -l) )); then
echo "Agent quality below threshold: $SUCCESS_RATE"
exit 1
fi
scripts/test_headless.shTest headless mode works correctly with project configuration.
Usage:
.Codex/skills/headless-runner/scripts/test_headless.sh
What it tests:
Codex command is availablescripts/run_with_retry.sh <prompt> [max_retries]Run headless command with automatic retry on failure.
Usage:
.Codex/skills/headless-runner/scripts/run_with_retry.sh "Run eval baseline" 3
Features:
Use for: Simple, one-off tasks
Codex -p "Generate changelog from git log since v0.3.13"
Use for: Multi-step workflows where each step depends on previous
#!/bin/bash
set -euo pipefail
# Step 1
Codex -p "Step 1" --output-format json > step1.json
ARTIFACT1=$(jq -r '.artifact' step1.json)
# Step 2 (uses Step 1 output)
Codex -p "Step 2 using $ARTIFACT1" --output-format json > step2.json
ARTIFACT2=$(jq -r '.artifact' step2.json)
# Step 3
Codex -p "Step 3 using $ARTIFACT2"
Use for: Multi-turn tasks that need context
#!/bin/bash
# Start conversation
RESULT=$(Codex -p "Analyze codebase for tech debt" --output-format json)
SESSION_ID=$(echo "$RESULT" | jq -r '.session_id')
# Continue conversation with context
Codex --resume $SESSION_ID "Focus on files over 800 lines"
Codex --resume $SESSION_ID "Generate refactoring plan"
Codex --resume $SESSION_ID "Estimate effort for top 3 items"
Use for: Independent tasks that can run concurrently
#!/bin/bash
# Start multiple tasks in parallel
Codex -p "Task A" --output-format json > taskA.json &
PID_A=$!
Codex -p "Task B" --output-format json > taskB.json &
PID_B=$!
Codex -p "Task C" --output-format json > taskC.json &
PID_C=$!
# Wait for all to complete
wait $PID_A $PID_B $PID_C
# Aggregate results
jq -s '{taskA: .[0], taskB: .[1], taskC: .[2]}' taskA.json taskB.json taskC.json
Output formats:
--output-format text (default) - Human-readable--output-format json - For automation (includes session_id, status, cost)--output-format stream-json - For real-time progressConfiguration:
.Codex/ configError handling:
if ! Codex -p "..." ; then ...echo "$RESULT" | jq -e '.status == "success"'run_with_retry.sh script)Tool permissions:
--allowedTools "Read,Grep,Glob"--allowedTools "Bash,Read"--allowedTools "Bash,Read,Write,Edit"For complete details, see CLI Reference and Troubleshooting.
Build autonomous agents using headless Codex + AILANG messaging:
For complete autonomous agent patterns (task claiming, handoffs, error handling), see:
resources/agent_workflows.md - Autonomous agent patterns with messagingQuick example:
# Agent checks inbox for tasks
MESSAGES=$(ailang agent inbox --unread-only my-agent)
MESSAGE_ID=$(echo "$MESSAGES" | grep "ID:" | head -1 | awk '{print $2}')
# Claim task
ailang agent ack $MESSAGE_ID
# Process with headless Codex
RESULT=$(Codex -p "Process task from inbox" --output-format json)
# On success: keep ack, send result
if [ "$(echo "$RESULT" | jq -r '.status')" = "success" ]; then
ailang agent send --to-user --from "my-agent" '{"status": "complete"}'
else
# On failure: return to queue
ailang agent unack $MESSAGE_ID
fi
See resources/agent_workflows.md for autonomous agent patterns with AILANG messaging system.
See resources/cli_reference.md for complete CLI flag documentation.
See resources/examples.md for comprehensive workflow examples.
See resources/troubleshooting.md for common issues and solutions.
This skill loads information progressively:
scripts/ directoryresources/ (detailed CLI reference, examples, troubleshooting)--output-format json → .cost field--resumeFrequently asked questions
Run Codex programmatically from scripts, CI/CD pipelines, and autonomous agent workflows. Headless mode automatically loads all project configuration (.Codex/ directory), giving you full access to skills, agents, hooks, and commands.
The source record exposes this install command: npx skills add https://github.com/sunholo-data/ailang --skill ".agents/skills/headless-runner". Inspect the command and pinned source before running it.
The pinned source record declares support for: codex.
Static rules flagged exec-script in the source; the page lists the matching lines and excerpts.
Alternatives
sunholo-data/ailang
Run Claude Code in headless/programmatic mode for automation, CI/CD, and agent workflows. Use when user asks about headless mode, programmatic execution, scripting Claude, or automating Claude workflows.
PramodDutta/qaskills
Gate RAG pipelines in CI with versioned golden eval sets, per-metric thresholds, baseline drift detection, and a build that fails when retrieval or answer quality regresses.
upex-galaxy/agentic-qa-boilerplate
Execute regression test suites via CI/CD, analyze results, classify failures, and produce GO/NO-GO release decisions. Use when running regression, smoke, or sanity suites through GitHub Actions, monitoring workflow runs, downloading Allure or Playwright artifacts, classifying failures (REGRESSION vs FLAKY vs KNOWN vs ENVIRONMENT vs NEW TEST), computing pass-rate and trend metrics, deciding release readiness, generating executive quality reports, or creating regression issues. Triggers on: run re
nyldn/claude-octopus
Multi-AI validation, scoring, and review using available external providers (Double Diamond Deliver phase)