Best for
- Task has a measurable numeric metric (size, time, count, score, coverage)
- A verify command exists that outputs the metric deterministically
- "Better" means the number going consistently in one direction
xoai/sage/skills/autoresearch/SKILL.md
Autonomous iteration toward a measurable outcome. Use when the user wants to optimize a numeric metric through repeated modify-verify cycles — reduce bundle size, increase test coverage, improve query time, lower readability score. Not for exploratory research, subjective judgment, or tasks without a verification command.
Decision brief
Autonomous iteration toward a measurable outcome. The agent modifies code, commits, runs a verify command, keeps improvements, reverts regressions — repeating until a target is hit, a budget is exhausted, or the user interrupts.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/xoai/sage --skill "skills/autoresearch"Inspect the Agent Skill "autoresearch" from https://github.com/xoai/sage/blob/f7cc487b393474030cef15d50efdbb195612b756/skills/autoresearch/SKILL.md at commit f7cc487b393474030cef15d50efdbb195612b756. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Each iteration follows 8 phases. Read references/loop-protocol.md for per-phase detail.
Task has a measurable numeric metric (size, time, count, score, coverage)
Subjective goals ("make the UI prettier")
Before the loop can start, capture these (skip if already provided):
An optional Python runtime handles the deterministic phases (COMMIT, VERIFY, DECIDE, LOG, REPEAT); the agent handles the creative phases (REVIEW, IDEATE, MODIFY). The runtime was extracted from core into its own package in Phase 3 (like sage-memory) — install it with sage add xo…
Permission review
The documentation asks the agent to run terminal commands or scripts.
| 5 | VERIFY | runtime | Run verify command with wall-clock budget |The documentation asks the agent to run terminal commands or scripts.
python3 -m autoresearch run --brief .sage/work/<slug>/brief.md --project .Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 84/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 25 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Autonomous iteration toward a measurable outcome. The agent modifies code, commits, runs a verify command, keeps improvements, reverts regressions — repeating until a target is hit, a budget is exhausted, or the user interrupts.
Core principles (from Karpathy's autoresearch pattern):
Before the loop can start, capture these (skip if already provided):
| Field | Required | Example |
|---|---|---|
| Goal | Yes | "Reduce bundle below 200KB" |
| Metric name | Yes | bundle_kb |
| Direction | Yes | lower or higher |
| Target | Optional | 200 |
| Verify command | Yes | pnpm build && measure.sh |
| Writable scope | Recommended | src/**/*.ts |
| Frozen scope | Recommended | package.json, *.lock |
| Per-run budget | Yes (default 120s) | 120 seconds |
| Max iterations | Optional | 100 |
| Termination | Auto | target if target given, else interrupt |
Present as a brief for user approval:
Sage: Autoresearch session configured.
Goal: [goal statement]
Metric: [name] ([direction]), target: [target or "none — runs until interrupted"]
Verify: [command]
Scope: writable [globs], frozen [globs]
Budget: [seconds]s per run, [max iterations or "unlimited"]
[A] Start — begin autonomous iteration
[R] Revise — change configuration
Each iteration follows 8 phases. Read references/loop-protocol.md
for per-phase detail.
| # | Phase | Actor | What happens |
|---|---|---|---|
| 1 | REVIEW | agent | Read current state, recent history (last 20 iterations from JSONL) |
| 2 | IDEATE | agent | Propose ONE change, ≤1 sentence. If stuck, load references/stuck-recovery.md |
| 3 | MODIFY | agent | Make the change. Stay within writable scope. |
| 4 | COMMIT | runtime | git add -A && git commit on autoresearch/<slug> branch |
| 5 | VERIFY | runtime | Run verify command with wall-clock budget |
| 6 | DECIDE | runtime | Parse METRIC, compare to best → keep / discard / crash |
| 7 | LOG | runtime+agent | Append JSONL, rebuild TSV, agent updates living doc |
| 8 | REPEAT | runtime | Check termination → loop or exit |
Decision rules (Phase 6):
crash, reset to HEADcrash, resetcrash, resetkeep, advance branchdiscard, resetAn optional Python runtime handles the deterministic phases (COMMIT, VERIFY,
DECIDE, LOG, REPEAT); the agent handles the creative phases (REVIEW, IDEATE,
MODIFY). The runtime was extracted from core into its own package in Phase 3
(like sage-memory) — install it with sage add xoai/sage-autoresearch.
Running the runtime — probe first, and degrade LOUDLY if it is absent (announce + log to decisions.md; never silently skip):
if python3 -c 'import autoresearch' 2>/dev/null; then
python3 -m autoresearch run --brief .sage/work/<slug>/brief.md --project .
else
echo "Sage: autoresearch runtime not installed — running in degraded (manual)"
echo "mode. For the deterministic runtime: sage add xoai/sage-autoresearch"
fi
Harness contract: The verify command must print METRIC name=number
to stdout. See references/harness-conventions.md.
All state lives in .sage/work/<YYYYMMDD-slug>/:
| File | Role |
|---|---|
brief.md | Configuration (goal, metric, scope, budget) |
autoresearch.md | Living doc — ideas tried, wins, dead ends |
autoresearch.jsonl | Structured log (one line per iteration) |
results.tsv | Human-readable view (derived from JSONL) |
runs/NNNN-*.log | Per-iteration stdout+stderr |
.autoresearch-state.json | Crash recovery state (not committed) |
On resume (new session, context reset, platform switch):
autoresearch.md for high-level contextautoresearch.jsonl for recent historygit log on the branchSee references/session-continuity.md for full protocol.
Session end: Store a structured summary in sage-memory:
Session start: Search sage-memory for priors on this repo + metric. Inject into IDEATE as "known-good starting points" and "known dead ends."
| Gate | When | Check |
|---|---|---|
| scope | After MODIFY | Changed files ⊆ writable, frozen untouched |
| pre-verify | After COMMIT | git status is clean |
| metric-parseable | After VERIFY | At least one METRIC line in stdout |
| budget | During VERIFY | Wall-clock ≤ per_run_seconds |
Gates are enforced by the runtime, not by prose. The agent cannot bypass them.
references/loop-protocol.md — per-phase inputs, outputs, failure modesreferences/metric-design.md — what makes a good metricreferences/harness-conventions.md — METRIC line contractreferences/stuck-recovery.md — escape local minimareferences/crash-handling.md — retry vs skip decision treereferences/session-continuity.md — resume protocolAlternatives
github/awesome-copilot
Autonomous iterative experimentation loop for any programming task. Guides the user through defining goals, measurable metrics, and scope constraints, then runs an autonomous loop of code changes, testing, measuring, and keeping/discarding results. Inspired by Karpathy's autoresearch. USE FOR: autonomous improvement, iterative optimization, experiment loop, auto research, performance tuning, automated experimentation, hill climbing, try things automatically, optimize code, run experiments, auton
Orchestra-Research/AI-Research-SKILLs
Orchestrates end-to-end autonomous AI research projects using a two-loop architecture. The inner loop runs rapid experiment iterations with clear optimization targets. The outer loop synthesizes results, identifies patterns, and steers research direction. Routes to domain-specific skills for execution, supports continuous agent operation via Claude Code /loop and OpenClaw heartbeat, and produces research presentations and papers. Use when starting a research project, running autonomous experimen
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist