Best for
- GEPA's economy (minibatch gate + frontier) pays off precisely when evaluations are costly and feedback is informative; otherwise the bookkeeping doesn't earn its keep.
skillberry-ai/cap-evolve/skills/algorithms/gepa/SKILL.md
Runs the real GEPA optimization loop (arXiv:2507.19457) — sample-efficient reflective Pareto search. Use when rollouts are expensive and the scorer gives informative per-task feedback, and you want the most quality per evaluation. Each iteration samples a parent from a per-instance Pareto frontier, evaluates it on a cheap minibatch of train tasks with full traces, builds a reflective dataset over the failures, asks the optimizer for one targeted component edit, re-checks the child on the same mi
Decision brief
GEPA (Agrawal et al., 2025, arXiv:2507.19457) is the highest-ceiling member of the family. Its power comes from a two-stage economy that spends cheap rollouts to decide whether a candidate is worth an expensive honest evaluation, plus reflection on traces (not scalars) and a per…
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/skillberry-ai/cap-evolve --skill "skills/algorithms/gepa"Inspect the Agent Skill "gepa" from https://github.com/skillberry-ai/cap-evolve/blob/4bb97c4e190c4795326d1834b3a5cea3cd3d499a/skills/algorithms/gepa/SKILL.md at commit 4bb97c4e190c4795326d1834b3a5cea3cd3d499a. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Requires baseline first (reads the seed's full-val result from baseline.json). Reports the frontier/pool, best candidate, accepts, merges, and metric-calls spent; test stays sealed for finalize.
1. Select a parent by sampling the per-instance Pareto frontier frequency- weighted — each non-dominated candidate's weight is how many val instances it is best at, so a specialist that uniquely tops one task is kept (seeded RNG, logged). 2. Sample a minibatch of --minibatch-siz…
GEPA's economy (minibatch gate + frontier) pays off precisely when evaluations are costly and feedback is informative; otherwise the bookkeeping doesn't earn its keep.
For a single-file / monolithic capability there is only one component; round-robin and all coincide, and the system-aware merge skips gracefully (nothing independent to recombine) rather than producing a degenerate child.
--max-metric-calls (default 0 = unlimited): PRIMARY budget — total rollouts.
Permission review
The documentation asks the agent to run terminal commands or scripts.
python scripts/check.py # behavioral, offline (mock optimizer + synthetic adapter)The documentation asks the agent to run terminal commands or scripts.
python scripts/run.py --run-dir .capevolve/run_X --project .capevolve/project \Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 84/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 36 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
GEPA (Agrawal et al., 2025, arXiv:2507.19457) is the highest-ceiling member of the
family. Its power comes from a two-stage economy that spends cheap rollouts to
decide whether a candidate is worth an expensive honest evaluation, plus reflection
on traces (not scalars) and a per-instance Pareto frontier that keeps
specialists alive. This skill is a thin wrapper over cap_evolve.gepa.gepa_loop;
all honesty-critical machinery (splits, gate, seal, stats, cache) is the engine's.
--minibatch-size (default 4) train ids.REFLECTION.md in the
optimizer workdir, plus a round-robin component focus as FOCUS.md. Invoke
the optimizer.sum(child) > sum(parent). This is the economy: a proposal that doesn't even help the
minibatch is rejected here, before any full-val cost.--merge-cadence accepts, up to --max-merges):
find two frontier dominators sharing a common ancestor both beat, recombine
component-by-component (each component from whichever descendant changed it),
minibatch-gate, then full-val + standard gate.Budget is in rollouts/metric-calls (--max-metric-calls, primary) — both
minibatch and full-val evals count — with --max-iterations as a secondary cap.
The test split is never touched; minibatch/merge evals draw from train/val only.
| Situation | Use |
|---|---|
| Rich per-task feedback + expensive rollouts; want max quality/eval | gepa |
| First run / need a yardstick baseline | hill-climb (--focus all) |
| Binary pass/fail, no diagnosis in feedback | hill-climb (reflection has little to chew on) |
| Tiny task set (frontier collapses to 1–2 points) | hill-climb |
| Want a fixed edit-budget schedule + epoch slow-update | skillopt |
| Single global-best lineage is fine and merges add no value | hill-climb / skillopt |
GEPA's economy (minibatch gate + frontier) pays off precisely when evaluations are costly and feedback is informative; otherwise the bookkeeping doesn't earn its keep.
--component-selector round_robin (default): each iteration focuses ONE
component (cycled across the parent's editable files), so every proposal is a
small, attributable change — the unit the merge later recombines.--component-selector all: list every component in FOCUS.md; the optimizer
may edit anywhere. Use for monolithic capabilities or when changes must span files.For a single-file / monolithic capability there is only one component; round-robin
and all coincide, and the system-aware merge skips gracefully (nothing independent
to recombine) rather than producing a degenerate child.
--max-metric-calls (default 0 = unlimited): PRIMARY budget — total rollouts.--max-iterations (default 50): secondary cap on propose→gate iterations.--minibatch-size (default 4): train ids per cheap local gate.--n-trials (default 1): rollouts/task on the full-val eval (raise under noise so
the significance gate is trustworthy).--component-selector (round_robin | all), --selection-strategy
(default pareto_per_instance), --max-merges (default 2), --merge-cadence
(default 3).--gate-mode / --k-se: the val acceptance bar (paired significance by default).--no-regression: reject a child that breaks any previously-passing val task.--seed: seeds the parent-sampling + minibatch RNG (logged for reproducibility).--resume: reconstruct the pool/lineage/frontier from the run dir (a
gepa_state.json checkpoint + each accepted candidate's rollouts) and continue the
Pareto search where it stopped, instead of restarting from the seed. Preserved spend
keeps the budget honest; the parent-sampling RNG stream restarts (selection is
stochastic by design, so the resumed run is not byte-identical).python scripts/check.py # behavioral, offline (mock optimizer + synthetic adapter)
python scripts/run.py --run-dir .capevolve/run_X --project .capevolve/project \
--optimizer 'python .../run-optimizer/scripts/run.py --name mock --workdir {workdir} --prompt {prompt}' \
--max-metric-calls 400 --minibatch-size 4 --component-selector round_robin
Requires baseline first (reads the seed's full-val result from baseline.json).
Reports the frontier/pool, best candidate, accepts, merges, and metric-calls spent;
test stays sealed for finalize.
When orchestration_mode: agent, drive gepa yourself: maintain the candidate pool/Pareto frontier; each round pick a parent (per gepa's selection), reflect on its val feedback to propose an edit, evaluate on val via cap-evolve, gate Δ>k·SE, accept→snapshot & add to the frontier / reject→drop. Metric-calls is the primary budget. Log rounds to the run dir; between rounds verify rollouts+results landed so the dashboard reflects the frontier. Re-read stop_condition; stop on it/budget. Seal once with cap-evolve finalize, then report.
references/concepts.md — the GEPA economy, reflective dataset / actionable side
information, per-instance frequency-weighted frontier, system-aware merge, the
metric-call budget, and the relation to the hill-climb / skillopt siblings.
Cites arXiv:2507.19457.