skillberry-ai/cap-evolve/skills/phases/intake/SKILL.md
intake
Phase 1 of the pipeline — collect inputs and scaffold the run. Use at the very start of any optimization. Interviews the user to decide what capability to optimize, which runner/optimizer/algorithm to use, and where the data is; scaffolds .capevolve/project/ (adapter stub, capevolve.yaml, PROJECT.md); and for every NEEDED input that is missing, asks the user (quoting path, how to retrieve it, alternatives) rather than fabricating it.
- Source repository stars
- 36
- Declared platforms
- 0
- Static risk flags
- 2
- Last source update
- 2026-08-04
- Source checked
- 2026-08-04
Decision brief
What it does—and where it fits
The first phase. Its job is to turn a vague wish ("make this agent better at X") into a concrete, runnable project: a filled capevolve.yaml, an adapter ready to implement, and every NEEDED input resolved before any budget is spent. Intake is cheap; a botched intake is not — an u…
Not for
- Tasks that require unconfirmed production actions or broad system permissions.
- Environments where the pinned source and install steps cannot be inspected.
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/skillberry-ai/cap-evolve --skill "skills/phases/intake"Inspect the Agent Skill "intake" from https://github.com/skillberry-ai/cap-evolve/blob/4bb97c4e190c4795326d1834b3a5cea3cd3d499a/skills/phases/intake/SKILL.md at commit 4bb97c4e190c4795326d1834b3a5cea3cd3d499a. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
Step 0 — Inspect before asking
Before any question, inspect the target so you can PROPOSE defaults instead of asking blind: 1. Look at the repo/benchmark/agent: entrypoint, how one eval runs, where traces/scores live. 2. Detect candidate metrics (what the scorer emits), a natural train/val/test split, and cos…
Look at the repo/benchmark/agent: entrypoint, how one eval runs, where traces/scores live.Detect candidate metrics (what the scorer emits), a natural train/val/test split, and cost caps.Run gh auth status to know whether GitHub is available. - 02
How to run
The script is purely mechanical: it copies the template and prints the next steps. The judgment — interviewing, choosing components, and running the ask-if-missing loop — is yours, driven by this SKILL.md and inputs/INPUTS.md.
The script is purely mechanical: it copies the template and prints the next steps. The judgment — interviewing, choosing components, and running the ask-if-missing loop — is yours, driven by this SKILL.md and inputs/INP…Then implement adapters/adapter.py, fill capevolve.yaml, and proceed to implement-and-check. Together, intake → implement-and-check is the full integration: scaffold → implement the 3 required adapter methods → cap-evol…Worked example (onboard a new benchmark from a prompt): see examples/ for an end-to-end onboarding. The intake/onboarding step installs the benchmark (clones + installs it) and the optimizer agent wires the adapter from… - 03
Inputs / outputs (manifest tokens)
Downstream, implement-and-check consumes project; baseline consumes project + tasks. If intake under-delivers either token, the hard gate in implement-and-check fails loudly rather than silently optimizing against a stub.
needs: (nothing) — intake is the pipeline entry point.provides: project (the scaffolded .capevolve/project/) and tasks (the- needs: (nothing) — intake is the pipeline entry point. - provides: project (the scaffolded .capevolve/project/) and tasks (the evaluation dataset, resolved either to a path or to the adapter's tasks()). - 04
Metrics
Which metrics should the dashboard show? (detected: ) — multiple choice + free text → metricsdisplay.
Which metrics should the dashboard show? (detected: ) — multiple choice + free text → metricsdisplay.Which ONE gates accept/reject? — single choice → metricprimary. (This is the only metric the gate uses.)For each shown metric, is higher or lower better? → metricdirections (parallel to metricsdisplay). - 05
GitHub integration
gh auth status = authed? Offer: mirror the algorithm's work items as issues + ship winner as PR (Closes n) → githubintegration: true; else offer gh auth login or skip → false. WHAT gets mirrored is algorithm-specific (t…
gh auth status = authed? Offer: mirror the algorithm's work items as issues + ship winner as PR (Closes n) → githubintegration: true; else offer gh auth login or skip → false. WHAT gets mirrored is algorithm-specific (t…- gh auth status = authed? Offer: mirror the algorithm's work items as issues + ship winner as PR (Closes n) → githubintegration: true; else offer gh auth login or skip → false. WHAT gets mirrored is algorithm-specific…
Permission review
Static risk signals and limitations
Reads files
The documentation asks the agent to read local files, directories, or repositories.
Both path forms matter, each for its own reader: YOU read it in this repo atRuns scripts
The documentation asks the agent to run terminal commands or scripts.
python scripts/run.py --base .capevolve # scaffold .capevolve/projectEvidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 90/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 36 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- skillberry-ai/cap-evolve
- Skill path
- skills/phases/intake/SKILL.md
- Commit
- 4bb97c4e190c4795326d1834b3a5cea3cd3d499a
- License
- Apache-2.0
- Collected
- 2026-08-04
- Default branch
- main
View the original SKILL.md
intake — collect inputs, scaffold the project
The first phase. Its job is to turn a vague wish ("make this agent better at X")
into a concrete, runnable project: a filled capevolve.yaml, an adapter ready to
implement, and every NEEDED input resolved before any budget is spent. Intake
is cheap; a botched intake is not — an unresolved input discovered three phases
later means a wasted optimization run and a meaningless number.
Inputs / outputs (manifest tokens)
- needs: (nothing) — intake is the pipeline entry point.
- provides:
project(the scaffolded.capevolve/project/) andtasks(the evaluation dataset, resolved either to a path or to the adapter'stasks()).
Downstream, implement-and-check consumes project; baseline consumes
project + tasks. If intake under-delivers either token, the hard gate in
implement-and-check fails loudly rather than silently optimizing against a stub.
Step 0 — Inspect before asking
Before any question, inspect the target so you can PROPOSE defaults instead of asking blind:
- Look at the repo/benchmark/agent: entrypoint, how one eval runs, where traces/scores live.
- Detect candidate metrics (what the scorer emits), a natural train/val/test split, and cost caps.
- Run
gh auth statusto know whether GitHub is available.
Then ask the FEWEST questions — present detected metrics/splits/caps as multiple-choice
defaults with a free-text escape, keeping the ask-user-if-missing discipline for NEEDED
inputs. The metric / GitHub / stop-condition questions below feed directly into the
capevolve.yaml spec keys, so ask them here (after inspecting, before scaffolding).
Metrics
- Which metrics should the dashboard show? (detected:
<list>) — multiple choice + free text →metrics_display. - Which ONE gates accept/reject? — single choice →
metric_primary. (This is the only metric the gate uses.) - For each shown metric, is higher or lower better? →
metric_directions(parallel tometrics_display).
GitHub integration
gh auth status= authed? Offer: mirror the algorithm's work items as issues + ship winner as PR (Closes #n) →github_integration: true; else offergh auth loginor skip →false. WHAT gets mirrored is algorithm-specific (thealgorithm_skilldefines it — e.g. evo-graph → weaknesses). GitHub is mirror-only; the run dir stays authoritative.
Orchestration mode
- Ask: deterministic or agent? →
orchestration_mode(defaultdeterministic).deterministic— cap-evolve sequences intake→…→algorithm→finalize; honesty is code-enforced. Best when a deterministic engine exists for the algorithm.agent— the coding agent drives the loop itself (reads the algorithm's "Agent-mode loop"), self-policing honesty, and seals via the finalize phase. Required for agent-only algorithms. The purpose-built fully-agentic algorithm isalgorithm_skill: agent-optimize(free-form loop; it also does a Phase-0 understand-the-benchmark step). In agent mode also collectstop_condition.
Stop condition (agent mode)
- Free-text halt rule re-read each round →
stop_condition. Deterministic mode leaves it blank and uses budget knobs.
What it does
- Interview (driven by this SKILL.md): pick the capability skill (what is optimized), the optimizer (which coding agent proposes edits), the algorithm (the search loop), the dataset, the splits, and the budget.
- Scaffold
.capevolve/project/from the template (scripts/run.py): adapter stub,inputs/,capevolve.yaml,PROJECT.md, and the optimizer-prompt templateoptimizer/INSTRUCTIONS.md(the wholetemplates/project/tree is copytree'd verbatim, so this file is already in place — confirm it exists). - Resolve inputs per
inputs/INPUTS.md— the contract below. - Wire trajectories + scoring into the adapter. From the trajectories path
and metric extraction / scoring source inputs:
- implement
adapter.trajectories(split)to RETURN the runner's native trajectory directory for the last eval ofsplit(any structure/format; it is copied verbatim into the optimizer's./trajectories/). ReturnNoneonly if there is genuinely no separate native store (cap-evolve then falls back to its per-rollout JSON) — note that choice inPROJECT.md. - implement
score()to extract the OBJECTIVE metric from a rollout, matching the benchmark's own scoring source; verify it reproduces the benchmark's number. - Make
score()'s feedback ARGUMENT-LEVEL — it IS the learning signal. A tool-name-only signal ("action X was wrong") is too coarse for the optimizer to localize a fix; it pattern-matches to prose rules and the run plateaus. For EACH failing check, the feedback must point at the specific argument/value/step that was wrong: name the wrong ARGUMENT key and the agent's OWN wrong value (not the gold value), name the wrong target id, and for communication/omission misses name the value or field the agent failed to state when it is derivable from the agent's own state (e.g. an un-stated computed total). This is gold-SAFE: derive everything from the agent's own messages/tool-calls/observed state (and the user's own profile/db state the agent saw) — use the gold record ONLY to learn WHICH check/argument failed (key names are safe; gold VALUES must never be read or printed). When a piece is not safely derivable, fall back to the coarser tool-name message. Keepscore()deterministic (the check gate requires it).
- implement
- Author the optimizer instructions for THIS benchmark — SCOPED TO THE SELECTED
CAPABILITIES. Customize the scaffolded
.capevolve/project/optimizer/INSTRUCTIONS.md. Keep the{{...}}placeholders intact ({{FOCUS_SUMMARY}},{{FAILURES}},{{CAP_BRIEF}},{{ALGO_BRIEF}},{{BENCH_REPO}}— the harness fills them per iteration). Keep the authored static guidance short on meta-narration, explicit and DEMANDING on iteration depth, and make it capability-scoped: include guidance, skill references, and edit-space ONLY for the capabilities actually listed incapevolve.yaml: capabilities.- DEPTH MANDATE — address ALL failure clusters each iteration. The authored instructions must demand a substantial multi-root-cause pass: diagnose ALL clusters and fix as many as possible in ONE candidate, spanning EVERY edit class the selected capabilities offer (each capability's SKILL.md enumerates its own). Scope each fix to protect passing tasks; do NOT trade breadth for caution. State plainly that a single small edit is an under-used iteration. For the concrete per-capability wording of this mandate, use the selected capability's own optimizer playbook (see the capability-playbook rule below).
- State the GOAL up front: maximize the eval score — make the largest improvement you can this iteration, grounded in the trajectories.
- The authored INSTRUCTIONS MUST encode all of these (generic, capability-scoped):
- STEP-0 reading mandate. Before diagnosing, the optimizer must READ
./guidance/<cap>/SKILL.md(for EACH selected capability) and the optimizer features reference under./guidance/optimizer/. State this as an explicit first step. - Each selected capability's own optimizer playbook. For EVERY capability in
capevolve.yaml: capabilities, load that capability's playbook from its skill. Both path forms matter, each for its own reader: YOU read it in this repo atskills/capabilities/<cap>/references/(e.g.tools→optimizer-playbook.md), but any path you WRITE INTO the authored INSTRUCTIONS must be the runtime-materialized form./guidance/<cap>/references/— the repo-relative path does not resolve in the optimizer's working dir. Fold its edit-depth demands and subagent-fan-out pattern into the authored INSTRUCTIONS. Do NOT restate one capability's playbook here — it lives with that capability and evolves with it. - The NON-OVERFITTING guardrail. Demand that every prompt/tool edit encode
a GENERAL rule/policy/validation that generalizes across the whole class of
inputs — NEVER hardcode a specific task's id/value/date/name/answer. A guard
must fire on the general condition (e.g. "id not in the user's profile"), not
match a literal value (NOT
if id == "<TASK_SPECIFIC_ID>"). A literal special-case that only helps one task is forbidden — it overfits, fails the held-out gate, and hurts other tasks. Per-task specifics are for understanding the failure CLASS only; the fix must be general. - EXPLOIT ground-truth/eval present in the trajectories (diagnosis only).
Tell the optimizer that when
./trajectories/include ground-truth / expected actions / a reward breakdown, it should USE them during diagnosis to localize the exact defect (expected vs actual action/argument/value) — and if not present, infer from the traces + feedback. State plainly that ground truth informs the failure class only; the resulting edit must still be GENERAL (guardrail 3) and never copy a gold value.
- STEP-0 reading mandate. Before diagnosing, the optimizer must READ
- Capability-scoping (the key rule): reference
./guidance/<cap>/SKILL.mdand present the editable artifacts for the selected caps only. If onlytoolsis selected, do NOT include any prompt-editing guidance, do NOT reference thesystem-promptskill, and do NOT present the prompt/policy file as editable. If onlysystem-promptis selected, do not surface the tools file as editable. Each capability's "What you can change here" lives in its./guidance/<cap>/SKILL.md— point the optimizer there rather than restating it. - Always include (capability-agnostic): READ the four cross-iteration files in
the working dir FIRST —
./LEDGER.md(framework facts: each iteration's outcome + tasks broken/fixed), the whole./JOURNAL.md(the optimizer's append-only handover across the run), and./RUNMAP.md+./prior_iterations/<id>/(every prior iteration's PROCESS.md + capability diff) — and never re-propose an approach the journal/ledger shows was rejected as implemented (a better-designed version may still work — not a permanent ban); READ and USE the diagnose skill./guidance/diagnose/SKILL.md(incl. its KNOWLEDGE / BEHAVIORAL / CAPABILITY-GAP tags), the optimizer features reference./guidance/optimizer/<name>.md, this step's./trajectories/, and any./guidance/sources/data-model files. Each iteration the optimizer MUST: fill./PROCESS.md(the required explainability template — ranked issues + tags, every edit + class, verify-the-fix, subagents/features used, what to preserve, what was skipped) and APPEND its entry to./JOURNAL.mdbelow the marker (what was tried / worked / regressed / refuted / plateau-signal / focus-next). Ship MULTIPLE edit classes and ADD a new code-bearing tool whenever a CAPABILITY-GAP/stall cluster is present.
- Set the spec keys in
capevolve.yaml:runner_repo_path— the benchmark/runner source, surfaced read-only to the optimizer.optimizer_instructions_file— point at the customized template (defaultoptimizer/INSTRUCTIONS.md).capability_sources— the benchmark's data-model / types source files that the tools import (resolved relative to the project dir; copied verbatim into the optimizer's./guidance/sources/), so the optimizer can write correct code against the real types. Set this whenever a selected capability's code imports a shared types/data-model module; leave the default[]when there is none.target_model(+ optionaltarget_profile_file) — the runtime/CONSUMING LLM the agent reads the capabilities with, DISTINCT fromoptimizer_model. A model id or a tier (frontier|strong|mid|weak); steers the optimizer prompt + capability guidance to optimize FOR that reader. Ask which model the agent runs at runtime; leave blank (profile-agnostic) if unknown. Seeinputs/INPUTS.md→target_model.
Ask-the-user-if-missing (mandatory — the core discipline)
Read inputs/INPUTS.md. It classifies every input as NEEDED or
RECOMMENDED. For each NEEDED input that is not already present:
ASK THE USER. Quote (a) the exact path where it is expected, (b) the command or option that produces it, and (c) any alternatives. Then wait.
Never invent a NEEDED input. Fabricating a dataset, a scorer, or a gold answer does not unblock the run — it produces a number that measures nothing and hides that fact. A missing tasks file is a question for the user, not a gap for you to paper over. This is the single most important behavior of this phase.
RECOMMENDED inputs have sane defaults and may be skipped — but log every skip
in PROJECT.md (e.g. "num_trials defaulted to 1 — scores will be single-trial,
so the significance gate will correctly reject marginal gains"), so the honesty
cost of each default is visible at report time.
Block on a missing NEEDED input (never fabricate)
The action when a NEEDED input is absent depends on the run mode:
- Interactive / chat mode — ASK THE USER and wait: quote what is needed, why
it is needed (what breaks without it), and how to provide it (the exact path /
command / option / alternatives from
INPUTS.md). Do not proceed past the missing input. - Non-interactive mode (
cap-evolve run/ orchestrate, no human to ask) — do NOT fabricate. WRITE a clearly delimited section intoPROJECT.md:BLOCKED: <input> — why it is needed — how to provide it, then STOP with a non-zero exit. A blocked-but-honest stop is correct; a green run on a guessed input is not.
This extends the ask-if-missing discipline above — it is the same rule, with an explicit non-interactive fallback so a headless run fails loud and recorded instead of silently inventing a dataset, scorer, trajectories path, or scoring source.
Why a contract, not a guess
The classic failure mode of "auto-optimize my agent" tooling is to start running
with whatever it can find and backfill assumptions. That yields a green run and a
worthless result. Splitting inputs into NEEDED (blocking → ask) vs RECOMMENDED
(default → log) makes the only legitimate way to proceed-without-an-input an
explicit, recorded default — never a silent fabrication. Treat INPUTS.md as the
spec; this SKILL.md is just the procedure for honoring it.
Dual-mode
This phase runs two ways from the same SKILL.md: standalone as the slash command /cap-evolve:intake (the argument-hint shows its run.py args), and orchestrator-callable — cap-evolve run / the orchestrate skill invokes the same scripts/run.py headlessly and threads the run dir between phases.
How to run
python scripts/run.py --base .capevolve # scaffold .capevolve/project
The script is purely mechanical: it copies the template and prints the next
steps. The judgment — interviewing, choosing components, and running the
ask-if-missing loop — is yours, driven by this SKILL.md and inputs/INPUTS.md.
Then implement adapters/adapter.py, fill capevolve.yaml, and proceed to
implement-and-check. Together, intake → implement-and-check is the full
integration: scaffold → implement the 3 required adapter methods →
cap-evolve check green, before any budget is spent. The using-agent (e.g. the chosen optimizer) can run
this whole integration autonomously.
Worked example (onboard a new benchmark from a prompt): see
examples/for an end-to-end onboarding. The intake/onboarding step installs the benchmark (clones + installs it) and the optimizer agent wires the adapter from the stub untilcap-evolve checkpasses, then optimizes the selected capability. The example'ssetup.shis the executable transcript of that onboarding;run.shruns the full optimization with the live dashboard.
What good vs bad intake looks like
- Good: every NEEDED input resolved to a real path or
"adapter"; splits and budget chosen deliberately; each defaulted RECOMMENDED input logged inPROJECT.md;capevolve.yamlfully filled; the user answered every blocking question before the scaffold was declared done. - Bad: a tasks file that "looked plausible" was synthesized; the scorer leaks the gold answer into feedback; test == train with no note; budget left at a default that cannot possibly find a gain; the run proceeded past a missing NEEDED input "to keep moving".
- Bad: authored INSTRUCTIONS that omit the STEP-0 reading mandate, or that skip a selected capability's own optimizer playbook — so an iteration can pass on a cosmetic edit that the capability's playbook calls under-used.
References
references/concepts.md— the inputs contract, NEEDED vs RECOMMENDED rationale, the 3 required adapter methods, and split/trial/budget guidance with sources.