Best for
- User says "post-release tasks", "update dashboard", "run benchmarks"
- After successful release (once GitHub release is published)
- User asks about eval baselines or benchmark results
sunholo-data/ailang/.claude/skills/post-release/SKILL.md
Run automated post-release workflow (eval baselines, dashboard, docs) for AILANG releases. Runs the tier-based benchmark suite (core+stretch+frontier by default) for standard + agent evals with validation and progress reporting. Use when user says "post-release tasks for vX.X.X" or "update dashboard". Fully autonomous with pre-flight checks.
Decision brief
Run post-release tasks for an AILANG release: evaluation baselines, dashboard updates, and documentation.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/sunholo-data/ailang --skill ".claude/skills/post-release"Inspect the Agent Skill "post-release" from https://github.com/sunholo-data/ailang/blob/9944e264e3b9043881978731dccd258f561082a3/.claude/skills/post-release/SKILL.md at commit 9944e264e3b9043881978731dccd258f561082a3. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Review the “Quick Start” section in the pinned source before continuing.
If release doesn't exist, run release-manager skill first.
tools/publish-unified-dashboard.sh vX.X.X
Use the data above first. Only re-run these commands manually if the injected context is empty or you need to refresh after making changes.
Review the “User says: "Run post-release tasks for v0.3.14"” section in the pinned source before continuing.
Permission review
The documentation includes network, browsing, or remote request actions.
Visit: http://localhost:3000/ailang/docs/benchmarks/performanceThe documentation asks the agent to run terminal commands or scripts.
git tag -l vX.X.XThe documentation asks the agent to run terminal commands or scripts.
git add docs/static/codebase_stats.jsonThe documentation includes network, browsing, or remote request actions.
curl -s https://ailang-dev-dashboard-ejjw6zt3bq-ew.a.run.app/benchmarks/os/latest.json \The documentation asks the agent to create, modify, or delete local files.
# Create test file and verify behavior matches documentationEvidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 91/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 33 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Run post-release tasks for an AILANG release: evaluation baselines, dashboard updates, and documentation.
Use the data above first. Only re-run these commands manually if the injected context is empty or you need to refresh after making changes.
Most common usage:
# User says: "Run post-release tasks for v0.3.14"
# This skill will:
# 1. Run eval baseline (extended_suite: 18 production models + lang-harness sweep) - ALWAYS USE --full FOR RELEASES
# 2. Update website dashboard (JSON with history preservation)
# 3. Update axiom scorecard KPI (if features affect axiom compliance)
# 4. Extract metrics and UPDATE CHANGELOG.md automatically
# 5. Move design docs from planned/ to implemented/
# 6. Run docs-sync to verify website accuracy (version constants, PLANNED banners, examples)
# 7. Commit all changes to git
🚨 CRITICAL: For releases, ALWAYS use --full flag by default
Invoke this skill when:
scripts/run_eval_baseline.sh <version> [--full] [--cross-harness]Run evaluation baseline for a release version.
🚨 CRITICAL: ALWAYS use --full for releases!
Usage:
# ✅ RECOMMENDED release baseline (standard + agent + 4-language Explorer sweep)
.claude/skills/post-release/scripts/run_eval_baseline.sh 0.15.0 --full --lang-harness
# Standard + agent only (no 4-lang Explorer sweep)
.claude/skills/post-release/scripts/run_eval_baseline.sh v0.15.0 --full
# Major release: includes cross-harness comparison (gpt5-5 + opencode-gpt5-5 etc.)
.claude/skills/post-release/scripts/run_eval_baseline.sh 0.15.0 --full --cross-harness
# ❌ Dev only — 3 cheap models, AILANG lang only (quick testing/validation)
.claude/skills/post-release/scripts/run_eval_baseline.sh 0.15.0
Cost: see the "Cost & time" table below — the per-mode figures live in ONE place there (and are currently flagged stale/understated). Don't duplicate cost numbers here.
Output:
Running eval baseline for 0.3.14...
Tier scope: --tier core,stretch,frontier (56 benchmarks)
Mode: FULL (extended_suite, 18 models incl. claude-fable-5 + claude-opus-5, + agent_suite)
Expected cost: run `ailang eval-suite --full --tier core,stretch,frontier --dry-run` for a real
computed estimate — recent full releases banked $98-135 combined; see Cost & time below.
Expected time: ~45-90 minutes
Standard-eval budget cap: $150 (see `ailang eval-suite --dry-run` for a real per-run estimate)
[Running benchmarks...]
✓ Baseline complete
Results: eval_results/baselines/0.3.14
Files: 726 result files
What it does:
extended_suite (--full, 18 models — roster as of 2026-07-27; models.yml is the source of truth, this list is a convenience copy):
dev_models (default): gpt5-4-mini, claude-haiku-4-5, gemini-3-flashgemini-3-5-flash-lite is in the suite as a language-improvement gauge, not for capability (v0.30.0-era baseline to lift: core 14/19, stretch 9/21, frontier 1/16).claude-fable-5 refuses a chunk of the Python side (17/53 at v0.30.0), so its Python row under-counts by design.CAVEATS.md).agent_suite (6 cloud weak+reference models, run in parallel): gpt5-6-luna (codex — OpenAI weak/fast), claude-haiku-4-5 (claude — Anthropic weak), claude-sonnet-4-6 (claude — longitudinal anchor), opencode-or-deepseek-v4-pro (OS agent champion), opencode-or-deepseek-v4-flash (OS best-value), opencode-or-glm-5-2 (GLM-5.2 — SETTLED 2026-07-19: beats 5.1 in agent mode 77% vs 73%; opencode-or-glm-5-1 retired to opt-in).dev.ailang.os-rotation-filler → eval_results/rotation/os-rolling, --agent --bank-by-version). For the on-device-vs-cloud agent table, aggregate the rotation's per-version GPU data with this suite's results.agent_suite, but the "it hangs" reason is obsolete*: the 2026-06-04 removal (0 completions, orphaned subprocesses) was fixed by the ollama-loop-convergence work, and motoko was re-added to ollama_suite 2026-06-15 (motoko-local-qwen3-6-35b-a3b-mxfp8, validated on the profile matrix, $0). It stays out of agent_suite because that suite is the cloud weak+reference set — local motoko is covered by ollama_suite + the rig rotation. See MOTOKO.md before touching any motoko checkout.smoke (23), core (23), stretch (25), frontier (8), vision (9) — counts as of the v0.32.0 curation cycle (4 stretch→core promotions + 8 frontier→stretch demotions, both audited from the v0.32.0 baseline against CURATION.md §5; see design_docs/implemented/v0_32_0/m-eval-standard-confidence-gating.md). Frontier halved from 16→8 because exactly half no longer had any of the 3 flagship models (gpt5-6-sol/claude-opus-5/gemini-3-1-pro) failing them in standard mode — worth re-auditing again once the roster has settled for a release or two.smoke for cloud models. Smoke is the cheap/fast sanity tier for the local OS-model iteration loop (the nightly rig, Ollama, de-flaking). Cloud/API models (Anthropic, OpenAI, Google, OpenRouter) go straight to core,stretch,frontier — smoke would just spend API budget re-confirming saturated benchmarks every model already passes, with zero added signal. The only time smoke joins a cloud run is an explicit --tier smoke,core,stretch,frontier full audit.core,stretch,frontier — Core is the headline metric, Stretch is harder mixed results, Frontier is the top-end discriminator (release baselines are its only routine data source)core 70%+ for AILANG; vision intentionally lowlang_harness_suite: claude-haiku-4-5, gemini-3-flash, gpt5-4-mini, opencode-haikucore only (19 benchmarks) — stretch/frontier are skipped here even if a wider --tier was set globallycontract_bst_validate, contract_roman_numeral, effect_composition, effect_tracking_io_fs) and auto-skip on JS/Go runsharness_suite (8 models, paired across harnesses)
opencode-gpt5-6-sol pair)eval_results/baselines/vX.X.X/⚠️ Note: gpt5-5-pro is in models.yml but not in any default suite — agent mode is blocked
(codex rejects with ChatGPT account, opencode returns 0 tool calls). Don't add it to suites.
scripts/update_dashboard.sh <version>Update website benchmark dashboard with new release data.
Usage:
.claude/skills/post-release/scripts/update_dashboard.sh 0.3.14
Output:
Updating dashboard for 0.3.14...
1/5 Generating Docusaurus markdown...
✓ Written to docs/docs/benchmarks/performance.md
2/5 Generating dashboard JSON with history...
✓ Written to docs/static/benchmarks/latest.json (history preserved)
3/5 Validating JSON...
✓ Version: 0.3.14
✓ Success rate: 0.627
4/5 Clearing Docusaurus cache...
✓ Cache cleared
5/5 Summary
✓ Dashboard updated for 0.3.14
✓ Markdown: docs/docs/benchmarks/performance.md
✓ JSON: docs/static/benchmarks/latest.json
Next steps:
1. Test locally: cd docs && npm start
2. Visit: http://localhost:3000/ailang/docs/benchmarks/performance
3. Verify timeline shows 0.3.14
4. Commit: git add docs/docs/benchmarks/performance.md docs/static/benchmarks/latest.json
5. Commit: git commit -m 'Update benchmark dashboard for 0.3.14'
6. Push: git push
What it does:
scripts/extract_changelog_metrics.sh [json_file]Extract benchmark metrics from dashboard JSON for CHANGELOG.
Usage:
.claude/skills/post-release/scripts/extract_changelog_metrics.sh
# Or specify JSON file:
.claude/skills/post-release/scripts/extract_changelog_metrics.sh docs/static/benchmarks/latest.json
Output:
Extracting metrics from docs/static/benchmarks/latest.json...
=== CHANGELOG.md Template ===
### Benchmark Results (M-EVAL)
**Overall Performance**: 59.1% success rate (399 total runs)
**By Language:**
- **AILANG**: 33.0% - New language, learning curve
- **Python**: 87.0% - Baseline for comparison
- **Gap: 54.0 percentage points (expected for new language)
**Comparison**: -15.2% AILANG regression from 0.3.14 (48.2% → 33.0%)
=== End Template ===
Use this template in CHANGELOG.md for 0.3.15
What it does:
scripts/cleanup_design_docs.sh <version> [--dry-run] [--force] [--check-only]Move design docs from planned/ to implemented/ after a release.
Features:
Usage:
# Check-only: Report issues without making changes
.claude/skills/post-release/scripts/cleanup_design_docs.sh 0.5.9 --check-only
# Preview what would be moved/deleted/relocated
.claude/skills/post-release/scripts/cleanup_design_docs.sh 0.5.9 --dry-run
# Execute: Move implemented docs, delete duplicates, relocate misplaced
.claude/skills/post-release/scripts/cleanup_design_docs.sh 0.5.9
# Force move all docs regardless of status
.claude/skills/post-release/scripts/cleanup_design_docs.sh 0.5.9 --force
Output:
Design Doc Cleanup for v0.5.9
==================================
Checking 5 design doc(s) in design_docs/planned/v0_5_9/:
Phase 1: Detecting issues...
[DUPLICATE] m-fix-if-else-let-block.md
Already exists in design_docs/implemented/v0_5_9/
[MISPLACED] m-codegen-value-types.md
Target: v0.5.10 (folder: v0_5_9)
Should be in: design_docs/planned/v0_5_10/
Issues found:
- 1 duplicate(s) (can be deleted)
- 1 misplaced doc(s) (wrong version folder)
Phase 2: Processing docs...
[DELETED] m-fix-if-else-let-block.md (duplicate - already in implemented/)
[RELOCATED] m-codegen-value-types.md → design_docs/planned/v0_5_10/
[MOVED] m-dx11-cycles.md → design_docs/implemented/v0_5_9/
[NEEDS REVIEW] m-unfinished-feature.md
Found: **Status**: Planned
Summary:
✓ Deleted 1 duplicate(s)
✓ Relocated 1 doc(s) to correct version folder
✓ Moved 1 doc(s) to design_docs/implemented/v0_5_9/
⚠ 1 doc(s) need review (not marked as Implemented)
What it does:
--check-only to only report issues, --dry-run to preview actions, --force to move allgit tag -l vX.X.X
gh release view vX.X.X
If release doesn't exist, run release-manager skill first.
The docs/static/codebase_stats.json file drives the Codebase Statistics page (LOC/token/commit growth chart).
⚠️ The deploy workflow regenerates this file at build time but does NOT commit it back — it only bakes the result into the Pages artifact. The generator (generate_codebase_stats.sh) appends only the version it runs on to the committed history. So if a release is not captured + committed here, that version is permanently skipped from the history chart, producing visible gaps (e.g. the 0.14.1 → 0.25.0 jump that skipped every 0.15–0.24 graduation). You must run + commit this every release.
# Generate stats for THIS release and append to the committed history
AILANG_VERSION=vX.X.X bash tools/generate_codebase_stats.sh
# Sanity check: current == this release, and the new entry is in history
jq -r '.current.version, (.history[-1].version)' docs/static/codebase_stats.json
# Commit (CI will NOT do this for you)
git add docs/static/codebase_stats.json
git commit -m "data(stats): codebase statistics for vX.X.X"
git push
codebase_stats.json current shows vX.X.Xhistory includes a vX.X.X entry (no gap vs. previous release)Backfilling a missed gap: check out each missed tag in a throwaway
git worktree, run the same counting logic, and merge the entries intohistorysorted by version. See the v0.15→0.24 backfill (June 2026) for the pattern.
The dashboard fetches the benchmark JSONs at runtime from the GCS bucket via the dashboard's
/benchmarks/ route (M-EVAL-DATA-HOSTING-DECOUPLE), so the rig no longer commits them every
45 minutes — W5 retired that churn. The committed copies now serve two purposes only:
Between releases the working tree carries these files as modified-but-uncommitted; that drift is expected ("the bucket has newer data than HEAD"). This step is what sweeps it up. Skip it and the fallback copy silently ages another release, so an outage would serve stale numbers.
# The rig regenerates these continuously; commit the current state as the release snapshot.
git add docs/static/benchmarks/latest.json \
docs/static/benchmarks/os/latest.json \
docs/static/benchmarks/os/history.json
git commit -m "data(bench): release provenance snapshot for vX.X.X"
git push
# Verify the runtime path agrees with what you just committed
curl -s https://ailang-dev-dashboard-ejjw6zt3bq-ew.a.run.app/benchmarks/os/latest.json \
| jq -r '.ailang_version, .version, .generated'
latest.json ailang_version shows vX.X.X (not the previous release)ailang_version (bucket and git agree at release time)If the runtime route and the committed copy disagree, the bucket sync is failing — check
bucket synclines in/tmp/ailang-os-filler.logbefore shipping the release notes.
🚨 CRITICAL: ALWAYS use --full for releases!
Correct workflow:
# ✅ RECOMMENDED for releases — adds 4-language Explorer data for ~$7 more
.claude/skills/post-release/scripts/run_eval_baseline.sh X.X.X --full --lang-harness
# Minimum acceptable for releases
.claude/skills/post-release/scripts/run_eval_baseline.sh X.X.X --full
This runs all 18 production models (extended_suite) with both AILANG and Python.
Tier scope for releases (counts re-centered after the v0.32.0 curation cycle: 4 stretch→core promotions, 8 frontier→stretch demotions, audited from the v0.32.0 baseline against CURATION.md §5):
--tier core,stretch,frontier — 56 benchmarks (23 core + 25
stretch + 8 frontier). This is the tier for every cloud/API model — never smoke.
Smoke is the local OS-model iteration tier only (see the 🚫 note above); a cloud model
added to a release baseline (e.g. a new Anthropic/OpenAI/Google model) goes straight to
core,stretch,frontier.frontier (8, halved from 16 this cycle): the anti-saturation discriminator tier — a
frontier benchmark's defining property is that at least one frontier model FAILS it in
standard mode; if every frontier model passes it, it demotes back to stretch
(CURATION.md §5). Release baselines are the ONLY routine source of frontier-failure
data (its authoring-time failure validation was parked as API-billed), so keep it in
every release run — that data doubles as the check that the tier still discriminates.--tier core — 23 benchmarks (Core is the headline metric)--tier smoke,core,stretch,frontier — 79 benchmarks; only add smoke when you
deliberately want the local sanity tier in the sweep, not for routine cloud baselinesvision benchmarks are research-grade and excluded by default — opt in explicitly
with --tier vision if you want to publish those numbers.Override tier via the script's --tier flag (see run_eval_baseline.sh --help). If
unsure, the default is tuned to produce a release-ready baseline in ~30–60 minutes.
Cost & time (default tier core,stretch,frontier = 56 benchmarks):
Get a real number before you spend, not a guessed one. ailang eval-suite --dry-run
(with the same --models/--tier/--benchmarks you're about to run) now prints a computed
$ estimate — historical mean tokens per benchmark from the most recent baseline × models.yml
pricing (M-EVAL-STANDARD-CONFIDENCE-GATING). Pairs with no history are flagged explicitly, not
silently folded into the total. Run it first; the table below is orientation, not a substitute.
Real banked cost (from cost_usd on actual result files, not projected):
| Baseline | Standard | Agent | Combined |
|---|---|---|---|
| v0.30.0 | $98.33 (1810 files) | $15.07 (280 files) | $113.40 |
| v0.29.2 | $102.89 (1853 files) | $32.33 (386 files) | $135.22 |
Standard-eval cost is concentrated: in v0.30.0, the top 5 of 18 models were 71% of spend
(claude-fable-5 alone was 32%). --lang-harness adds ~$7-10; --cross-harness roughly
triples the agent step (~3x its base cost). Time: ~45-90 min for --full, +15-20 min per
extra step (--lang-harness/--cross-harness).
Confidence-gated standard eval (default-on as of v0.32.0): --confidence-gate /
--no-confidence-gate. --full runs skip re-confirming Trivial-band (ELO-saturated)
core/stretch benchmarks per observatory.db ratings — frontier always stays full (its
curation contract requires routine full-coverage failure data), and any model with no
standard-mode rating history (a new/swapped flagship) always gets the full tier regardless.
Falls open to today's exact full-tier behavior if the ratings DB is unseeded, missing, or
corrupt. Use --no-confidence-gate to force the old full-tier behavior for the periodic
full-audit cadence below, or to hand-audit a roster change. First dry-run projection: full-tier
$49.57 vs gated $35.28, ~29% — directional, not yet validated against a live release's actual
spend (that validation happens automatically as releases accumulate real observatory.db
history). See m-eval-standard-confidence-gating.md's
Verification section for the full caveat. Cadence: even once trusted, force a full
(non-gated) run at least quarterly or on any extended_suite roster change — confidence-gating
is benchmark-centric, not model-centric, so a stable model regressing on a previously-Trivial
benchmark needs a periodic full re-audit to catch, not just the new-model rule.
Budget cap: --budget-usd (run_eval_baseline.sh flag, defaults to $150 for --full
runs — ~15% headroom over the highest real combined baseline above; dev-mode runs are uncapped
unless set explicitly). On breach, in-flight trials finish, no new trial is scheduled, and
baseline.json gets budget_stopped: true — check for that key before treating a baseline as
complete. Override with --budget-usd <N>, or --budget-usd 0 to disable the cap entirely.
❌ WRONG workflow (what happened with v0.3.22):
# DON'T do this for releases!
.claude/skills/post-release/scripts/run_eval_baseline.sh X.X.X # Missing --full!
# Then try to add production models later with --skip-existing
# Result: Confusion, multiple processes, incomplete baseline
If baseline times out or is interrupted:
# Resume with ALL extended_suite models (maintains --full semantics)
ailang eval-suite --full --langs python,ailang --parallel 5 \
--output eval_results/baselines/X.X.X --skip-existing
The --skip-existing flag skips benchmarks that already have result files, allowing resumption of interrupted runs. But ONLY use this for recovery, not as a strategy to "add more models later".
Use the automation script:
.claude/skills/post-release/scripts/update_dashboard.sh X.X.X
IMPORTANT: This script automatically:
Test locally (optional but recommended):
cd docs && npm start
# Visit: http://localhost:3000/ailang/docs/benchmarks/performance
Verify:
Commit dashboard updates:
git add docs/docs/benchmarks/performance.md docs/static/benchmarks/latest.json
git commit -m "Update benchmark dashboard for vX.X.X"
git push
The local Ollama rig (opencode + pi + motoko on local qwen, via the os-rotation-filler) measures whether each AILANG release moves the needle for local models.
This step is AUTOMATED since 2026-07-20: the os-rotation-filler's release-pickup
step (3b) detects the std/VERSION bump on origin/dev within one cycle (~45 min),
pulls, reinstalls the binary, and runs the snapshot/reset itself. Check
/tmp/ailang-os-filler.log for a release pickup complete line. Only run the
manual command below if the log shows snapshot/reset ... failed or the rig was
down at release time:
# MANUAL FALLBACK — normally done by the filler's release pickup automatically.
# Snapshot the rig's current numbers as this release, AND reset the active-model
# accumulator so the rotation re-measures fresh against the NEXT release.
# (Retired models — anything not matching ACTIVE_PATTERN, default qwen3-6 — stay
# frozen at their last version.)
tools/os-release-snapshot.sh vX.X.X --reset
git add docs/static/benchmarks/os/history.json docs/static/benchmarks/os/latest.json
git commit -m "Snapshot local-rig OS leaderboard for vX.X.X (longitudinal)"
git push
This appends a vX.X.X entry to docs/static/benchmarks/os/history.json (deduped
by version) and clears the active-model rolling files. The website's
local-rig trend chart reads os/history.json; the live table still reads
os/latest.json. Run without --reset if you only want to refresh a version's
numbers while they're still filling (safe, idempotent).
The whole point of the on-device roster is comparison with cloud. Step 3 above
publishes the cloud baseline into the main latest.json; this step folds the local
rig's rotation results for the same release into that SAME leaderboard so the
on-device models (qwen/gemma) appear alongside the cloud frontier in the main tables
(ELO leaderboard, gap-trend, per-model-trend). It re-runs eval-report on the cloud
baseline with --merge pointed at this release's rotation dir:
# Regenerates docs/static/benchmarks/latest.json with cloud + local UNIFIED.
# Auto-skips the merge if no rotation dir exists yet for this release
# (then it's just a cloud-only refresh, identical to step 3).
tools/publish-unified-dashboard.sh vX.X.X
git add docs/static/benchmarks/latest.json
git commit -m "Unify local rotation into main leaderboard for vX.X.X"
git push
Under the hood this is:
ailang eval-report eval_results/baselines/vX.X.X vX.X.X --merge eval_results/rotation/os-rolling/vX.X.X --format=json
(the wrapper adds --merge only when that rotation dir exists). Do not redirect
its stdout — eval-report writes latest.json itself and preserves history. Verify
ratings.agent.byLang.ailang.models now lists BOTH cloud (claude-*, opencode-or-*)
and local (*-qwen*, *gemma*) models, while ratings.standard still has the cloud
roster. Note the daily os-rotation-filler also runs this automatically once a release's
cloud baseline exists, so this step mainly guarantees the unify happens at release.
Review and update the axiom scorecard:
# View current scorecard
ailang axioms
# The scorecard is at docs/static/benchmarks/axiom_scorecard.json
# Update scores if features were added/improved that affect axiom compliance
When to update scores:
Update history entry:
{
"version": "vX.X.X",
"date": "YYYY-MM-DD",
"score": 18,
"maxScore": 24,
"percentage": 75.0,
"notes": "Added capability budgets (A9 +1)"
}
Generate metrics template:
.claude/skills/post-release/scripts/extract_changelog_metrics.sh X.X.X
This outputs a formatted template with:
Update changelog automatically:
changelogs/ (find with ls changelogs/ | grep current)CHANGELOG.md is an index file — do NOT write entries thereFor CHANGELOG template format, see resources/version_notes.md.
v0.8.0+ (chain-based - recommended):
# Find the chain ID from the latest eval run
ailang eval-chains list
# View per-benchmark pass/fail with cost and turns
ailang eval-chains view <chain-id>
# Pass rate breakdown
ailang eval-chains stats <chain-id>
# Show failures with error details
ailang eval-chains failures <chain-id>
# Generate chain-based report
ailang eval-report --from-chain <chain-id> X.X.X --format=json
Legacy (file-based):
# Get KPIs (turns, tokens, cost by language)
.claude/skills/eval-analyzer/scripts/agent_kpis.sh eval_results/baselines/X.X.X
Target metrics: Avg Turns ≤1.5x gap, Avg Tokens ≤2.0x gap vs Python.
For detailed agent analysis guide, see resources/version_notes.md.
After dashboard + changelog metrics, run the curation analysis primitives. These inform
what to keep, demote, or promote for the next release — the suite is curated, not
accumulated. See benchmarks/CURATION.md for the full philosophy.
# Tag coverage: which of the 12 canonical tags are thin or over-represented?
ailang eval-matrix eval_results/baselines/vX.X.X vX.X.X --by-tags
# AILANG-only wins: AILANG beats Python by ≥ 10pp — protect these from regressions
ailang eval-matrix eval_results/baselines/vX.X.X vX.X.X --ailang-wins
# Dual-mode saturation audit (RECOMMENDED for demotion decisions):
# Lists every benchmark's standard AND agent pass rate per language.
# Demote candidates require ≥95% on ALL 4 dimensions (std-AI, std-Py, agent-AI, agent-Py).
# Standard-only saturation is misleading: many benchmarks are 100/100 in standard
# but drop to 12-50% in agent mode — those are still valuable signal, KEEP IN CORE.
.claude/skills/benchmark-manager/scripts/audit_saturation.sh vX.X.X
# Built-in saturated check (uses topline successRate only — less reliable, prefer the script above):
ailang eval-matrix eval_results/baselines/vX.X.X vX.X.X --show-saturated
# Optional: compare this release's tag deltas against the previous baseline
ailang eval-report eval_results/baselines/vX.X.X vX.X.X --format=json
Record three things in the release notes (or the design doc retro):
stretch or be
retired in the next sprint.This step is cheap (seconds), has no external dependencies, and produces the input for the next release's benchmark-manager / eval-gap-finder work.
Step 1: Check for issues (duplicates, misplaced docs):
.claude/skills/post-release/scripts/cleanup_design_docs.sh X.X.X --check-only
Step 2: Preview all changes:
.claude/skills/post-release/scripts/cleanup_design_docs.sh X.X.X --dry-run
Step 3: Check any flagged docs:
[DUPLICATE] - These will be deleted (already in implemented/)[MISPLACED] - These will be relocated to correct version folder[NEEDS REVIEW] - Update **Status**: to Implemented if done, or leave for next versionStep 4: Execute the cleanup:
.claude/skills/post-release/scripts/cleanup_design_docs.sh X.X.X
Step 5: Commit the changes:
git add design_docs/
git commit -m "docs: cleanup design docs for vX_Y_Z"
The script automatically:
prompts/ with latest AILANG syntax (ailang prompt)prompts/devtools/ with latest toolchain docs (ailang devtools-prompt)
ailang devtools-prompt | grep "new-command"docs/) with latest featuresdocs/guides/evaluation/ if significant benchmark improvementsdocs/LIMITATIONS.md:
# Test examples from LIMITATIONS.md
# Example: Test Y-combinator still fails (should fail)
echo 'let Y = \f. (\x. f(x(x)))(\x. f(x(x))) in Y' | ailang repl
# Example: Test named recursion works (should succeed)
ailang run examples/factorial.ail
# Example: Test polymorphic operator limitation (should panic with floats)
# Create test file and verify behavior matches documentation
git add docs/LIMITATIONS.md && git commit -m "Update LIMITATIONS.md for vX.X.X"Run docs-sync to verify website accuracy:
# Check version constants are correct
.claude/skills/docs-sync/scripts/check_versions.sh
# Audit design docs vs website claims
.claude/skills/docs-sync/scripts/audit_design_docs.sh
# Generate full sync report
.claude/skills/docs-sync/scripts/generate_report.sh
What docs-sync checks:
docs/src/constants/version.js match git tagIf issues found:
Commit docs-sync fixes:
git add docs/
git commit -m "docs: sync website with vX.X.X implementation"
See docs-sync skill for full documentation.
release-manager runs make brain-index-syntax-reset immediately after the
tag pushes. This step verifies that the reset actually populated the
brain with chunks tagged with the new release version — protects against
silent indexer failures that would leave Claude pulling stale snippets.
Spot-check ≥5 chunks reference the active version:
EXPECTED_VERSION="$(ailang prompt --version-active)"
ailang cache search --namespace ailang-syntax --limit 5 "string" \
| grep -c "version:${EXPECTED_VERSION}" \
|| { echo "FAIL: μRAG corpus does not reference $EXPECTED_VERSION"; exit 1; }
Quick stat check:
ailang cache stats | grep -E "ailang-(syntax|builtins|examples)"
# Expect: ailang-syntax >= 50, ailang-builtins >= 250, ailang-examples >= 100
Append a one-line audit to release notes (or eval_results/baselines/<version>/notes.md):
μRAG corpus reindex verified: <ailang-syntax count>, <ailang-builtins count>, <ailang-examples count> chunks @ <version>.
If verification fails:
make brain-index-syntax-reset from the project root.ailang prompt --version-active returns the new release tag.See resources/post_release_checklist.md for complete step-by-step checklist.
ailang binary installed (for eval baseline)The agent baseline (agent_suite) routes each model to a CLI harness. To run all of
them on the Studio rig, these must be installed and authenticated:
Harness (agent_cli) | Install | Auth | Used by |
|---|---|---|---|
opencode | npm i -g opencode-ai | OPENROUTER_API_KEY | OS models (glm/minimax/deepseek) — reliable |
claude | npm i -g @anthropic-ai/claude-code | logged-in Claude | sonnet anchor — ⚠️ can hang on long runs |
codex | npm i -g @openai/codex | one-time: printenv OPENAI_API_KEY | codex login --with-api-key (env var alone gives 401 — codex defaults to ChatGPT-OAuth) | gpt5-4-mini |
managed_agents | (none — Vertex API) | gcloud auth application-default login | gemini agent (no gemini CLI executor exists) |
pi | npm i -g @mariozechner/pi-coding-agent | per-provider | optional minimal harness |
motoko | see MOTOKO.md — do NOT go install blind, the checkout matters | OPENROUTER_API_KEY | ollama_suite (local qwen3.6) + harness_suite; not in agent_suite (cloud-only suite) |
API keys live in ~/.config/ailang/secrets.env (sourced from ~/.zshenv). Pull the
cloud-managed ones with ~/.config/ailang/pull-secrets.sh (Anthropic/OpenAI/Google from
Secret Manager). OpenRouter is the manual key. Gemini uses Vertex ADC, not the API key.
Verify everything resolves before a release run:
for c in opencode claude codex pi motoko; do which $c >/dev/null && echo "✓ $c" || echo "✗ $c"; done
# Note: NO gemini CLI — the @google/gemini-cli is deprecated/unused. Google agent mode
# goes through managed_agents (Vertex API via `gcloud auth application-default login`).
ailang eval-suite --agent --models agent_suite --benchmarks fizzbuzz --langs ailang --dry-run # all should route, none "<none>"
Symptom: standard/ has holes for claude-* models, and/or agent/ has NO claude-*
rows at all (only opencode-*). This is what happened to the v0.30.0 baseline: 43 standard
holes and zero Claude agent runs, from an Anthropic quota exhaustion that ran to 2026-08-01.
Do NOT read the resulting low claude-* pass rates as a model or AILANG regression — they
are coverage artifacts.
Before starting a release run: confirm Anthropic quota headroom, since extended_suite now
carries four Anthropic rows (opus-5, fable-5, sonnet-5, sonnet-4-6) and agent_suite two more.
If it happens anyway: record it in the baseline's CAVEATS.md, and resume the missing rows
with --skip-existing once quota returns rather than publishing the partial numbers as-is.
Solution: Use --skip-existing flag to resume:
bin/ailang eval-suite --full --skip-existing --output eval_results/baselines/vX.X.X
Cause: Wrong JSON file (performance matrix vs dashboard JSON)
Solution: Use update_dashboard.sh script, not manual file copying
Cause: Stale build cache
Solution: Run cd docs && npm run clear && rm -rf .docusaurus build
Cause: Didn't run update_dashboard.sh with correct version
Solution: Re-run update_dashboard.sh X.X.X with correct version
This skill loads information progressively:
scripts/ directory (automation)resources/post_release_checklist.md (detailed checklist)Scripts execute without loading into context window, saving tokens while providing powerful automation.
For historical improvements and lessons learned, see resources/version_notes.md.
Key points:
--validate flag to check configuration before running--tier core,stretch,frontier default) — benchmarks
are resolved from benchmarks/*.yml by tier, not a hardcoded listailang eval-matrix --by-tags/--show-saturated/ --ailang-wins in Step 5b--full flag for release baselines (all production models)--check-only to report issues without changes--dry-run to preview all actions--force to move all regardless of statusFrequently asked questions
Run post-release tasks for an AILANG release: evaluation baselines, dashboard updates, and documentation.
The source record exposes this install command: npx skills add https://github.com/sunholo-data/ailang --skill ".claude/skills/post-release". Inspect the command and pinned source before running it.
Static rules flagged network, exec-script, write-files in the source; the page lists the matching lines and excerpts.
Alternatives
sunholo-data/ailang
Run automated post-release workflow (eval baselines, dashboard, docs) for AILANG releases. Runs the tier-based benchmark suite (core+stretch+frontier by default) for standard + agent evals with validation and progress reporting. Use when user says "post-release tasks for vX.X.X" or "update dashboard". Fully autonomous with pre-flight checks.
oaslananka/kicad-mcp-pro
Use this skill for GitHub Copilot pull request and code reviews in oaslananka/kicad-mcp-pro. Review Python MCP server changes, KiCad adapter and tool-contract changes, tests, npm/package wrappers, Tauri/Rust desktop code, GitHub Actions, security controls, documentation, generated metadata, and compatibility/release surfaces. Use it whenever reviewing a PR or diff in this repository, especially changes under src/, tests/, packages/, src-tauri/, .github/workflows/, or public MCP metadata/configur
UiPath/skills
Always invoke for `.xaml` or `.cs` workflow files. UiPath RPA — create, edit, build, run, debug `.cs` coded workflows and `.xaml` workflows. UI automation with Object Repository selectors, test case authoring, Integration Service connector calls. Live desktop/browser UI exploration and control. Deploy via `.uipx`→uipath-solution. Non-solution Orchestrator ops→uipath-platform. Test reports→uipath-test. Agents→uipath-agents.
upex-galaxy/agentic-qa-boilerplate
Walks new users through this repo's QA flow — Playwright + KATA + Allure + Xray stack, Jira QA workflow (Backlog → Shift-Left QA → Estimation → Ready For Dev → Ready For QA → In Test → QA Approved → Ready For Release → Deployed to Production), /shift-left-testing for pre-sprint AC refinement on backlog Stories, /sprint-testing for in-sprint manual QA, /test-documentation for TMS test cases, /test-automation for KATA-compliant E2E/API tests, /regression-testing for CI suite execution, /framework-