Best for
- "When did the X feature stop working?" — pass the feature keyword.
- "Has feature Y improved?" — see the broken/working trend over time.
- Before shipping a fix — sanity check that the regression is reproducible.
sonichi/sutando/skills/regression-search/SKILL.md
Search phone-call history for when a feature regressed (find-regression.py) and drill into a single call to see what went wrong (diagnose-call.py). Skips reading 100+ transcripts by hand.
Decision brief
Two scripts for hunting down bad calls without reading every transcript:
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/sonichi/sutando --skill "skills/regression-search"Inspect the Agent Skill "regression-search" from https://github.com/sonichi/sutando/blob/6a8f0fccd32e5aa620a3572c8885544f144bb6fe/skills/regression-search/SKILL.md at commit 6a8f0fccd32e5aa620a3572c8885544f144bb6fe. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Flags: - --since YYYY-MM-DD — only show calls on/after this date - --json — machine-readable output - --show-snippet — print a one-line transcript snippet for each call
"When did the X feature stop working?" — pass the feature keyword.
A call is broken for a query if any of: - Sutando refuses ("I can't", "I'm not able", "I'm unable", "sorry I cannot") - Sutando reports an error ("error", "failed", "didn't work", "something went wrong") - The user repeats the same request 2+ times in a row (Sutando didn't respo…
Keyword matching only. "recording doesn't stop" vs "recording won't start" both match record. The issue calls this out as future work.
Permission review
The documentation asks the agent to run terminal commands or scripts.
python3 skills/regression-search/scripts/find-regression.py "record"The documentation asks the agent to run terminal commands or scripts.
python3 skills/regression-search/scripts/find-regression.py "summon" --since 2026-04-01Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 86/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 359 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Two scripts for hunting down bad calls without reading every transcript:
find-regression.py — search results/calls/calls.jsonl for calls touching a feature, classify each as working/broken, print a sorted timeline.diagnose-call.py — drill into a single call by SID, report refusals/errors/silences/repeated requests, optionally show metrics from data/call-metrics.jsonl.Closes #188.
python3 skills/regression-search/scripts/find-regression.py "record"
python3 skills/regression-search/scripts/find-regression.py "summon" --since 2026-04-01
python3 skills/regression-search/scripts/find-regression.py "play" --json
Flags:
--since YYYY-MM-DD — only show calls on/after this date--json — machine-readable output--show-snippet — print a one-line transcript snippet for each callA call is broken for a query if any of:
Otherwise the call is working if Sutando's response includes the feature keyword and isn't flagged broken.
These are intentionally crude — the goal is "good enough to find the regression window without reading 163 transcripts." Tune as you find false positives.
record. The issue calls this out as future work.python3 skills/regression-search/scripts/diagnose-call.py de1f04733fc2
python3 skills/regression-search/scripts/diagnose-call.py CA701fc4129779... --metrics
python3 skills/regression-search/scripts/diagnose-call.py de1f04733fc2 --json
Accepts a full SID or just the last 12 characters. Reports turn counts, refusals, errors, silences, repeated user requests, and the ending style (normal vs abrupt user end vs sutando silence). With --metrics, also pulls per-event tool-call timeline from data/call-metrics.jsonl (requires PR #223). Exit code 1 if any issues are found, 0 if clean — useful for CI.
Typical workflow: run find-regression.py to surface broken candidates, then diagnose-call.py <sid> to drill into the worst one.