Source profileQuality 92/100Review permissions

baphuongna/pi-crew/skills/real-test-pi-crew/SKILL.md

real-test-pi-crew

End-to-end verification for pi-crew changes: fast critical tests, 3-path kill-switch proof, bundle md5 sync, live TUI probing, smoke team runs, and a live feature-action battery (team tool + subagent tools).

Source repository stars
50
Declared platforms
1
Static risk flags
3
Last source update
2026-08-24
Source checked
2026-08-25

Decision brief

What it does: where it fits

End-to-end verification discipline for pi-crew changes. Distilled from the broker Phase-4 rollout (commits 1cb2dca → d599578 → 612e18b → 4186284, July 2026). The pain this skill prevents: shipping code that compiles + unit-tests-green but breaks in the user's live Pi session, or…

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorDeclaredSource recordInstall path and trigger
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/baphuongna/pi-crew --skill "skills/real-test-pi-crew"
    Safe inspection promptEditorial

    Inspect the Agent Skill "real-test-pi-crew" from https://github.com/baphuongna/pi-crew/blob/519a5e4e374c1ff6bf02dd05830b11359d1302b1/skills/real-test-pi-crew/SKILL.md at commit 519a5e4e374c1ff6bf02dd05830b11359d1302b1. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      team action='run' team='fast-fix' workflow='fast-fix' goal='...' async=false

      team: action: run run | status | events | cancel | retry | ... team: fast-fix team (a role-set): default / fast-fix / implementation / parallel-research / research / review workflow: fast-fix workflow (a phase DAG): default / fast-fix / plan-execute / implementation / review / r…

      team: action: run run | status | events | cancel | retry | ... team: fast-fix team (a role-set): default / fast-fix / implementation / parallel-research / research / review workflow: fast-fix workflow (a phase DAG): def…
    2. 02

      Core principle: disk ≠ live Pi

      Two locations hold pi-crew state:

      Source (src/, test/, package.json, workflows/, src/runtime/plan-templates.ts) — git-tracked, git diff shows it.Bundle (dist/index.mjs) — pre-built, loaded by Pi at extension cold-start only.Two locations hold pi-crew state:
    3. 03

      Prerequisites

      Before running any tier, verify these are available:

      Before running any tier, verify these are available:Working directory should be the pi-crew repo root:The skill maps to existing CI gates as follows:
    4. 04

      CI integration

      The skill maps to existing CI gates as follows:

      The skill maps to existing CI gates as follows:To add Tier 1 to a pre-commit hook:
    5. 05

      .git/hooks/pre-commit (or via husky / pre-commit framework)

      Review the “.git/hooks/pre-commit (or via husky / pre-commit framework)” section in the pinned source before continuing.

      Review and apply the “.git/hooks/pre-commit (or via husky / pre-commit framework)” source section.

    Permission review

    Static risk signals and limitations

    Runs scripts

    medium · line 67

    The documentation asks the agent to run terminal commands or scripts.

    npm run test:critical || {

    Runs scripts

    medium · line 132

    The documentation asks the agent to run terminal commands or scripts.

    npm run test:critical

    Reads files

    low · line 396

    The documentation asks the agent to read local files, directories, or repositories.

    | Session load model | Same file: "dist/index.mjs (pre-built bundle) if present — DEFAULT since v0.9.17" |

    Writes files

    medium · line 484

    The documentation asks the agent to create, modify, or delete local files.

    | Trusting a team-run agent not to edit the repo under test | Agents spawned by `team`/`Agent`/`crew_agent` inherit the session cwd and have `edit`/`write` tools — a proactive LLM (observed with deepseek) will make **unauthorized source edi

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score92/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars50SourceRepository attention, not individual Skill quality
    Compatibility1 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    baphuongna/pi-crew
    Skill path
    skills/real-test-pi-crew/SKILL.md
    Commit
    519a5e4e374c1ff6bf02dd05830b11359d1302b1
    License
    MIT
    Collected
    2026-08-25
    Default branch
    main
    View the original SKILL.md

    real-test-pi-crew

    End-to-end verification discipline for pi-crew changes. Distilled from the broker Phase-4 rollout (commits 1cb2dcad599578612e18b4186284, July 2026). The pain this skill prevents: shipping code that compiles + unit-tests-green but breaks in the user's live Pi session, or hangs the verifier worker.

    When to use: after any change to src/runtime/crew-broker*.ts, src/ui/, src/config/, src/extension/registration/lifecycle-handlers.ts, src/runtime/child-pi-spawn.ts, src/runtime/plan-templates.ts, src/schema/team-tool-schema.ts (or any Type.Unsafe({...}) schema definition), src/extension/registration/team-tool.ts, workflows/*.workflow.md, or before any commit touching these paths. Schema changes additionally require Tier 9 (feature battery) because the team tool's TypeBox schema is validated by pi-ai BEFORE the handler runs — a too-strict or malformed schema breaks every action silently.

    Core principle: disk ≠ live Pi

    Two locations hold pi-crew state:

    1. Source (src/, test/, package.json, workflows/, src/runtime/plan-templates.ts) — git-tracked, git diff shows it.
    2. Bundle (dist/index.mjs) — pre-built, loaded by Pi at extension cold-start only.

    The 3-way resolution order for dist/index.mjs (per index.ts:5-22):

    1. dist/index.mjs (pre-built bundle) if present  ← DEFAULT since the v0.9.17 bundle-as-default rollout
    2. Inline strip-types loading — fallback when bundle missing
       OR PI_CREW_USE_BUNDLE=0
    

    Note on version pins: this skill mentions specific versions (v0.9.17, v0.9.46, v0.9.47) as anchors for when a behavior was introduced, not as a constraint on which version the skill applies to. The verification discipline (Tiers 1–8) applies to every pi-crew release. Verify the version pin is still accurate via git log --oneline -- index.ts and git log --oneline -- src/ui/run-dashboard.ts.

    Workflow files are runtime dataworkflows/*.workflow.md and task prompt strings inside src/runtime/plan-templates.ts are loaded per-call, NOT bundled. Edits take effect immediately, no rebuild needed.

    The most common silent-failure mode: edit src/, run npm test (pass!), rebuild bundle (good md5!), but the session still has the old code because Pi wasn't /quit-ed + reopened.

    Prerequisites

    Before running any tier, verify these are available:

    ToolUsed inCheck
    node (>=22)Tiers 1, 2, 3node --version
    npmAll tiersnpm --version
    bashAll tiersecho $BASH_VERSION
    md5sumTiers 3, 4, 8which md5sum (or md5 on macOS)
    tmuxTier 5which tmux (optional — Tier 6 is the fallback)
    python3Tier 6python3 --version (optional — Tier 5 is the fallback)
    pi in PATHTiers 5, 6which pi (must be installed via npx pi install .)
    gitReference lookupsgit log --oneline -1 should work

    Working directory should be the pi-crew repo root:

    cd ${PWD}
    ls package.json  # must exist
    

    CI integration

    The skill maps to existing CI gates as follows:

    CI gateSkill tierFile
    npm test:critical (manual / pre-commit)Tier 1n/a — not in CI by default
    PI_CREW_BROKER=0 npm run test:criticalTier 2 (env kill switch path)n/a — manual
    npm run typecheckTier 3.github/workflows/*.yml (every PR)
    Bundle-staleness checkTier 3 last stepscripts/check-bundle-staleness.mjs
    Multi-OS CIn/a (skill is local).github/workflows/*.yml — Linux + macOS + Windows
    Full npm test (>5 min)n/a — too slow for in-loopCI only

    To add Tier 1 to a pre-commit hook:

    # .git/hooks/pre-commit (or via husky / pre-commit framework)
    npm run test:critical || {
      echo "✋ test:critical failed — fix before commit"
      exit 1
    }
    

    To add Tier 1 to CI as a fast-feedback gate (under 30s):

    # .github/workflows/fast.yml
    - name: Critical unit tests
      run: npm run test:critical
    - name: Disabled-path proof
      run: PI_CREW_BROKER=0 npm run test:critical
    - name: Explicit-on proof
      run: PI_CREW_BROKER=1 npm run test:critical
    

    Tier 1 — Critical unit tests (~25s, 101 tests, the only suite you need for broker/UI changes)

    What: run the curated 14-file fast subset.

    Why this exists: full npm run test:unit runs 642 files, >4 minutes. Verifier worker timeout is 300s → worker killed mid-run, run = "hang". The fix (introduced in commit 1cb2dca) splits out a test:critical subset covering exactly what changed in the broker/UI work.

    How:

    time npm run test:critical
    

    Expected output: # tests 101 # pass 101 # fail 0 # duration_ms ~26000. (Count was 97 at v0.9.46; 101 since the model-routing merge — verify with the actual run; the skill's hard-coded numbers drift between releases.)

    References:

    WhatWhere
    Script definitionpackage.json:67 — list of 14 files passed to node scripts/test-runner.mjs
    Introduced in commit1cb2dca fix(verifier): use test:critical instead of test:unit to avoid worker timeout
    Runner wrapperscripts/test-runner.mjs — injects --test-force-exit, forwards to tsx --test
    The 14 filesbroker: crew-broker-{handshake,stale-socket,feature-flag,server-gate,client-fallback,mailbox-observer,close-during-reconnect,steer-dedup,symlink-steering}.test.ts; UI: keybinding-map.parity.test.ts, pi-tui-dispatch-probe.test.ts, session-utils-extract.test.ts; config: config-schema-sync.test.ts, child-pi-env-spread.test.ts
    Failure mode that motivates itWorker timeout in src/runtime/child-pi-constants.ts:23 (RESPONSE_TIMEOUT_MS = DEFAULT_CHILD_PI.responseTimeoutMs = 300000); verifier LLM ran npm test and got killed at 300s with exit 143 (SIGTERM)

    Run after: any edit to src/runtime/crew-broker*.ts, src/ui/, src/config/, src/extension/registration/lifecycle-handlers.ts, or src/runtime/child-pi-spawn.ts.


    Tier 2 — Three-path kill-switch proof

    What: prove all three precedence paths in effectiveEnabled() still resolve correctly.

    Why: any change to DEFAULT_BROKER (in src/config/defaults.ts:169) or effectiveEnabled() (in src/extension/registration/lifecycle-handlers.ts:819-833) can silently break the precedence chain. The chain:

    PI_CREW_BROKER=0     → disabled (env always wins)
    broker.enabled=false → disabled (config)
    PI_CREW_BROKER unset → enabled (DEFAULT_BROKER=Phase 4 default-on)
    PI_CREW_BROKER=1     → enabled (explicit; redundant under default-on)
    

    How:

    # 1. default path (whatever DEFAULT_BROKER.enabled is right now)
    npm run test:critical
    # 2. env kill switch
    PI_CREW_BROKER=0 npm run test:critical
    # 3. env explicit-on (must still work under default-on)
    PI_CREW_BROKER=1 npm run test:critical
    

    All three must show # pass 101 # fail 0. Measured times in this session (2026-08-11): ~26s for default, ~26s for PI_CREW_BROKER=0, ~26s for PI_CREW_BROKER=1 (varies ±1-2s run-to-run).

    References:

    WhatWhere
    DEFAULT_BROKER constantsrc/config/defaults.ts:169-173 (Phase 4: enabled: true)
    Precedence functionsrc/extension/registration/lifecycle-handlers.ts:819-833 (return cfg?.enabled !== false; at line 828)
    resolveBrokerEnvOverridesrc/config/defaults.ts:186-193
    Env-precedence unit teststest/unit/crew-broker-feature-flag.test.ts:31 (default-on assertion), :54-110 (env=1/env=0/unset/arbitrary cases at lines 54, 66, 78, 90, 103)
    Controller-gate teststest/unit/crew-broker-server-gate.test.ts:78 (env kill switch under default-on), :143 (env=1 with no config)
    Decision docdocs/decisions/2026-07-22-broker-phase4-gated-on.md
    Superseded docdocs/decisions/2026-07-21-broker-phase4-default-on.md (marked SUPERSEDED in commit 4186284)
    Default flip commit612e18b feat(broker): Phase 4 gated ON — flip broker.enabled default to true

    Tier 3 — Typecheck + bundle rebuild + md5 sync

    What: prove the bundle actually contains the source you just edited.

    How:

    npm run typecheck    # ~20s, exits 0 with "strip-types import ok"
    npm run build:bundle # <1s, prints "[build-bundle] dist/index.mjs NNNN KB in NNN ms"
    md5sum dist/index.mjs
    

    Compare the printed md5 against what the user's Pi session loaded. If they differ → the session is running stale bundle.

    References:

    WhatWhere
    typecheck scriptpackage.json "typecheck" — runs tsc --noEmit && node --experimental-strip-types -e "await import('./index.ts'); ..."
    build:bundle scriptpackage.json "build:bundle" — runs node scripts/build-bundle.mjs
    Bundle builderscripts/build-bundle.mjs (esbuild-based, bundles index.bundle.tsdist/index.mjs)
    Bundle resolution ruleindex.ts:5-22 (entrypoint docstring); also scripts/build-bundle.mjs:14-20 (entrypoint preference); symlink is live for source files but the bundled dist/index.mjs is loaded
    Postinstall hookscripts/postinstall.mjs:43 — best-effort bundle rebuild; falls back to strip-types if esbuild missing
    Bundle md5 after Phase-4 commit1cc4d55e18add7b9a036c569143320b6 (~2.78 MB at the time; check current: md5sum dist/index.mjs. As of v0.9.66 I-batch 2026-08-11: 16e29d053bd370e24f40df147dadcb79 ~2.81 MB)

    Tier 4 — Bundle sync into a live Pi session

    What: ensure the user's running Pi sees your changes.

    The immediate-vs-rebuild rule (which edits take effect without a rebuild):

    • workflows/*.workflow.md edits → immediate, no rebuild, no restart
    • src/runtime/plan-templates.ts taskTemplate strings → immediate, runtime data
    • Everything else (src/ edits, package.json) → must npm run build:bundle THEN user /quit + reopen Pi

    How to verify in this session:

    md5sum dist/index.mjs
    # then in Pi session, the user runs `md5sum` in a shell tool
    # if they differ, user needs to /quit + reopen
    

    How to verify in a fresh pty/tmux session without disturbing the user's main Pi:

    tmux -S /tmp/sock new-session -d -x 160 -y 50 -s pi \
      "cd ${PWD} && exec pi 2>&1"
    

    References:

    WhatWhere
    Bundle resolutionindex.ts:5-22 — "dist/index.mjs (pre-built bundle) if present AND not explicitly disabled — DEFAULT since v0.9.17"
    Bundle size impact after Phase-4 flipdocs/decisions/2026-07-22-broker-phase4-gated-on.md §Verification: "2.78 MB before and after the flip; the broker code was already in the bundle; only the default boolean changed"
    Symlink confirmationThe symlink lives in the CONSUMING project, not inside pi-crew itself. From the pi-crew repo, check the parent: readlink ../node_modules/pi-crew (returns ../pi-crew for dev clones). For global installs: readlink "$(npm root -g)"/pi-crew. Pattern is always <consumer>/node_modules/pi-crew → <pi-crew-repo>.

    Tier 5 — Live TUI probe via tmux send-keys

    What: drive a real Pi session's keystrokes from the shell, capture screen state.

    Why tmux and not raw pty: tmux gives you a clean separation — session persists across your bash commands, capture-pane gives ASCII screenshot, send-keys with hex escapes covers \x1b[A (legacy CSI), \x1bOA (app-cursor-mode), and Kitty-protocol variants.

    How:

    # Spawn (160x50 fits ~standard TUI)
    tmux -S /tmp/sock new-session -d -x 160 -y 50 -s pi \
      "cd ${PWD} && exec pi 2>&1"
    
    # Wait for pi to start
    sleep 2
    
    # Send slash command
    tmux send-keys -t pi '/team-help' Enter
    sleep 1
    tmux capture-pane -t pi -p | tail -40
    
    # Send raw escape sequence (app-cursor-mode up arrow)
    tmux send-keys -t pi $'\x1bOA'
    sleep 0.5
    tmux capture-pane -t pi -p > /tmp/screen-after-up.txt
    

    Key gotcha: terminals send arrow keys as one of 3 byte sequences. pi-crew's matchesKey() helper (src/ui/key-utils.ts:37-42, the keyOf() function) normalizes all of them — but verify it does in your probe:

    ModeUp arrowDown arrowSource
    Legacy CSI\x1b[A\x1b[Bvt100, xterm
    App-cursor-mode\x1bOA\x1bOBvim, less, full-screen apps
    Kitty protocol\x1b[1;2A (Shift+Up) etc.modern terminals (kitty, foot, ghostty)

    References:

    WhatWhere
    keyOf() helpersrc/ui/key-utils.ts:37-42 (import + type alias at lines 16-18)
    Dispatch pathsrc/ui/keybinding-map.ts (migrated to matchesKey() in commit f05a10d)
    Golden snapshot testtest/unit/keybinding-map.parity.test.ts — 7 it() blocks asserting parity against a generated golden snapshot; BINDINGS table has 27 entries (src/ui/keybinding-map.ts:132-180)
    Live probe testtest/unit/pi-tui-dispatch-probe.test.ts — direct probe of dispatch (3 tests)
    Probe commit84944f7 test(probe): add invalidate() to control object so typecheck passes
    Tab/Space bindsrc/ui/run-dashboard.ts + commit 15a0ffe fix(ui): also bind Tab/Space/Enter/S to select in dashboard dispatch
    Tmux session file/tmp/sock (created on first new-session -S)

    Tier 6 — Live TUI probe via Python pty (bulk keys + diag)

    What: send many keys in sequence + capture per-keystroke diag output.

    When to use: when you need to probe dispatch across multiple keypresses, or want to verify each key reached the component's handleInput.

    How (simplified inline example — for the full hardened script with zombie reaping, non-blocking read, and escape-sequence decoding, use scripts/pty_probe.py directly):

    #!/usr/bin/env python3
    """pty_probe.py — bulk-key + diag probe for pi-crew TUI components."""
    import os, sys, time
    
    CMD = ['pi']
    ENV = dict(os.environ)  # keystroke diag env var REMOVED (see note below)
    
    pid, fd = pty.fork()
    if pid == 0:
        os.execvpe(CMD[0], CMD, ENV)
    else:
        time.sleep(2)  # initial pi startup
        keys = [
            'j', 'j', 'k',                      # vim nav (run dashboard)
            '\x1b[A',                            # legacy CSI up
            '\x1b[B',                            # legacy CSI down
            '\x1bOA',                            # app-cursor-mode up
            '\x1bOB',                            # app-cursor-mode down
            'q', 'q',                            # quit (double-tap)
        ]
        for k in keys:
            os.write(fd, k.encode())
            time.sleep(0.3)
        time.sleep(1)
        sys.stdout.write(os.read(fd, 65536).decode(errors='replace'))
    

    ⚠️ The inline code above is a teaching example. For real use, run the bundled script (scripts/pty_probe.py, 161 lines) which adds zombie reaping (_reap_child), non-blocking read (select.select with 5s timeout), exec error handling, and --keys escape-sequence decoding:

    python3 scripts/pty_probe.py [--keys '\x1bOA,q,q'] [--cwd /path] [--startup-sleep 3]
    

    The inline code works for a quick one-off but leaks a zombie pi process on exit.

    Keystroke diag env var REMOVED (2026-08-10): PI_CREW_BROKER_DIAG_UI=1 made run-dashboard's handleInput write a [PI-CREW-DIAG] line to stderr per keystroke. It was removed in e3ee6fe2 (PR-B5: remove TEMP DIAGNOSTIC from run-dashboard, UI-8) — there is no replacement in src/. To prove keystroke arrival now, rely on screen-change evidence (Tier 5 tmux capture-pane before/after each key, or the pty output diff): a key that changes screen state reached the TUI; a key that does not was consumed or never arrived. Capture the probe output to a file with 2>&1 | tee /tmp/pty-probe.log and diff the rendered frames.

    References:

    WhatWhere
    Keystroke diag env varREMOVEDe3ee6fe2 (PR-B5/UI-8). No replacement; use screen-change evidence
    Reduced-noise commit00e8ba0 chore(broker): strip diagnostic noise from focused-field fix — diag calls left in but no longer noisy (pre-removal)
    Original probe84944f7 test(probe): add invalidate() to control object so typecheck passes

    Tier 7 — Smoke team run (verifier prompt doesn't hang)

    What: prove the verifier worker completes within RESPONSE_TIMEOUT_MS (300s).

    Why this is its own tier: test:critical covers unit-level invariants, but the verifier LLM is a separate failure mode — it reads the verifier prompt from src/runtime/plan-templates.ts:143, 190 (taskTemplate strings) or from workflows/*.workflow.md:24, 30, 31 (workflow verifier sections), then decides which bash command to run. If the prompt says "Run tests" without specifying which, the LLM runs npm test and the worker hangs at 300s with exit 143.

    How (from parent Pi session — team is a tool, not a shell command):

    # illustrative — the actual tool takes positional + named params:
    #   team action='run' team='fast-fix' workflow='fast-fix' goal='...' async=false
    team:
      action: run              # run | status | events | cancel | retry | ...
      team: fast-fix           # team (a role-set): default / fast-fix / implementation / parallel-research / research / review
      workflow: fast-fix       # workflow (a phase DAG): default / fast-fix / plan-execute / implementation / review / research / parallel-research / pipeline / chain
      goal: "Smoke-verify <X>. Run `npm run test:critical && npx tsc --noEmit` once, cache output, report exact pass/fail counts + total time. Confirm verifier completes without hang (must be <300s)."
      async: false             # synchronous: wait for completion before returning
    

    The team tool is described in the agent's system prompt. Use team action='status' <runId> to inspect mid-run, team action='events' <runId> <limit> for the event log, team action='cancel' <runId> to abort.

    Real measured outcomes from this session:

    Run IDGoalResultWall-clock
    team_20260722083504_cae04a2804a24d79smoke full-implementation3/4 phases, 04_verify hung on npm test572s
    team_20260722095143_2e58fce2ce91af19first smoke-fix smoke3/3 PASS, verifier used fast path but ran multiple LLM turns (think→bash→observe→respond) totaling ~907s cumulative907s
    team_20260722100811_9bf95bebff2b052are-smoke after workflow prompt fix3/3 PASS, verifier used test:critical cache449s

    References:

    WhatWhere
    verificationCommand for plan-templatessrc/runtime/plan-templates.ts:146, 193 — both templates now npm run test:critical && npx tsc --noEmit
    taskTemplate for verifiersrc/runtime/plan-templates.ts:143, 190 — explicit "Do NOT run npm test" + "<2 min" budget
    Workflow verifier promptsworkflows/fast-fix.workflow.md:24, workflows/default.workflow.md:31, workflows/plan-execute.workflow.md:30, workflows/review.workflow.md:31
    Verifier fix commit (plan-templates)1cb2dca fix(verifier): use test:critical instead of test:unit to avoid worker timeout
    Verifier fix commit (workflows)d599578 fix(workflows): specify fast test:critical command in verifier prompts
    Watchdog constantsrc/runtime/child-pi-constants.ts:23RESPONSE_TIMEOUT_MS = DEFAULT_CHILD_PI.responseTimeoutMs
    Cache directiveRun FAST checks ONCE (cache output to .crew/cache/) — anti-re-run safeguard baked into all 4 workflow verifier prompts
    Decision docdocs/decisions/2026-07-22-broker-phase4-gated-on.md §Verification (mentions the smoke run team_20260722100811_9bf95bebff2b052a)

    Two known failure modes for verifier:

    1. Verifier LLM runs npm test (full unit + integration suite, >4 min) instead of npm run test:critical. Symptom: worker killed with exit 143 after exactly 300s. Fix: rewrite the verifier prompt to specify the exact fast command AND include "Do NOT run npm test or npm run test:unit".
    2. Verifier LLM improvises with a clean-cache npm test run anyway. The cache directive ("cache to .crew/cache/", "do NOT re-run") catches this — the second worker that observes a cached log should not re-run.

    Tier 8 — Bundle-vs-session md5 sync (operational check)

    What: the concrete md5 comparison step that proves Tier 4's claim. Tier 4 explains when you need a rebuild; Tier 8 is the command you run to confirm the session picked it up. Run Tier 8 as the final integrity check after Tier 3-4.

    How:

    # Disk
    md5sum dist/index.mjs
    
    # Session (ask user to run in their pi shell tool)
    # The symlink is in the CONSUMING project, not inside pi-crew:
    readlink ../node_modules/pi-crew/dist/index.mjs 2>/dev/null \
      || readlink "$(npm root -g)"/pi-crew/dist/index.mjs \
      || md5sum "$(npm root -g)"/pi-crew/dist/index.mjs
    # (the consuming project loads pi-crew via this symlink — see index.ts:5-22)
    

    If the two md5s match → session is on the latest code. If not → user must /quit + reopen Pi.

    Agent-inside-session caveat: when the agent doing the testing runs inside the very Pi session under test, the agent cannot restart its own session — only the user can. Pattern that works: (1) edit source + rebuild bundle, (2) ask the user to /quit + reopen, (3) on resume, re-check md5sum dist/index.mjs then issue a probe tool call (e.g. team action='list'). If the probe returns the old error (e.g. Unknown type, Validation failed for tool team), the session did NOT reload — there may be multiple pi PIDs and the user reopened a different one. Verify with ps -eo pid,lstart,tty,args | grep pi which PID is yours (the one whose session log is being appended to right now).

    References:

    WhatWhere
    Symlink pathindex.ts:5-22the symlink lives in the CONSUMING project (parent dir or global prefix), not inside pi-crew itself. From the repo: readlink ../node_modules/pi-crew (dev) or readlink "$(npm root -g)"/pi-crew (global). Verify with readlink + npm root -g.
    Session load modelSame file: "dist/index.mjs (pre-built bundle) if present — DEFAULT since v0.9.17"

    Tier 9 — Feature battery (live action coverage)

    What: drive the team tool + subagent tools through a spread of actions from the parent Pi session to prove the full surface works end-to-end, not just one smoke run.

    Why this exists: Tier 7 proves one team run completes. But pi-crew has ~50 team actions plus 4 subagent tools (Agent, crew_agent, get_subagent_result, crew_agent_steer), dispatched through several code paths (sync run, async run, chain, parallel, direct subagent). A schema or registration regression can break some paths while others still pass. The battery catches path-specific breakage.

    When required: any change to src/schema/team-tool-schema.ts, src/extension/registration/team-tool.ts, src/extension/team-tool/*.ts (handler dispatch), or the subagent-tool registration. Optional but cheap for any change — the read-only actions are free.

    How (run from the parent Pi session — these are tool calls, not shell):

    1. 9a. Read-only actions (free, no subagent spawn — run these first as a fast battery):
      • team action='list' — teams/workflows/agents
      • team action='recommend' goal='...' — planner routing
      • team action='health' — run-state scan
      • team action='doctor' focus='zombies' — orphan subagent scan (read-only)
      • team action='status' runId='<recent>' details=false — compact
      • team action='events' runId='<recent>' — full event lifecycle
      • team action='summary' runId='<recent>' — cost/by-role report
      • team action='get' resource='workflow' team='implementation' — resource inspect
      • team action='explain' runId='<recent>' — markdown render
      • team action='worktrees' runId='<recent>' — workspace listing
    2. 9b. Spawn paths (cost tokens — one probe each is enough):
      • team action='run' sync (fast-fix, trivial goal) — proves sync run + child-pi spawn + provider-extension loading
      • team action='run' async=true — proves background dispatch
      • team action='run' chain='"A" -> "B"' — proves sequential handoff (chain runner). Omit workflow — passing workflow:'chain' forwards it to each step and fails fast (~58ms silent; issue #44).
      • Agent direct subagent — proves the direct-subagent tool
      • crew_agent run_in_background=true then get_subagent_result — proves background subagent lifecycle
    3. Acceptance: every action returns without Unknown type / Validation failed for tool team / empty error text; every spawn path completes with consistency=1 and the expected probe token in the agent output.

    Real measured outcome (this session, after the v0.9.57 schema fix): 9a (15 team actions) + 9b (4 subagent tools / 3 run paths) exercised; all green; the two silent-failure modes that motivated this tier (Unknown type from Type.Unsafe without Kind, and Validation failed for tool team from empty-string-strict schema) were caught ONLY by this battery — Tier 1-8 all passed while the team tool was broken live. The session also surfaced the unauthorized-agent-edit anti-pattern (a chain-run agent edited chain-runner.ts mid-smoke) — see Anti-patterns.

    Not covered by the cheap battery above — the actions below need extra setup, cost, or user confirmation. Run them only when the change touches their code path, and prefer a throwaway cwd / config so you don't mutate the user's real state. Organised by cost/safety. As of 2026-08-11 (extended battery, run report real-test-2026-08-11-scratchpad-I-batch.md), 9c/9e/9f have been exercised live once each — they are no longer unproven, but still require explicit scope+confirmation to re-run.

    9c. Lifecycle / recovery (needs a running run — start an async run, then exercise these against its runId):

    • team action='wait' runId='...' — block until completion
    • team action='steer' runId='...' message='...' — inject a steering note mid-run
    • team action='status' runId='...' details=true — full dump mid-run
    • team action='cache' subAction='...' runId='...' — snapshot cache ops
    • team action='checkpoint' runId='...' — state checkpoint
    • team action='cancel' runId='...' — ⚠️ destructive (kills the run); use a throwaway run
    • team action='invalidate' runId='...' — cache invalidation
    • team action='resume' runId='...' / retry — resume a completed/failed run
    • team action='respond' taskId='...' message='...' — mailbox reply (needs a waiting task)
    • subagent steering: crew_agent run_in_background=true a long task (e.g. sleep 60), then crew_agent_steer while it runs, then get_subagent_result — proves the steer arrived (timing-sensitive; assert the agent's output reflects the steer)

    9d. Destructive (⚠️ requires explicit user confirmation per the delegation policy — never run unprompted):

    • team action='prune' keep=<N> — delete old finished runs
    • team action='cleanup' — sweep stale workspaces/state
    • team action='forget' runId='...' — delete one run's state
    • team action='doctor' focus='zombies' is READ-ONLY (safe) but the follow-up kill <PID> it suggests is destructive — confirm with the user before killing

    9e. Admin / mutation (mutates config or workflow files — use a scratch project cwd or back up first):

    • team action='create' resource='team' ... / update / delete — manage teams/agents/workflows
    • team action='init' / config / validate / autonomy / settings — project setup
    • team action='workflow-create' / workflow-save / workflow-delete / workflow-get / workflow-list — workflow CRUD
    • team action='import' / imports / export — run data portability
    • team action='parallel' tasks=[...] — parallel dispatch (spawn path, costs tokens per task)

    9f. Background / scheduled (expensive or niche):

    • team action='run' runKind='goal-loop' — the goal loop runs many turns judging an objective; smoke with a trivial objective + low maxTurns (e.g. 2) to prove dispatch without burning budget
    • team action='schedule' cron='...' ... / scheduled / subAction='remove' — cron; assert the job registers then remove it (cleanup)
    • team action='auto-summarize' / anchor / auto_boomerang — background features; assert no-throw on a completed run
    • team action='api' — programmatic surface

    Acceptance for 9c–9f: the action returns a structured result (not Unknown type / not an empty error), and for spawn/lifecycle paths the run reaches the expected terminal status. For 9d/9e, the mutation is reversible or confined to scratch state.


    Anti-patterns (the cost is real, observed in this session)

    Anti-patternCostWhere fixedReference
    npm test in verifier prompt300s worker timeout, run = "hang"1cb2dcasrc/runtime/plan-templates.ts:143, 190 + 4 workflow files
    npm run test:unit for in-loop verify>4 min, same hang1cb2dcapackage.json:67 (test:critical script)
    Default-off assumption in testsBreak when default flips612e18btest/unit/crew-broker-feature-flag.test.ts:31 (DEFAULT_BROKER.enabled === true)
    Test using real loadConfig() to mock configFlaky when env / disk config changes612e18btest/unit/crew-broker-server-gate.test.ts:78 (use brokerEnv: "0" instead of flagOn: false)
    Source edit seen immediatelyNo, requires bundle rebuild + reloadn/a (permanent)index.ts:5-22 — bundle resolution rules
    Skip disabled-path proofeffectiveEnabled() regression slips throughn/a (permanent)Tier 2 above
    npm run test:unit against 642 files>4 min; mis-judges verifier runtimen/a (permanent)Tier 1 above
    Skip typecheckTS errors slip past test:critical (which uses --test-timeout=30000)n/a (permanent)Tier 3 above
    Run pi from a stale bundleSession shows old behavior despite src/ editsn/a (permanent)scripts/check-bundle-staleness.mjs — CI gate
    Test by reading codeProves nothing about runtimen/a (permanent)All tiers above
    makeFakeCtx({ flagOn: false }) without brokerEnv: "0"makeFakeCtx deletes PI_CREW_BROKER env if brokerEnv is undefined612e18b (test fix)test/unit/crew-broker-server-gate.test.ts:78 — pass brokerEnv: "0" to preserve env
    Trust green CI on one OSmacOS/Windows regressions slip throughn/a (permanent).crew/knowledge.md — "CI runs 3 OSes ... A flake on one OS IS a real bug"
    Trusting a team-run agent not to edit the repo under testAgents spawned by team/Agent/crew_agent inherit the session cwd and have edit/write tools — a proactive LLM (observed with deepseek) will make unauthorized source edits to pi-crew during a trivial smoke run (e.g. "improving" chain-runner.ts while parsing a chain string). The edit can be correct + green-tested yet still be unintended scope creep that silently lands in your commit.n/a (permanent)After EVERY team/subagent run: git status and verify each changed file was authored by you. Diff + review any surprise change before staging. Consider workspaceMode: 'worktree' for parallel/risky runs to isolate mutations.
    Armed-role tool-surface bug (found live 2026-08-11): an opt-in tool (e.g. scratchpad) is armed via ROLE_TOOL_CONFIGS[role].scratchpad=true AND env PI_CREW_SCRATCHPAD=1, but NEVER appears in the worker surface. Root cause: resolveToolPolicy (src/agents/agent-config.ts:165) falls back roleConfig.tools ?? agent.tools when the role has no tools allowlist, and the builtin agents/{executor,verifier,test-engineer}.md frontmatter tools: did not list scratchpad → pi got --tools read,grep,find,ls,bash,edit,write and hard-filtered scratchpad. Env was correct; the tool was silently dropped by the --tools allowlist. Reproduce: pi -p --tools read,grep,find,ls,bash,edit,write "list tools" → no scratchpad; with scratchpad added → present.f753be30Fix: keep armed-role agents/*.md frontmatter tools: lists in sync with ROLE_TOOL_CONFIGS (QW17 pins the pinned roles; add the new tool to BOTH frontmatter AND role config for pinned roles, frontmatter-only for vacuous roles). A smoke run that claims the tool is "not available" in the worker is a REAL signal — verify the worker's actual --tools allowlist, not just env vars.
    Type.Unsafe({ anyOf/type }) schema field without [TypeBox.Kind] symbolValue.Check throws Unknown type the first time a model emits that field (e.g. skill, config) — every team action returns isError:true text "Unknown type". Tier 1-8 stay green because unit tests never send the offending field.v0.9.57src/schema/team-tool-schema.tsSkillOverride/FreeformConfig switched from Type.Unsafe to TypeBox-native Type.Union/Type.Record. See Tier 9.
    Schema too strict for model-emitted empty strings (runId:"", workspaceMode:"", budgetTotal:0)pi-ai validateToolArguments runs BEFORE the pi-crew handler and rejects "" against Literal unions / patterns → Validation failed for tool team → model loops.v0.9.57src/schema/team-tool-schema.ts — added Literal("") to unions, `^$
    Claiming "all 9 tiers pass" while 9c–9f were never runOverclaim — once reported "9 tiers pass" when only 9a (8/10) + 9b (4/5) had actually run; 9c–9f were skipped. Past runs then become unverifiable ("did it really pass 9 tiers?"). 2026-08-11 repeat: an initial report said "9c–9f skipped" yet the summary read as full coverage until the gap was called out.n/a (process)Fill REPORT-TEMPLATE.md per-tier DURING the run. "Tier 9 pass" = 9a AND 9b AND the applicable 9c–9f, each with evidence. Round-up-to-pass is the anti-pattern this row exists to prevent. If 9c–9f are skipped, SAY SO in the verdict and do not phrase it as "all pass".
    chain run with workflow:"chain" forwarded to stepsEvery chain step fails in ~58ms with an EMPTY error string — looks like a parse failure but isn't. chain-dispatch forwards params.workflow ("chain") into executor overrides; each step then runs the "chain" workflow via the normal executeTeamRun path and fails fast + silently.Open (issue #44)Omit workflow when invoking action:'run' chain=... — chain then runs 2/2 success (~308s). See docs/bugs/chain-workflow-forward-quirk.md.

    Failure symptoms + recovery

    When a tier fails, the recovery is usually quick. Match the symptom to the cause:

    SymptomLikely causeRecovery
    test:critical returns # fail N>0Regression in touched sourceRead the failing test's name + assertion; fix the source; rerun
    test:critical hangs >60sOne test opened a socket/pty that didn't closeRun individual file: node --import tsx/esm --test --test-force-exit test/unit/<file>.test.ts; check for missing await or unclosed handle
    typecheck fails with TS2xxxTS type drift after src/ editFix the type error; do not commit until exit 0
    build:bundle failsesbuild error in index.bundle.tsRun npx esbuild --bundle src/index.bundle.ts --outfile=dist/index.mjs for the verbose error
    md5sum dist/index.mjs differs from sessionStale bundle in user's PiUser must /quit + reopen Pi; new extension cold-start loads new bundle
    Tmux probe: keys not reaching componentWrong terminal encodingCheck pi-tui env; use both \x1b[A and \x1bOA; check matchesKey is wired in the dispatched class
    pty_probe.py errors OSError: [Errno 6] No such devicePty already closedReduce --startup-sleep or check pi actually launched
    Smoke team: 04_verify exits with 143Verifier ran slow command (typically npm test)Read worker transcript for actual command run; fix the verifier prompt per Tier 7
    Smoke team: worker times out at 300sEither verifier command slow OR LLM thinking capCheck RESPONSE_TIMEOUT_MS (300s); bump only if you verified the command itself finishes <300s
    stale-ctx error in worker outputExtension ctx is stale after session replacementThis is runtime noise, not a regression; ignore. (Source: .crew/knowledge.md "Process Safety" notes)
    Bundle md5 not changing after rebuildStale dist/ cache or esbuild no-oprm -rf dist/ && npm run build:bundle; verify new md5
    Team tool returns Unknown type (isError:true, short text)Value.Check in the handler hit a Type.Unsafe({...}) schema node with no [TypeBox.Kind] symbol — only triggered when the model actually sends that field. Tier 1-8 pass; only Tier 9 (feature battery) catches it.Replace the Type.Unsafe with a TypeBox-native constructor (Type.Union, Type.Record, Type.Any). Reproduce with node --input-type=module -e "import {Value} from '@sinclair/typebox/value'; import {TeamToolParams} from './src/schema/team-tool-schema.ts'; Value.Check(TeamToolParams, {action:'list', skill:'', config:{}})" — a throw = the bug.
    Validation failed for tool "team": ... must be equal to constantpi-ai validateToolArguments (@earendil-works/pi-ai/dist/utils/validation.js) rejects model-emitted ""/0/false defaults against Literal unions / patterns / minimums — it runs BEFORE the pi-crew handler, so handler-side normalization is too late.Loosen the schema to accept the unset marker (Literal(""), pattern `^$
    User says "restarted" but the probe still shows the OLD errorMultiple pi PIDs open; the user reopened a different terminal than the one the agent runs in; the agent's session never reloaded the bundle.ps -eo pid,lstart,tty,args | grep pi to list PIDs; match the agent's session log (the .jsonl being appended right now) to its PID; have the user reopen THAT session, or move the work into the freshly-opened one.

    Performance budget (per-tier soft limits)

    TierSoft limitHard limitWhat happens over hard limit
    1 (test:critical)25s60sWorker likely hung — cancel + bisect by file
    2 (3-path proof, total)75s180sSame as above
    3 (typecheck + build:bundle)25s60stypecheck regression — check imports
    4 (md5 sync check)<1s5sDisk/symlink issue
    5 (tmux spawn)5s15stmux server issue
    6 (pty probe)5s15spi not in PATH
    7 (smoke team)60s (verifier only)300s (worker hard limit)Worker killed by RESPONSE_TIMEOUT_MS
    8 (final md5 sync)<1s5sDisk/symlink issue
    9 (feature battery)30s (read-only batch) + ~120s per spawn probe300s per spawn probe (worker hard limit)Spawn probe hung or returned Unknown type/Validation failed — a schema or registration regression; see Tier 9 + Failure symptoms

    If a tier runs over the hard limit, stop and investigate — don't bump the budget silently. The budget exists precisely so regressions in test runtime (which usually means a regression in test setup/teardown) are caught early.


    Edge cases

    macOS specifics

    TopicLinuxmacOSAction
    md5sumyesno (use md5 -r)The Prerequisites table notes this.
    XDG_RUNTIME_DIR/run/user/<uid>unset by defaultpi-crew falls back to os.tmpdir() (per-user /var/folders/.../T/). Broker works the same.
    Unix abstract socketyesnoThe broker uses concrete paths under $XDG_RUNTIME_DIR, so it works on both.
    tmuxusually preinstalledbrew install tmuxSame commands; the pty_probe.py works on both.
    /tmp/socktmpfs/tmp is nodeboot-protected (cleared on reboot but not on logout)Same.

    Non-standard paths

    The skill assumes pi-crew is at ${PWD} (the directory you cd'd into). If you have it elsewhere:

    export PI_CREW_ROOT=/path/to/pi-crew
    cd $PI_CREW_ROOT
    # Now ${PWD} resolves correctly inside the skill
    

    The cd ${PWD} calls appear in the Prerequisites section, Tier 4, Tier 5, and the Quick reference section — all use the same path. Once you cd into the repo once, all commands that reference ${PWD} resolve correctly. Tier 6 uses scripts/pty_probe.py --cwd instead, and Tier 8 uses readlink (no cd needed).

    No-tmux fallback

    If tmux is not installed, use Tier 6 (Python pty) instead. Tier 6 doesn't depend on tmux; it spawns pi directly under a pty. The trade-off: Tier 5 gives you capture-pane for ASCII screenshots; Tier 6 gives you per-keystroke diag output.

    Stale /tmp/sock (tmux session already exists)

    If a previous Tier 5 run left a /tmp/sock server running, tmux new-session -S /tmp/sock will reuse it instead of creating a fresh session. The new pi instance attaches to the existing session, which may have leftover state. To force a fresh session:

    tmux -S /tmp/sock kill-server 2>/dev/null  # clean up
    tmux -S /tmp/sock new-session -d -x 160 -y 50 -s pi "cd ${PWD} && exec pi 2>&1"
    

    Multiple concurrent Pi sessions

    When the user has multiple Pi sessions open (e.g., main + scratch), each loads the same dist/index.mjs. The md5sum check is global — if any session loaded the old bundle, you need to restart ALL of them, not just the one you're testing in. Tier 8 covers this only for the user's "main" Pi; warn them about siblings.

    Broker on Windows

    broker.enabled=true is silently no-op on native Windows (no unix-domain socket). Users on WSL1/2 get full broker behavior. Don't waste time running Tier 7 smoke tests on native Windows — the verifier will run fine but the broker won't actually do anything. Use PI_CREW_BROKER=0 to skip the broker entirely.


    Cross-skill notes

    This skill overlaps with these built-in/project skills. Pick the right one:

    SkillWhen to use instead
    test (built-in)When you want generic test execution guidance (not pi-crew-specific)
    lint (built-in)When you only need lint + format (Tier 3's typecheck replaces it for TypeScript)
    verify-before-complete (project)When claiming "done" without specific tier discipline; this skill's Tier 1-8 are stricter and pi-crew-specific
    code-optimizer (built-in)When auditing for perf, not for verification
    iterative-audit (project)When doing a multi-round codebase audit; this skill's "review kỹ" rounds are a different beast — they're verification, not audit
    review / security-review (built-in)When reviewing someone else's PR diff; this skill is for verifying YOUR OWN changes

    The "skill stack" for a typical pi-crew change:

    1. Edit src/
    2. tier 1 (test:critical)        ← this skill
    3. tier 2 (3-path proof)         ← this skill, if broker change
    4. tier 3 (typecheck + bundle)   ← this skill
    5. tier 5/6 (live TUI)           ← this skill, if ui change
    6. tier 7 (smoke team)           ← this skill, if plan/workflow change
    7. commit + push
    8. verify-before-complete        ← make the "done" claim with evidence
    

    Maintenance

    The skill mentions specific commits, line numbers, and version pins. As the code evolves, these will drift. Maintenance playbook:

    WhatWhenHow
    Verify line refs after each src/ commitEvery commit touching the cited filegit log -p -- src/extension/registration/lifecycle-handlers.ts | grep effectiveEnabled — if line moved, update the skill
    Verify commit hashes still existQuarterly or before major editsgit log --oneline -1 <hash> — if gone, find the equivalent newer commit
    Verify version pins (v0.9.46, etc.)Each releasegit log --oneline -- src/ui/run-dashboard.ts | head -5 — confirm diag removal history (e3ee6fe2) still accurate
    Verify test:critical still has 14 filesEach src/runtime/crew-broker*.ts editcat package.json | grep test:critical — adjust the file list
    Verify Tier 7 verifier prompts still say test:criticalEach workflow file editgrep "Run FAST checks" workflows/*.workflow.md

    The skill does NOT need to be updated for every commit — only when the cited lines/files move. Consider it a "living reference" not a "live spec".


    Quick reference — exact commands

    # Tier 1 (critical unit, ~25s, 101 tests)
    npm run test:critical
    # Tier 2 (3-path proof, broker changes only)
    PI_CREW_BROKER=0 npm run test:critical
    PI_CREW_BROKER=1 npm run test:critical
    # Tier 3 (compile + bundle)
    npm run typecheck
    npm run build:bundle
    md5sum dist/index.mjs
    # Tier 4 (sync check — symlink is in the CONSUMING project)
    readlink ../node_modules/pi-crew  # dev: → ../pi-crew
    readlink "$(npm root -g)"/pi-crew  # global install
    # Tier 5 (tmux probe)
    tmux -S /tmp/sock new-session -d -x 160 -y 50 -s pi \
      "cd ${PWD} && exec pi 2>&1"
    tmux send-keys -t pi '<key>' ; sleep 0.5
    tmux capture-pane -t pi -p
    # Tier 6 (pty probe)
    python3 scripts/pty_probe.py 2>&1 | tee /tmp/diag.log
    # Tier 7 (smoke team)
    # from parent Pi session only — uses the `team` tool, not shell
    # Tier 8 (final md5 sync — compare disk vs loaded bundle)
    md5sum dist/index.mjs
    md5sum "$(npm root -g)"/pi-crew/dist/index.mjs 2>/dev/null \
      || md5sum ../node_modules/pi-crew/dist/index.mjs
    # Tier 9 (feature battery — from parent Pi session, tool calls not shell)
    #   read-only: team action=list / recommend / health / doctor / status / events / summary / get / explain / worktrees
    #   spawn:     team action=run (sync) ; team action=run async=true ; team action=run chain='"A" -> "B"'
    #              Agent (direct) ; crew_agent run_in_background=true + get_subagent_result
    #   reproduce the two silent schema failures:
    #   node --input-type=module -e "import {Value} from '@sinclair/typebox/value'; import {TeamToolParams} from './src/schema/team-tool-schema.ts'; Value.Check(TeamToolParams, {action:'list', skill:'', config:{}})"  # throws 'Unknown type' = Type.Unsafe-without-Kind bug
    

    Done-criteria checklist

    Before claiming "tested":

    • Tier 1: test:critical fresh-run, all pass (<25s). Count varies by release — was 97 at v0.9.46, 101 since the model-routing merge (v0.9.66); record the actual count in the report.
    • Tier 2: 3-path proof all pass — required if you touched src/config/defaults.ts or src/extension/registration/lifecycle-handlers.ts
    • Tier 3: npm run typecheck exit 0, npm run build:bundle exit 0
    • Tier 4: bundle md5 matches what the session loaded (or user has /quit-ed + reopened)
    • Tier 5/6: live TUI smoke for any src/ui/ change — keystroke reached handleInput
    • Tier 7: smoke team run for any src/runtime/plan-templates.ts or workflows/*.workflow.md change — completed, no hang, verifier output under 60s
    • Tier 8: final md5 sync check passed
    • Tier 9: feature battery — required if you touched src/schema/team-tool-schema.ts, src/extension/registration/team-tool.ts, any Type.Unsafe({...}) schema, or any armed-role tool list (agents/*.md / src/config/role-tools.ts). 9a read-only batch all return clean; one probe per 9b spawn path (sync / async / chain / Agent / crew_agent+get_subagent_result) completes with consistency=1. Run 9c–9f only when the change touches their code path; at least one full 9c/9e/9f sweep per release is recommended so the battery stays proven (see real-test-2026-08-11-scratchpad-I-batch.md); 9d (destructive) requires explicit user confirmation. After every run: git status to catch unauthorized agent edits.
    • Output report: save docs/real-test/reports/real-test-<YYYY-MM-DD>-<slug>.md from skills/real-test-pi-crew/REPORT-TEMPLATE.md, filled DURING the run with per-tier evidence (counts/md5/runId) — not reconstructed from memory afterward. This is what makes past runs verifiable instead of trust-the-summary.

    "All 9 tiers pass" is a claim that needs per-row evidence. Tier 9 means 9a and 9b and whichever of 9c–9f applies to the change — not "9a passed, therefore 9 passed". If any required item above is unchecked or lacks concrete evidence (a number, an md5, a runId), the answer to "is it tested?" is no — say so explicitly instead of rounding up to "pass".


    File-anchored references (full index)

    Decision docs:

    • docs/decisions/2026-07-21-broker-phase4-default-on.md — interim default-off (SUPERSEDED)
    • docs/decisions/2026-07-22-broker-phase4-gated-on.md — default-on flip + risk + monitoring + rollback
    • docs/decisions/2026-07-21-broker-windows-perms.md — Windows named-pipe perms + Phase-4 update note

    Source files (critical paths):

    • src/config/defaults.ts:155-187DEFAULT_BROKER + resolveBrokerEnvOverride
    • src/extension/registration/lifecycle-handlers.ts:819-833effectiveEnabled() (precedence)
    • src/runtime/child-pi-constants.ts:23RESPONSE_TIMEOUT_MS = 300_000
    • src/runtime/plan-templates.ts:143, 146, 190, 193 — verifier taskTemplate + verificationCommand
    • src/runtime/crew-broker.ts — broker server (per-connection gate, NDJSON framing)
    • src/runtime/crew-broker-client.ts — client (isEventFrame() distinguishes event vs response frames)
    • src/runtime/crew-broker-tokens.tsBrokerTokenRegistry with timingSafeEqual
    • src/runtime/broker-issuer.ts — per-run broker issuer (env injection at spawn)
    • src/runtime/crew-broker-child.ts — child-side broker client wiring
    • src/ui/key-utils.ts:37-42keyOf() using pi-tui matchesKey()
    • src/ui/keybinding-map.ts — dispatch using matchesKey() (commit f05a10d)

    Test files (the 14 in test:critical):

    • test/unit/crew-broker-{handshake,stale-socket,feature-flag,server-gate,client-fallback,mailbox-observer,close-during-reconnect,steer-dedup,symlink-steering}.test.ts
    • test/unit/keybinding-map.parity.test.ts
    • test/unit/pi-tui-dispatch-probe.test.ts
    • test/unit/session-utils-extract.test.ts
    • test/unit/config-schema-sync.test.ts
    • test/unit/child-pi-env-spread.test.ts

    Integration tests (Tier 1 covers none — these are for full E2E):

    • test/integration/crew-broker-msg.test.ts — 5 tests (Phases 1)
    • test/integration/crew-broker-phase2-3.test.ts — events.subscribe + task.waitStatus + steer.push + escalate

    Workflow files:

    • workflows/fast-fix.workflow.md:24 — verifier prompt (commit d599578)
    • workflows/default.workflow.md:31 — verifier prompt
    • workflows/plan-execute.workflow.md:30 — verifier prompt
    • workflows/review.workflow.md:31 — verifier prompt

    Commits (chronological, the patterns they introduced):

    • 1cb2dcatest:critical script + plan-templates verifier fix
    • d599578 — 4 workflow verifier prompt fixes
    • 612e18b — Phase 4 default-on flip (code + decision doc)
    • 4186284 — mark default-off doc SUPERSEDED + index update

    Real team runs (Tier 7 outcomes):

    • team_20260722083504_cae04a2804a24d79 — full-implementation, 3/4 phases done, 04_verify hung (root cause investigation)
    • team_20260722095143_2e58fce2ce91af19 — first fast-fix smoke, 3/3 PASS (after test:critical introduced)
    • team_20260722100811_9bf95bebff2b052a — final fast-fix smoke, 3/3 PASS, verifier used cached output (449s wall-clock)

    Frequently asked questions

    What to verify before installation and use

    What does the real-test-pi-crew source document cover?

    End-to-end verification discipline for pi-crew changes. Distilled from the broker Phase-4 rollout (commits 1cb2dca → d599578 → 612e18b → 4186284, July 2026). The pain this skill prevents: shipping code that compiles + unit-tests-green but breaks in the user's live Pi session, or…

    How do I install real-test-pi-crew?

    The source record exposes this install command: npx skills add https://github.com/baphuongna/pi-crew --skill "skills/real-test-pi-crew". Inspect the command and pinned source before running it.

    Which Agent platforms does the source record declare?

    The pinned source record declares support for: cursor.

    Which permission-related actions were detected?

    Static rules flagged exec-script, read-files, write-files in the source; the page lists the matching lines and excerpts.

    Alternatives

    Compare before choosing

    Computed 9420

    upex-galaxy/agentic-qa-boilerplate

    test-automation

    Plan, write, and review automated tests following KATA (Komponent Action Test Architecture) on Playwright + TypeScript, or explain existing automated tests in a sealed read-only mode. Use when writing E2E or API/integration tests, creating Page or Api components, designing ATCs, parameterizing test data, registering fixtures, reviewing test code for KATA compliance, or requesting break-down-tests / a plain-English test breakdown. The explain mode reads source and reports assertions without enter

    Computed 9420

    upex-galaxy/agentic-qa-boilerplate

    test-documentation

    Analyze, prioritize, and document test cases in TMS (Jira/Xray), or repair an existing Story-ATS-ATP-ATR-TC cascade through a sealed explicit mode. Use for Test/ATP/ATR artifacts, ROI and automation verdicts, maintaining traceability, fix-traceability, or broken TMS links. The repair-traceability mode audits, plans, waits for explicit approval, applies, and verifies without launching the general documentation workflow. Do NOT use for writing test code (test-automation) or running suites (regress

    Computed 9324,921

    alirezarezvani/claude-skills

    chaos-engineering

    Use when planning, running, or learning from chaos engineering experiments. Triggers on "chaos experiment", "fault injection", "gameday", "resilience test", "blast radius", "steady state", "abort criteria", "Chaos Toolkit", "Chaos Mesh", "Litmus", "Gremlin", "AWS FIS", or any deliberate failure-injection question. Ships experiment designer, blast-radius calculator, and postmortem generator (all stdlib Python), 4 references on chaos principles + experiment design + attack taxonomy + tooling lands

    Computed 93203

    PramodDutta/qaskills

    Angry User Simulator

    Simulate aggressive user behavior patterns including rapid clicking, random navigation, form abuse, tab spamming, and unexpected interaction sequences to find UI resilience issues