Source profileQuality 85/100

majiayu000/spellbook/skills/app-user-story-qa/SKILL.md

app-user-story-qa

End-to-end app feature inventory and user-story testing workflow with a canonical tracker. Use when the user asks to audit every feature, derive expected behavior from code, test user journeys, or explicitly fix and retest documented UX or logistical defects.

Source repository stars
249
Declared platforms
0
Static risk flags
1
Last source update
2026-08-02
Source checked
2026-08-04

Decision brief

What it does—and where it fits

End-to-end app feature inventory and user-story testing workflow with a canonical tracker.

Best for

  • Use this skill to turn a broad "test every feature in this app" request into a controlled loop with one source of truth. Anchor every feature to code, track expected behavior in a canonical spreadsheet, run tests agains…

Not for

  • Tasks that require unconfirmed production actions or broad system permissions.
  • Environments where the pinned source and install steps cannot be inspected.

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/majiayu000/spellbook --skill "skills/app-user-story-qa"
Safe inspection promptEditorial

Inspect the Agent Skill "app-user-story-qa" from https://github.com/majiayu000/spellbook/blob/01c5d88b0139a80ac38bfe7206ea99f28b0fc999/skills/app-user-story-qa/SKILL.md at commit 01c5d88b0139a80ac38bfe7206ea99f28b0fc999. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Workflow

    Map all user-visible surfaces first:

    CLI commands and flagsWeb or desktop UI routesAPI endpoints
  2. 02

    Purpose

    Use this skill to turn a broad "test every feature in this app" request into a controlled loop with one source of truth. Anchor every feature to code, track expected behavior in a canonical spreadsheet, run tests against each user story, and classify failures. Apply fixes only w…

    Use this skill to turn a broad "test every feature in this app" request into a controlled loop with one source of truth. Anchor every feature to code, track expected behavior in a canonical spreadsheet, run tests agains…
  3. 03

    Route

    Use planfirst for most runs. Switch to clarifyfirst only when the app boundary, writable checkout, production risk, or acceptance criteria are unclear enough that a wrong assumption would waste substantial work.

    Run a repo state snapshot: current path, branch, latest remote, dirty files, open PRs when relevant.If the checkout is dirty or user work is present, first decide which state the user asked to test. Preserve current user work in the isolated test worktree when the request targets the working tree; use a clean base onl…Read applicable repo instructions such as AGENTS.md, README, architecture docs, and feature entrypoints.
  4. 04

    Operating Contract

    Select one mode before editing production code:

    reportonly is the default for audit, inventory, test, diagnose, or tracker requests. It may create or update the requested canonical tracker and run tests, but it must not change product behavior.applyfixes requires the current user request to explicitly ask for fixes. It covers only defects already reproduced and classified within the agreed app boundary. Earlier approval and a generic request to "test everythi…Inventory local code and docs, create or update the single canonical tracker, run local tests, and document failures.
  5. 05

    Gotchas

    Do not create scattered notes or duplicate trackers. One canonical tracker is the audit log.

    Do not create scattered notes or duplicate trackers. One canonical tracker is the audit log.Do not invent expected behavior from product hopes. Expected behavior comes from current code, docs, and visible user surfaces.Do not mark a feature passed from stale output. Fresh evidence from the current session is required.

Permission review

Static risk signals and limitations

Reads files

low · line 16

The documentation asks the agent to read local files, directories, or repositories.

Read applicable repo instructions such as `AGENTS.md`, `README`, architecture docs, and feature entrypoints.

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score85/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars249SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
majiayu000/spellbook
Skill path
skills/app-user-story-qa/SKILL.md
Commit
01c5d88b0139a80ac38bfe7206ea99f28b0fc999
License
MIT
Collected
2026-08-04
Default branch
main
View the original SKILL.md

App User Story QA

Purpose

Use this skill to turn a broad "test every feature in this app" request into a controlled loop with one source of truth. Anchor every feature to code, track expected behavior in a canonical spreadsheet, run tests against each user story, and classify failures. Apply fixes only when the current request explicitly authorizes them.

Route

Use plan_first for most runs. Switch to clarify_first only when the app boundary, writable checkout, production risk, or acceptance criteria are unclear enough that a wrong assumption would waste substantial work.

Before editing:

  1. Run a repo state snapshot: current path, branch, latest remote, dirty files, open PRs when relevant.
  2. If the checkout is dirty or user work is present, first decide which state the user asked to test. Preserve current user work in the isolated test worktree when the request targets the working tree; use a clean base only when the user asked for the base branch.
  3. Read applicable repo instructions such as AGENTS.md, README, architecture docs, and feature entrypoints.
  4. Search for existing QA trackers before creating a new one.
  5. Define the app boundary by user-facing surfaces, not internal helper functions.

Operating Contract

Select one mode before editing production code:

  • report_only is the default for audit, inventory, test, diagnose, or tracker requests. It may create or update the requested canonical tracker and run tests, but it must not change product behavior.
  • apply_fixes requires the current user request to explicitly ask for fixes. It covers only defects already reproduced and classified within the agreed app boundary. Earlier approval and a generic request to "test everything" do not authorize fixes.

If the mode is ambiguous, use report_only and record proposed fixes in the tracker.

Direct actions:

  • Inventory local code and docs, create or update the single canonical tracker, run local tests, and document failures.
  • In apply_fixes only, fix narrow in-scope logistical or UX defects and add focused coverage when user-observable behavior changes.

Escalate before:

  • Testing production systems, using paid external services, changing credentials or deployment state, deleting user data, or broadening the app boundary beyond the user's request.
  • Editing shared test infrastructure, weakening assertions, or changing product requirements instead of the implementation.

Evidence-backed pushback:

  • Challenge requests to skip the tracker, mark untested rows as passed, or fix symptoms without reproducing the failure. Cite the tracker row, command output, code anchor, or missing environment precondition.

Feedback loop:

  • Promote repeated setup failures, brittle manual paths, or recurring test gaps into tracker notes, scripts, docs, or follow-up issues instead of leaving them only in chat.

Gotchas

  • Do not create scattered notes or duplicate trackers. One canonical tracker is the audit log.
  • Do not invent expected behavior from product hopes. Expected behavior comes from current code, docs, and visible user surfaces.
  • Do not mark a feature passed from stale output. Fresh evidence from the current session is required.
  • Do not "fix" environment blockers unless the user asked for environment repair.
  • Do not infer fix authorization from words such as "audit", "test", "check", or "create a tracker". Keep those runs in report_only.

Canonical Tracker

Create or update exactly one tracker. Prefer a real spreadsheet when the runtime supports it; otherwise use a CSV and treat it as the canonical spreadsheet. Do not scatter status across side reports.

Use these columns:

Feature ID,Surface,Feature / capability,User story,Expected behavior based on code,Code anchors,Initial test approach,Initial test command or route,Status,Test result,Errors,Fix status,Retest result,Notes

Rules:

  • Assign stable IDs such as F001, F002.
  • Require at least one code anchor per row. A row without an anchor is not complete.
  • Write expected behavior from current code and docs, not from aspirational specs.
  • Keep the tracker parseable: validate CSV/spreadsheet row widths after edits.
  • Update the tracker during the loop, not only at the end.

Status Values

Use a small, consistent status set:

  • Not Tested
  • Passed
  • Failed - Product
  • Failed - UX
  • Failed - Logistical
  • Failed - Test Infra
  • Blocked - Env
  • Fixed
  • Passed after fix

Use Failed - Logistical for repo-owned setup, script, port, packaging, or local workflow defects that block a valid user path. Use Blocked - Env for local machine issues such as missing credentials, occupied services outside the repo, stale PATH binaries, or unavailable optional runtimes. Do not "fix" the user's environment unless they explicitly ask.

Defect Taxonomy

Classify every failure before fixing:

  • Product: implemented behavior violates the user story or loses data.
  • UX: behavior works but the user path is confusing, brittle, or poorly surfaced.
  • Logistical: scripts, ports, setup, packaging, or local workflow make valid behavior hard to exercise.
  • Test Infra: the test itself is flaky, racy, or asserts the wrong contract.
  • Env: external setup blocks execution and is not a repo defect.

Fix only defects that are in scope for the task. Record out-of-scope defects in the tracker with clear rationale.

Workflow

1. Inventory

Map all user-visible surfaces first:

  • CLI commands and flags
  • Web or desktop UI routes
  • API endpoints
  • background jobs or hooks that affect users
  • plugin, package, install, or release surfaces
  • import/export, backup, migration, and governance flows

For each feature, record the user story as:

As a <user>, I want <capability>, so <outcome>.

Keep stories practical. Do not create rows for private helpers unless the user can observe the behavior through a surface.

2. Test

For every tracker row, choose the strongest feasible evidence:

  • direct smoke test for the user path
  • focused unit or integration test for the exact contract
  • broad regression suite when a feature is already covered there
  • manual or browser test when automation is missing

Record the command, route, or manual steps in the tracker. Fresh output from the current session is required before marking a row passed.

3. Fix

In report_only, record the reproduced defect, evidence, and proposed fix, then continue testing without editing production code.

In apply_fixes, state the exact defect and files being changed before editing. Keep fixes narrow:

  • For UX defects, fix production code and add or update focused coverage.
  • For product defects, record the evidence and escalate unless the user's request explicitly authorizes product behavior fixes.
  • For logistical defects, improve scripts, defaults, setup checks, or error messages.
  • For test-infra defects, preserve the behavior contract and fix the fixture or race.
  • Do not weaken assertions to make the suite pass.

Stop and re-evaluate after three failed attempts on the same defect.

4. Retest

After each fix:

  1. Rerun the focused failing test or smoke.
  2. Rerun the relevant broader gate for the touched surface.
  3. Update Fix status and Retest result.
  4. Keep the original error text or summary in Errors so the tracker remains an audit log.

Before completion, run the repo's required formatting, build, typecheck, lint, and test commands when practical.

Completion Gate

Finish only when:

  • every feature row has a user story, expected behavior, code anchors, and a status;
  • every Failed row is fixed and retested in apply_fixes, or is explicitly recorded as proposed, out of scope, or blocked with evidence in report_only;
  • every applied fix has a retest result;
  • the tracker validates structurally;
  • the final answer names the tracker path, changed files, defects found, fixes made, and verification commands.

If the repo has existing dirty work that is not yours, mention the isolated worktree or scope boundary in the final answer.

Alternatives

Compare before choosing

Computed 97106

AI-Unified-Process/marketplace

browserless-test

Creates Vaadin Browserless server-side unit tests for Vaadin views covering navigation, component interactions, form validation, grid operations, and notifications. Use when the user asks to "write Browserless tests", "write Vaadin UI unit tests", "unit test a Vaadin view without a browser", "create view tests with the official Vaadin testing framework", or mentions Browserless testing, SpringBrowserlessTest, browserless-test-junit6, UI Unit Testing, or server-side Vaadin testing.

Computed 976

mgiovani/cc-arsenal

team-review

Multi-agent review team: architecture, security, performance, testing, style, docs/UX, plus an adversary that cross-examines the other 6, for security-sensitive, architectural, or large PRs (15+ files) where a single-agent pass risks missing cross-cutting issues. Use for auth/payments/PII changes, schema/pattern changes, compliance sign-off, or when asked to 'get the review team on this' / 'multi-agent review' / 'thorough review before merge'. For a standard PR or a quick pre-merge check, use /r

Computed 957

event4u-app/agent-config

playwright-testing

Use when writing Playwright E2E tests — browser automation, visual regression testing, Page Objects, fixtures, and reliable test patterns.

Computed 94165

JasonColapietro/suede-creator-skills

suede-ai-eval

Design AI evals that catch regressions before users do: rubrics, test cases, failure modes, acceptance gates, and AI-SPEC artifacts.