eugenelim/agent-ready-repo/packs/experience-design/.apm/skills/design-review/SKILL.md
design-review
Evaluate an existing screen, flow, or mockup with a severity-rated findings list: quality-floor pass (states, a11y, motion), heuristic eval (Nielsen's 10), marketing clarity pass (tweet test, five-second scan, painkiller-first — fires on above-fold copy with a persuasion goal), and taste critique (grounded aesthetic reference + platform fit). Triggers on 'critique this design', 'review this screen', 'what is wrong with this mockup', 'do a heuristic eval', 'is this usable', 'does this fit our aes
- Source repository stars
- 15
- Declared platforms
- 0
- Static risk flags
- 0
- Last source update
- 2026-08-05
- Source checked
- 2026-08-05
Decision brief
What it does—and where it fits
Runs a structured evaluation of a screen, flow, or mockup and returns a prioritized, severity-rated findings list — each issue mapped to the recognized usability principle or aesthetic reference it violates, with one concrete, portable recommendation. The list is the artifact: i…
Not for
- Claiming to be a fresh-context reviewer. This skill is authoring-time self-review. The genuine fresh-context UX review is the forked-context experience-reviewer agent — it runs independently, between sessions, and does…
- Reprinting the aesthetic reference values. The taste critique points to the grounded referent; it never reprints palette entries, type scales, spacing values, or any literal from the reference. See references/taste-crit…
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/eugenelim/agent-ready-repo --skill "packs/experience-design/.apm/skills/design-review"Inspect the Agent Skill "design-review" from https://github.com/eugenelim/agent-ready-repo/blob/9563bc93aa5b0750b327be2fd95676ff2a5ec63b/packs/experience-design/.apm/skills/design-review/SKILL.md at commit 9563bc93aa5b0750b327be2fd95676ff2a5ec63b. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
Procedure
0. Surface inventory (multi-surface platforms only). If the review subject is a multi-surface platform (e.g., marketing site + documentation site, or app + marketing + docs): (a) enumerate every surface and label its genre (marketing / documentation / analytical / etc.); (b) con…
Surface inventory (multi-surface platforms only). If the review subject is a multi-surface platform (e.g., marketing site + documentation site, or app + marketing + docs): (a) enumerate every surface and label its genre…Frame the surface and load design-principles. Name the user, the primary task, and each step under review. This anchors every severity call that follows. Also load the design-principles artefact at docs/design/principle…Apply the shared floor first. Run the quality-floor checklist at references/quality-floor.md against the surface — handle all states, the accessibility floor, the reduced-motion principle. Each miss is a finding mapped… - 02
Output rendering
Severity list — Lead each finding with a severity glyph — 🟥 blocker, 🟧 major, 🟨 minor, ⚪ advisory — worst first, one finding per line, file:line anchor aligned.
Severity list — Lead each finding with a severity glyph — 🟥 blocker, 🟧 major, 🟨 minor, ⚪ advisory — worst first, one finding per line, file:line anchor aligned. - 03
When to invoke
Confirm all three before drafting; if any fails, resolve it first.
There is something concrete to review — a screen, flow, mockup, or described surface. A vibe with no artifact isn't ready; route to creative-direction.You know whose task you're judging — a critique needs a user and a goal. Without them, severity is unanchored guesswork; draw out the primary task first.You're evaluating, not creating — the ask is "is this good," not "make this." If it's deriving values or structuring a layout, hand to design-system or information-architecture. - 04
Genre-specific rubrics
After the quality-floor and heuristic passes, route to the genre-specific rubric that matches the surface's surface-genre: declaration (from the per-screen brief). If no genre is declared, elicit it — genre rubrics are not optional for genre-bearing surfaces; they surface issues…
Navigation tier match — does the navigation strategy match the page count tier? (≤30 pages: flat nav; 30–200: hub-and-spoke; 200: search-first.) A flat-nav structure on a 400-page docs site fails this item.Content typing — is every piece of content typed (tutorial / how-to / reference / explanation)? Does the page structure match its declared type? (A tutorial that contains a full API reference mid-step is typed incorrect…Landing page orientation and hub structure — does the docs landing page serve orientation rather than marketing copy? Verify all three hub jobs are present: (1) "Start Here" entry point — one link, one promise, above th… - 05
Documentation genre rubric
1. Navigation tier match — does the navigation strategy match the page count tier? (≤30 pages: flat nav; 30–200: hub-and-spoke; 200: search-first.) A flat-nav structure on a 400-page docs site fails this item. 2. Content typing — is every piece of content typed (tutorial / how-t…
Navigation tier match — does the navigation strategy match the page count tier? (≤30 pages: flat nav; 30–200: hub-and-spoke; 200: search-first.) A flat-nav structure on a 400-page docs site fails this item.Content typing — is every piece of content typed (tutorial / how-to / reference / explanation)? Does the page structure match its declared type? (A tutorial that contains a full API reference mid-step is typed incorrect…Landing page orientation and hub structure — does the docs landing page serve orientation rather than marketing copy? Verify all three hub jobs are present: (1) "Start Here" entry point — one link, one promise, above th…
Permission review
Static risk signals and limitations
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 86/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 15 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- eugenelim/agent-ready-repo
- Skill path
- packs/experience-design/.apm/skills/design-review/SKILL.md
- Commit
- 9563bc93aa5b0750b327be2fd95676ff2a5ec63b
- License
- Apache-2.0
- Collected
- 2026-08-05
- Default branch
- main
View the original SKILL.md
Skill: design-review
Runs a structured evaluation of a screen, flow, or mockup and returns a prioritized, severity-rated findings list — each issue mapped to the recognized usability principle or aesthetic reference it violates, with one concrete, portable recommendation. The list is the artifact: it turns "this feels off" into something a stakeholder can argue and a builder can act on.
Four modes, always run in this order:
- Quality-floor pass — mandatory; checks all states, accessibility, and reduced-motion.
- Heuristic evaluation — walks the surface against recognized usability principles.
- Marketing clarity pass (when the artifact includes above-fold copy with a persuasion/conversion goal — a landing page, marketing page, or product announcement) — checks the copy against the tweet test, five-second scan, and painkiller-first structure. Does not fire for internal tools, forms, settings screens, or content pages with no conversion goal.
- Taste critique (when a grounded aesthetic reference is present) — checks the screen against the grounded aesthetic reference and platform fit.
Authoring-time self-review. This skill is an interactive, authoring-time tool — it runs in the session, with the author. It is not a fresh-context pass and not an adversarial reviewer; a same-session critique marks its own homework. The genuine fresh-context UX review is the forked-context
experience-revieweragent — invoke it for an independent pass after the authoring session.
Output rendering
Severity list — Lead each finding with a severity glyph — 🟥 blocker, 🟧 major, 🟨 minor, ⚪ advisory — worst first, one finding per line, file:line anchor aligned.
When to invoke
Confirm all three before drafting; if any fails, resolve it first.
- There is something concrete to review — a screen, flow, mockup, or described surface. A vibe with no artifact isn't ready; route to
creative-direction. - You know whose task you're judging — a critique needs a user and a goal. Without them, severity is unanchored guesswork; draw out the primary task first.
- You're evaluating, not creating — the ask is "is this good," not "make this." If it's deriving values or structuring a layout, hand to
design-systemorinformation-architecture.
Procedure
-
Surface inventory (multi-surface platforms only). If the review subject is a multi-surface platform (e.g., marketing site + documentation site, or app + marketing + docs): (a) enumerate every surface and label its genre (marketing / documentation / analytical / etc.); (b) confirm which surface is under review in this pass; (c) note which other surfaces exist — they will need separate passes; (d) flag: cross-surface integration check required (see the marketing genre rubric for copy voice continuity and the information-architecture skill for cross-surface wayfinding). If the review subject is a single surface, skip to step 1.
-
Frame the surface and load design-principles. Name the user, the primary task, and each step under review. This anchors every severity call that follows. Also load the
design-principlesartefact atdocs/design/principles/<slug>.mdif one exists for this surface — every finding in this review must be mapped to the principle it was judged against. When a finding cannot be traced to any principle, route it to one of three places: (a) a quality-floor commitment it breaches (these are always valid regardless of whether principles exist — the floor applies unconditionally), (b) a recognized heuristic from the evaluation in step 3, or (c) a new-principle decision — flag it to the team as a gap in the design-principles artefact, not as a finding in this review. Pure aesthetic preferences with no principle backing and no floor/heuristic grounding go in a Director's notes section at the end of the findings list, clearly separated from the severity-rated findings. This is a mandatory procedure step — a design-review that skips design-principles integration does not produce a traceable findings list. -
Apply the shared floor first. Run the
quality-floorchecklist atreferences/quality-floor.mdagainst the surface — handle all states, the accessibility floor, the reduced-motion principle. Each miss is a finding mapped to the floor commitment it breaches; accessibility misses start at major. Then apply the surface-specific mobile checklist for the surface's genre:Marketing surface mobile: primary CTA — and secondary CTA if present — is visible above the fold on a small-phone viewport without scrolling; hero top padding does not consume more than one-fifth of a common phone-height viewport before any content appears; navigation drawer items are full-width touch targets — compact inline chips in a vertical list are not acceptable; drawer has an explicit open/close state signal on the toggle icon; stat/feature strip dividers reset correctly when items reflow to multiple rows; tab bars scroll horizontally or wrap without a broken 3+1 layout; grid minimum accounts for the narrowest target viewport usable width minus horizontal padding on both sides; install/code blocks scroll horizontally rather than clipping.
Documentation surface mobile: code blocks scroll horizontally (not clip); full-bleed sections that extend to the viewport edge must not produce a horizontal scrollbar on systems where scrollbars occupy layout space; sidebar collapses correctly and navigation is accessible without it at narrow-tablet widths and below; single-column reading width stays within a comfortable reading range; search is accessible from the mobile header without opening the sidebar.
-
Run the heuristic evaluation. Walk the surface against the recognized usability principles in
references/heuristics.md. For each problem, record what you observed before you judge it. -
Map and rate. Map each finding to the single best-fit principle (or floor commitment) and assign a 0–4 severity, naming the frequency × impact × persistence factors that set it. See
references/heuristics.md. -
Run the marketing clarity pass (when the artifact includes above-fold copy with a persuasion/conversion goal). For each of the three criteria, record what you observed, then map and rate: a. Tweet test — can the headline or tagline stand alone as a conviction statement? If you shared just that line with no surrounding context, would it communicate what this is and why it matters to the target reader? Failure: the line only describes the product, names a category without a reader benefit, or requires the page for meaning. b. Five-second scan — after 5 seconds on the above-fold, can a first-time visitor answer: what is this / who is it for / should I care? All three must be answerable from the visible content alone, not inferred. Failure: one or more answers are absent, ambiguous, or below the fold. c. Painkiller-first structure — does the copy lead with the reader's problem, pain, or desired outcome before naming the product's features? A painkiller solves a known hurt; a vitamin is a nice-to-have. Failure: copy leads with the author's feature list or product identity rather than the reader's recognized need.
Map each finding to the criterion it violates and assign a 0–4 severity using the frequency × impact × persistence rubric, where impact means conversion/persuasion cost — how badly the miss hurts the reader's ability to determine fit and take the intended action. (This is a deliberate application of the same rubric to a persuasion-cost dimension; it is not a separate scale.) Label source mode
marketing. A settings screen or internal tool that is out of scope for this pass produces no findings with this label. -
Run the taste critique (when a grounded aesthetic reference from
creative-directionis available). Seereferences/taste-critique.mdfor the full method. In brief: a. Check aesthetic alignment — for each named goal in the grounded reference, ask whether the screen advances, is neutral to, or contradicts it. Ground each verdict in the recorded referent (persona + precedent + standards), never in a fresh opinion. b. Check platform fit — verify the screen respects the platform surface's (responsive-web / iOS / Android / cross-platform) conventions; point to the platform standard as the warrant, never reprint its values. c. Map and rate taste findings — each taste finding maps to the aesthetic goal it contradicts or the platform convention it violates; rate 0–4 by the same severity rubric (frequency × impact × persistence), with 0 reserved for genuine disagreement where the referent does not clearly resolve the call. -
Prioritize and recommend. Merge all findings from all modes. Sort worst-first across modes, lead with a count-by-severity headline, and give each finding one concrete, portable recommendation expressed as design intent — never a stack-specific implementation. Label the source mode (
floor/heuristic/marketing/taste) so the reader knows which lens each finding came from.
Genre-specific rubrics
After the quality-floor and heuristic passes, route to the genre-specific rubric that matches the surface's surface-genre: declaration (from the per-screen brief). If no genre is declared, elicit it — genre rubrics are not optional for genre-bearing surfaces; they surface issues the generic passes miss.
Each rubric is a numbered checklist. Work through it in order; a "no" is a finding, mapped to the genre rubric item that failed it. Rate each finding with the standard 0–4 severity (frequency × impact × persistence). Label source mode genre-rubric.
Multi-surface routing: When the subject includes both a marketing surface and a documentation surface, the marketing genre rubric and the documentation genre rubric must each be run as separate passes. Do not collapse them into a single pass. Each pass gets its own findings list; the cross-surface integration check (copy voice continuity, cross-surface wayfinding) runs after both surface passes are complete.
Documentation genre rubric
- Navigation tier match — does the navigation strategy match the page count tier? (≤30 pages: flat nav; 30–200: hub-and-spoke; >200: search-first.) A flat-nav structure on a 400-page docs site fails this item.
- Content typing — is every piece of content typed (tutorial / how-to / reference / explanation)? Does the page structure match its declared type? (A tutorial that contains a full API reference mid-step is typed incorrectly.)
- Landing page orientation and hub structure — does the docs landing page serve orientation rather than marketing copy? Verify all three hub jobs are present: (1) "Start Here" entry point — one link, one promise, above the fold; (2) content-type entry points — one section per Diátaxis type (tutorial / how-to / reference / explanation), named by what the reader accomplishes, not by content-type label; (3) search above the fold with a placeholder naming a real example query. A landing page that leads with product benefits rather than reader navigation fails this item. For >200-page sites (see item 1, navigation tier match), search must be persistent and prominent — a top-right corner widget does not meet the search-first requirement.
- TTFV reachability — is the first-value moment achievable from the tutorial entry point? (Tutorial is scoped to ≤20 minutes of active work; prerequisites are stated before the reader starts; code samples work as pasted.)
- Machine-readability by design — are machine-readability requirements built into the IA? (Code blocks typed with language identifiers; API tables with consistent column structure; heading hierarchy that reflects content type.) These should be design decisions, not implementation afterthoughts.
Marketing genre rubric
- Hero approach fit — does the hero approach match the product's position and reader's awareness level? (Vision for underfunded markets; social-proof for mature markets; job-to-be-done for buyers who know the pain but not the product.) An approach mismatch is a conversion-strategy finding, not a cosmetic one.
- Above-fold spec — verify all six elements are present and correctly placed: (1) Headline: ≤10 words, IC-first (reader pain/goal before product name); (2) Subheadline: conviction-building (outcome or benefit), not a second problem statement; (3) Primary CTA: outcome language, not system action ("Install the core loop" not "Submit"); (4) Secondary CTA: only if primary asks meaningful commitment — absent is valid; (5) Proof signal: specific number, recognizable logo, or third-party rating, positioned adjacent to the CTAs; (6) Friction microcopy: one line removing the dominant objection to clicking the primary CTA ("No credit card", "Reversible", "1 command to try, 1 to remove") — absence is a blocker if the primary CTA implies commitment. Tone collision check: if the headline is a Statement, the subheadline must stay conviction-building — not pivot to problem-agitation (that belongs in a separate section below the fold).
- Scroll-story zone integrity — does each zone in the scroll story have a single job? A zone that simultaneously introduces a feature, shows a testimonial, and prompts a second CTA has no job — it has three, and does none of them well.
- Cross-surface copy voice continuity (only when a documentation surface is in scope for the same platform) — does the marketing copy voice carry through to the docs surface? Flag as minor: marketing uses precision/technical register but docs uses casual/tutorial register without explanation. Flag as major: marketing makes product claims the docs surface contradicts or doesn't support. This check requires reading at minimum the docs landing page and one how-to page alongside the marketing surface.
- Social proof tier calibration — is the social proof at the right tier for the product's maturity stage? (Early: named customer quotes; Growth: logos + metrics; Scaled: independent validation.) A startup listing analyst rankings it hasn't earned fails this item.
Analytical genre rubric
- Business-question traceability — does every widget trace to at least one of the 3–5 named business questions? A widget that answers no named question is a candidate for removal; flag it.
- Tier 1 KPI ceiling — are Tier 1 KPIs ≤9 and above the fold without scrolling? Does each KPI have a visible comparison baseline (not just a primary value)? "1,247 active users" with no comparison baseline fails this item.
- Filter adjacency — do filter controls live adjacent to the data they filter? A date range filter in a sidebar while the charts it controls are in the center creates a spatial mismatch; fails Shneiderman's zoom-and-filter principle.
- Widget state completeness — does every widget have a designed loading state, empty state, error state, and stale-data state? A widget designed only in its "populated" state fails this item.
Informational genre rubric
- Line-length constraint — is the reading column width constrained to produce 45–75 characters per line at the chosen type scale? A full-width text block on a wide viewport fails this item.
- Reading-pattern consistency — does the chosen pattern (F for dense/reference-heavy, Z for conversational/single-topic) apply consistently to heading placement, first-sentence construction, and layout decisions across the page?
- "What's next" intent clarity — does the post-article zone serve a defined reader intent (related topic / deeper dive / action / discovery) at a clear visual priority? A zone with four equal-weight "what's next" categories at identical visual weight serves no intent.
Marketplace genre rubric
- Card hierarchy decision weight — is the card hierarchy ordered by decision weight (primary identifier → key attribute → social proof signal → secondary attributes → CTA)? An attribute that drives the match decision (price, availability, compatibility) buried below secondary details fails this item.
- Filter architecture + buyer behavior match — is the filter architecture appropriate to the declared buyer behavior? (Browse-first: chip-based filters, immediately visible; Search-first: sidebar filters, complex taxonomy.) A browse-first surface with a sidebar-only filter fails this item.
- Transaction bridge context — does the transaction flow (cart, checkout, booking) keep marketplace context visible throughout? A checkout screen that shows only the cart line items, with no reference to the listing the buyer selected, fails this item.
Workspace genre rubric
- Last-location landing — does the surface land returning users at their last working context? A workspace that greets returning users with a generic dashboard instead of their last location fails this item.
- Interrupt escalation respect — is the interrupt escalation ladder respected? (Ambient for non-urgent; focal — modal, sound, motion overlay — only for time-sensitive + action-required.) A workspace that uses focal interrupts for non-urgent notifications fails this item.
- Agentic output review surface — for any agentic output (generated content, proposed code change, automated action), is there a review surface between the output and its application? An agent that applies output without a review step fails this item.
Anti-patterns to refuse
- Claiming to be a fresh-context reviewer. This skill is authoring-time self-review. The genuine fresh-context UX review is the forked-context
experience-revieweragent — it runs independently, between sessions, and does not mark its own homework. - Reprinting the aesthetic reference values. The taste critique points to the grounded referent; it never reprints palette entries, type scales, spacing values, or any literal from the reference. See
references/taste-critique.md. - Unrated opinions. A finding without a severity and a violated principle or aesthetic goal is taste, not a critique. Map it or drop it.
- Skipping the floor. The
quality-floorpass is mandatory, not optional polish. A surface can clear all ten heuristics and still fail the floor. - Prescribing the stack. Recommendations name the what and why as design intent. The moment you reach for a framework, value, or property, you've left the method.
- Burying the catastrophe. A flat or alphabetized list hides the blocker. Worst-first, always, with the headline up top.
Alternatives
Compare before choosing
mgiovani/cc-arsenal
team-review
Multi-agent review team: architecture, security, performance, testing, style, docs/UX, plus an adversary that cross-examines the other 6, for security-sensitive, architectural, or large PRs (15+ files) where a single-agent pass risks missing cross-cutting issues. Use for auth/payments/PII changes, schema/pattern changes, compliance sign-off, or when asked to 'get the review team on this' / 'multi-agent review' / 'thorough review before merge'. For a standard PR or a quick pre-merge check, use /r
wondelai/skills
improve-code-quality
Guided journey from a working-but-untested vibe-coded prototype to a production-ready product with tests, clean structure, a business-rules boundary, and resilience at scale. Orchestrates nine skills phase by phase - working-with-legacy-code, clean-code, refactoring-patterns, software-design-philosophy, clean-architecture, pragmatic-programmer, release-it, system-design, ddia-systems - asking the user questions at every decision point and recording results in the project docs/ folder (TESTING.md
wondelai/skills
remove-technical-debt
Guided journey from a large aged codebase everyone fears to touch to one that is safe to change, legible, bounded, and resilient - paid down in place without a rewrite. Orchestrates eight skills phase by phase - working-with-legacy-code, refactoring-patterns, clean-code, software-design-philosophy, clean-architecture, pragmatic-programmer, release-it, domain-driven-design - asking the user questions at every decision point and recording results in the project docs/ folder (TESTING.md, TECH-DEBT.
github/awesome-copilot
screen-recording
Create annotated animated GIF demos and screen recordings for pull requests and documentation. Covers frame capture, timing, imageio-based GIF creation, and per-frame annotation workflows.