Best for
- Use when evaluating code, designs, architectures, or comparing alternative approaches.
yonatangross/orchestkit/src/skills/assess/SKILL.md
Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.
Decision brief
Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Declared | Source record | Install path and trigger |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/yonatangross/orchestkit --skill "src/skills/assess"Inspect the Agent Skill "assess" from https://github.com/yonatangross/orchestkit/blob/4e5c1327b7d7902022ee69328e12db1f6a88f390/src/skills/assess/SKILL.md at commit 4e5c1327b7d7902022ee69328e12db1f6a88f390. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
xhigh silently falls back to high on a model that does not implement it: no error, no log line. /ork:doctor Category 14 reports this, and only when it can positively prove the active model lacks the tier.
Load: Read("${CLAUDEPLUGINROOT}/skills/chain-patterns/references/mcp-detection.md")
Review the “Phase Handoffs” section in the pinned source before continuing.
BEFORE creating tasks, clarify assessment dimensions:
Load details: Read("${CLAUDEPLUGINROOT}/skills/assess/references/orchestration-mode.md") for env var check logic, Agent Teams vs Task Tool comparison, and mode selection rules.
Permission review
The documentation asks the agent to read local files, directories, or repositories.
Read(file_path="$ARGUMENTS[0]") # If file pathThe documentation asks the agent to run terminal commands or scripts.
node "${CLAUDE_PLUGIN_ROOT}/skills/assess/scripts/render-spec.mjs" .claude/chain/assess-dashboard.json --checkThe documentation asks the agent to run terminal commands or scripts.
node "${CLAUDE_PLUGIN_ROOT}/skills/assess/scripts/render-spec.mjs" .claude/chain/assess-dashboard.jsonThe documentation asks the agent to create, modify, or delete local files.
After the composite and grade are final (post-refutation, Phase 2.5), ALWAYS write the machine-readable verdict — this is the stop-gate `/ork:implement` reads before Phase 1. Mirror the Phase 7b spec-emit pattern: build, write compact JSON,Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 92/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 223 | Source | Repository attention, not individual Skill quality |
| Compatibility | 1 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations.
/ork:assess backend/app/services/auth.py
/ork:assess our caching strategy
/ork:assess --model=opus the current database schema
/ork:assess frontend/src/components/Dashboard
xhigh)| Effort | Behavior |
|---|---|
low / medium | Subset of dimensions, faster turnaround |
high (default) | All six dimensions with pros/cons |
xhigh | All six dimensions + one additional assessor pass focused on uncertainty/caveats; emits confidence per dimension |
xhighsilently falls back tohighon a model that does not implement it: no error, no log line./ork:doctorCategory 14 reports this, and only when it can positively prove the active model lacks the tier.
TARGET = "$ARGUMENTS" # Full argument string, e.g., "backend/app/services/auth.py"
# $ARGUMENTS[0] is the first token (CC 2.1.59 indexed access)
# Model override detection (CC 2.1.72)
MODEL_OVERRIDE = None
for token in "$ARGUMENTS".split():
if token.startswith("--model="):
MODEL_OVERRIDE = token.split("=", 1)[1] # "opus", "sonnet", "haiku", "fable"
TARGET = TARGET.replace(token, "").strip()
Pass MODEL_OVERRIDE to all Agent() calls via model=MODEL_OVERRIDE when set. Accepts symbolic names (opus, sonnet, haiku, fable on harnesses whose Agent tool lists it; note fable is premium API spend after 2026-07-12) or full IDs (claude-opus-4-8) per CC 2.1.74.
Switching to Opus via
/model(CC 2.1.144+):/modelnow changes the model for the current session only, so picking Opus for an assess run no longer persists past it. Pressdin the picker only to set a default for new sessions.
$CLAUDE_EFFORT is the primary signal. CC 2.1.120 sets this env var from /effort or the model picker. --effort= token in $ARGUMENTS is the explicit override fallback (also covers older CC).
# Read env first (CC 2.1.120+), then check explicit override
EFFORT = os.environ.get("CLAUDE_EFFORT") # "low" | "medium" | "high" | "xhigh" | None
for token in "$ARGUMENTS".split():
if token.startswith("--effort="):
EFFORT = token.split("=", 1)[1] # explicit override wins
TARGET = TARGET.replace(token, "").strip()
EFFORT = EFFORT or "high" # default when CC < 2.1.120 and no flag
Use EFFORT to gate dimension count, agent count, and the optional xhigh uncertainty pass — see "Effort levels" table above. On CC < 2.1.120 the env var is unset; the explicit --effort= override is the only path. /ork:doctor Category 14 reports a provably unsupported xhigh request.
Load:
Read("${CLAUDE_PLUGIN_ROOT}/skills/chain-patterns/references/mcp-detection.md")
# 1. Probe MCP servers (once at skill start)
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")
# 2. Store capabilities
Write(".claude/chain/capabilities.json", {
"memory": probe_memory.found,
"skill": "assess",
"timestamp": now()
})
# 3. Check for resume
state = Read(".claude/chain/state.json") # may not exist
if state.skill == "assess" and state.status == "in_progress":
last_handoff = Read(f".claude/chain/{state.last_handoff}")
| Phase | Handoff File | Contents |
|---|---|---|
| 0 | 00-intent.json | Dimensions, target, mode |
| 1 | 01-baseline.json | Initial codebase scan results |
| 2 | 02-evaluation.json | Per-dimension scores + evidence |
| 3 | 03-report.json | Final report, grade, recommendations |
BEFORE creating tasks, clarify assessment dimensions:
AskUserQuestion(
questions=[{
"question": "What dimensions to assess?",
"header": "Dimensions",
"options": [
{"label": "Full assessment (Recommended)", "description": "All dimensions: quality, maintainability, security, performance"},
{"label": "Code quality only", "description": "Readability, complexity, best practices"},
{"label": "Security focus", "description": "Vulnerabilities, attack surface, compliance"},
{"label": "Quick score", "description": "Just give me a 0-10 score with brief notes"}
],
"multiSelect": false
}]
)
Based on answer, adjust workflow:
Load details: Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/orchestration-mode.md") for env var check logic, Agent Teams vs Task Tool comparison, and mode selection rules.
# 1. Create main task IMMEDIATELY
TaskCreate(
subject="Assess: {target}",
description="Comprehensive evaluation with quality scores and recommendations",
activeForm="Assessing {target}"
)
# 2. Create subtasks for each assessment phase
TaskCreate(subject="Understand target and gather context", activeForm="Understanding target") # id=2
TaskCreate(subject="Discover scope and build file list", activeForm="Discovering scope") # id=3
TaskCreate(subject="Rate quality across 6 dimensions", activeForm="Rating quality") # id=4
TaskCreate(subject="Analyze pros and cons", activeForm="Analyzing pros/cons") # id=5
TaskCreate(subject="Compare alternatives", activeForm="Comparing alternatives") # id=6
TaskCreate(subject="Generate improvement suggestions", activeForm="Generating suggestions") # id=7
TaskCreate(subject="Compile assessment report", activeForm="Compiling report") # id=8
# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"]) # Scope needs target understanding
TaskUpdate(taskId="4", addBlockedBy=["3"]) # Rating needs scoped file list
TaskUpdate(taskId="5", addBlockedBy=["4"]) # Pros/cons needs quality scores
TaskUpdate(taskId="6", addBlockedBy=["4"]) # Alternatives need quality scores
TaskUpdate(taskId="7", addBlockedBy=["5", "6"]) # Suggestions need analysis
TaskUpdate(taskId="8", addBlockedBy=["7"]) # Report needs suggestions
# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress") # When starting
TaskUpdate(taskId="2", status="completed") # When done — repeat for each subtask
| Question | How It's Answered |
|---|---|
| "Is this good?" | Quality score 0-10 with reasoning |
| "What are the trade-offs?" | Structured pros/cons list |
| "Should we change this?" | Improvement suggestions with effort |
| "What are the alternatives?" | Comparison with scores |
| "Where should we focus?" | Prioritized recommendations |
| Phase | Activities | Output |
|---|---|---|
| 1. Target Understanding | Read code/design, identify scope | Context summary |
| 1.5. Scope Discovery | Build bounded file list | Scoped file list |
| 2. Quality Rating | 6-dimension scoring (0-10) | Scores with reasoning |
| 3. Pros/Cons Analysis | Strengths and weaknesses | Balanced evaluation |
| 4. Alternative Comparison | Score alternatives | Comparison matrix |
| 5. Improvement Suggestions | Actionable recommendations | Prioritized list |
| 6. Effort Estimation | Time and complexity estimates | Effort breakdown |
| 7. Assessment Report | Compile findings | Final report |
Identify what's being assessed and gather context:
# PARALLEL - Gather context
Read(file_path="$ARGUMENTS[0]") # If file path
Grep(pattern="$ARGUMENTS[0]", output_mode="files_with_matches")
mcp__memory__search_nodes(query="$ARGUMENTS[0]") # Past decisions
Load Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/scope-discovery.md") for the full file discovery, limit application (MAX 30 files), and sampling priority logic. Always include the scoped file list in every agent prompt.
Output results incrementally as each evaluation phase completes:
| After Phase | Show User |
|---|---|
| 1. Target Understanding | Scope summary, file list, context |
| 1.5. Scope Discovery | Bounded file list (max 30 files) |
| 2. Quality Rating | Each dimension's score as the evaluating agent returns |
| 3. Pros/Cons | Balanced evaluation summary |
For Phase 2 parallel agents, show each dimension's score as soon as the evaluating agent returns — don't wait for all 4 agents. If any dimension scores below 4/10, flag it immediately as a priority concern requiring user attention.
Rate each dimension 0-10 with weighted composite score. Load Read("${CLAUDE_PLUGIN_ROOT}/skills/quality-gates/references/unified-scoring-framework.md") for dimensions, weights, grade interpretation, and per-dimension criteria. Load Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/quality-model.md") for assess-specific overrides.
Load Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/agent-spawn-definitions.md") for Task Tool mode spawn patterns and Agent Teams alternative.
Composite Score: Weighted average of all 6 dimensions (see quality-model.md).
The assessor that scores a dimension is also its only judge — self-preferential bias.
A separate blind refuter verifies decision-bearing scores before they reach the
composite. Effort gate: low/medium skip this phase entirely; high runs up-to-4
single refuters (advisory, no auto-swing); xhigh runs 3-refuter majority with auto-revise.
Load the protocol + assess bindings: Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/adversarial-refutation.md")
(which loads the shared engine ${CLAUDE_PLUGIN_ROOT}/shared/rules/adversarial-refutation.md).
Producer findings must first pass the evidence-replay gate before entering any score or verdict: Read("${CLAUDE_PLUGIN_ROOT}/shared/rules/evidence-replay.md").
When ORK_ALT_MODEL_CMD is configured and effort is high/xhigh, one quorum slot per high-weight or boundary-adjacent dimension score can route to a non-Claude model (Codex/GPT) for diverse failure modes. Off by default; substitutes one same-model slot, stamps refuter_model for provenance, cannot silently raise the grade (engine §7), owns no credentials/egress (shells out via ORK_ALT_MODEL_CMD, matches the egress guard #2533), and degrades to same-model on an absent command. Shares the review-pr operational doc: Read("${CLAUDE_PLUGIN_ROOT}/skills/review-pr/references/cross-model-refuter.md").
Runs after Phase 2 returns, before the composite/grade and Phases 3-7. Refuters are ALWAYS
isolated Agent(...) Task spawns (never team members, even in Agent Teams mode) fed only the
serialized claim — no producer score, identity, or prose. Revised scores recompute the
composite; the refutation ledger (02b-refutation.json) records survived/killed/downgraded
so wrong scores are auditable. Keep the producer-basis score AND a labeled post-refutation
score — refutation never silently raises the grade.
Load Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/phase-templates.md") for output templates for pros/cons, alternatives, improvements, effort, and the final report.
See also: Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/alternative-analysis.md") | Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/improvement-prioritization.md")
Parse --render= from $ARGUMENTS. Default is both.
| Mode | Behavior |
|---|---|
markdown | Current behavior — markdown assessment report only. No spec emitted. |
json-render | Emit .claude/chain/assess-dashboard.json only. Skip markdown report. |
both | Emit spec and markdown. Default — human reads the report, downstream skills parse the spec. |
When emitting a spec:
Read("${CLAUDE_PLUGIN_ROOT}/skills/assess/references/dashboard-spec.md"). Example: references/dashboard-example.json.Card, StatGrid, DataTable, StatusBadge, BarMeter, Markdown. Top-level fields composite (number) and grade (string) are required for assess specs.BarMeter per dimension scored. The verdict element is a StatusBadge with status success/warning/error mapped from grade (A/B → success, C → warning, D/F → error)..claude/chain/assess-dashboard.json with compact JSON.node "${CLAUDE_PLUGIN_ROOT}/skills/assess/scripts/render-spec.mjs" .claude/chain/assess-dashboard.json --check
If validation fails, fall back to markdown-only and surface the error. Never write a partial spec.
--render=both, render the markdown view from the spec:node "${CLAUDE_PLUGIN_ROOT}/skills/assess/scripts/render-spec.mjs" .claude/chain/assess-dashboard.json
This guarantees JSON spec and markdown report stay in sync.
xhigh effort: when effort=xhigh is active, add a sibling Markdown element per dimension containing confidence and caveats from the uncertainty pass. Reference list it in the dimensions Card's children alongside the BarMeter. See references/dashboard-spec.md for the exact pattern.
Downstream consumption: /ork:implement reads .claude/chain/assess-dashboard.json and pulls the lowest-scoring dimension and high-priority improvements (effort ≤ 2 AND impact ≥ 4) without parsing markdown tables. Measured: assess spec ≈ 830 tokens vs ~3500 token markdown for the same content.
When the assessment lands with a composite score, optionally persist scores + summary to the memory MCP knowledge graph as a typed entity. Future /ork:memory queries can then surface assessment lineage (which decisions did this codebase score 9/10 on testability? when did security regress below 7.0?).
python3 ${CLAUDE_PLUGIN_ROOT}/skills/assess/scripts/memory_writeback.py "<assessment-dir>"
<assessment-dir> is the dir containing assessment.json (typically the session's .claude/chain/). The script writes a memory-writeback.json handoff alongside it.
Auto-skip conditions (all exit 0, all WARN-logged):
| Skip reason | Trigger |
|---|---|
no composite score | assessment.json has no top-level composite numeric field |
yg-mcp-core not importable | yg-mcp-core>=0.3.0 not installed (orchestkit is public; yg-mcp-core lives on private pypi.yonyon.ai — HQ-only) |
memory MCP unreachable | memory MCP server down OR .mcp.json doesn't define memory |
The created entity has:
name: <slug-or-dir>@<timestamp> (stable across re-runs — re-runs create new entities)entityType: assessment (override with --entity-type <type>)observations: composite=X.XX, one <dim>=X.XX per scored dimension, optional summary: ... and topic: ...Mirrors Yonatan-HQ/hq-ext-plugin#194 (audio_podcast handler) and orchestkit#1886 (post-synthesis podcast) pattern. Unblocked by Yonatan-HQ/core#993 (yg-mcp-core 0.3.0).
After the composite and grade are final (post-refutation, Phase 2.5), ALWAYS write the machine-readable verdict — this is the stop-gate /ork:implement reads before Phase 1. Mirror the Phase 7b spec-emit pattern: build, write compact JSON, never write a partial file.
// .claude/chain/assess-verdict.json
{
"rubric": "ork-rubric/1.0",
"skill": "assess",
"verdict": "fail",
"composite": 5.1,
"dimension_scores": {"correctness": 7.0, "maintainability": 6.5, "performance": 5.5, "security": 3.2, "scalability": 6.0, "testability": 4.8, "compliance": 6.2},
"blockers": [
{"dimension": "security", "score": 3.2, "reason": "Unparameterized SQL in auth path (src/api/auth.ts:42)"}
],
"feature": "<assessment topic, e.g. first non-flag token of $ARGUMENTS>"
}
Verdict rules — thresholds come from ${CLAUDE_PLUGIN_ROOT}/skills/assess/rubric.json (schema: ${CLAUDE_PLUGIN_ROOT}/shared/rubric.schema.json):
verdict = "fail" when composite < min_pass (5.5) OR any dimension scores below its min_blocker. Otherwise "pass".min_blocker gets a blockers[] entry — dimension, score, one evidence-backed reason. blockers is [] on pass.Consumers: /ork:implement Step -0.5 blocks Phase 1 on verdict == "fail" (user must fix-first or explicitly override); Phase 7c memory writeback persists the verdict + dimension scores to the memory graph (add a verdict=pass|fail observation) for cross-session learning.
xhigh effort)Current-generation models report their own limits far better than older tiers did. When xhigh effort is active, enrich each dimension's rating with a confidence level and a list of caveats — things the model couldn't verify, assumptions it relied on, or cases it didn't test.
Output schema per dimension (JSON):
{
"dimension": "security",
"score": 7.2,
"confidence": "medium", // "low" | "medium" | "high"
"caveats": [
"Didn't execute the SQL queries against a real DB to confirm parameterization",
"Assumed NODE_ENV=production in deployment; didn't verify CI config",
"Reviewed 12 of 15 handlers; remaining 3 deferred by scope filter"
],
"evidence": ["src/api/auth.ts:42", "src/middleware/guard.ts:88"]
}
Rules:
confidence as an auto-gate. It's a signal for the human reader, not a pass/fail threshold.caveats must be specific. "Didn't check X" with file paths beats "uncertainty about security".score only — not weighted by confidence — to keep the number comparable across runs.Load Read("${CLAUDE_PLUGIN_ROOT}/skills/quality-gates/references/unified-scoring-framework.md") for grade thresholds and scoring criteria.
| Decision | Choice | Rationale |
|---|---|---|
| 6 dimensions | Comprehensive coverage | All quality aspects without overwhelming |
| 0-10 scale | Industry standard | Easy to understand and compare |
| Parallel assessment | 4 agents (6 dimensions) | Fast, thorough evaluation |
| Effort/Impact scoring | 1-5 scale | Simple prioritization math |
| Rule | Impact | What It Covers |
|---|---|---|
complexity-metrics (load ${CLAUDE_PLUGIN_ROOT}/skills/assess/rules/complexity-metrics.md) | HIGH | 7-criterion scoring (1-5), complexity levels, thresholds |
complexity-breakdown (load ${CLAUDE_PLUGIN_ROOT}/skills/assess/rules/complexity-breakdown.md) | HIGH | Task decomposition strategies, risk assessment |
Done means all of these hold:
high/xhigh effort, decision-bearing scores passed the adversarial refutation lane before entering the composite.claude/chain/assess-verdict.json written with verdict pass/fail and a blockers[] entry for every dimension below its min_blockerrender-spec.mjs --check and carries the required composite + grade fieldsork:verify - Post-implementation verificationork:code-review-playbook - Code review patternsork:quality-gates - Task complexity assessment, gate patternsVersion: 1.8.0 (June 2026) — optional cross-model adversarial refuter lane (provenance + cost gate, #2542)
Frequently asked questions
Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations.
The source record exposes this install command: npx skills add https://github.com/yonatangross/orchestkit --skill "src/skills/assess". Inspect the command and pinned source before running it.
The pinned source record declares support for: claude code.
Static rules flagged read-files, exec-script, write-files in the source; the page lists the matching lines and excerpts.
Alternatives
brucesongs/kali-claw
Insecure Design (OWASP A06:2025) focuses on security flaws in system architecture and design phases, rather than code implementation-level bugs.
vasilyu1983/AI-Agents-public
Scans public GitHub repos for agent skills, dev practices, and code patterns. Use when enriching skills, setting team policy, or researching a build domain.
brucesongs/kali-claw
Binary reverse engineering covers the complete chain from static analysis, dynamic debugging, to vulnerability discovery, exploit development, and malware analysis.
Jamie-BitFlight/claude_skills
Create high-quality Claude Code agents from scratch or by adapting existing agents as templates. Use when the user wants to create a new agent, modify agent configurations, build specialized subagents, or design agent architectures. Guides through requirements gathering, template selection, and agent file generation following Anthropic best practices (v2.1.63+).