Best for
- Use when tracking delivery, quality trajectories, cost, experience, pilots, scorecards, or leadership reporting.
vasilyu1983/AI-Agents-public/frameworks/shared-skills/skills/dev-ai-coding-metrics/SKILL.md
Measures AI coding impact and extension robustness. Use when tracking delivery, quality trajectories, cost, experience, pilots, scorecards, or leadership reporting.
Decision brief
Measures coding assistants and coding agents without collapsing results into vanity metrics or one blended score.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Declared | Source record | Install path and trigger |
| Claude Code | Declared | Source record | Install path and trigger |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/vasilyu1983/AI-Agents-public --skill "frameworks/shared-skills/skills/dev-ai-coding-metrics"Inspect the Agent Skill "dev-ai-coding-metrics" from https://github.com/vasilyu1983/AI-Agents-public/blob/53f6cb73ea53a2646e3e7d4665062ad66f3683ac/frameworks/shared-skills/skills/dev-ai-coding-metrics/SKILL.md at commit 53f6cb73ea53a2646e3e7d4665062ad66f3683ac. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
1. Define the decision. 2. Pick the program mode: assistant, agent, or mixed. 3. Build the minimum viable scorecard. 4. Choose the study design. 5. Produce one deliverable.
Review the “When to Use This Skill” section in the pinned source before continuing.
Review the “Defaults” section in the pinned source before continuing.
Review the “ASCII Flow” section in the pinned source before continuing.
Review the “Quick Reference” section in the pinned source before continuing.
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 91/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 82 | Source | Repository attention, not individual Skill quality |
| Compatibility | 2 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Measures coding assistants and coding agents without collapsing results into vanity metrics or one blended score.
The critical distinction is mode: assistants help inline or in chat; agents execute multi-step work and need task-level measurement. Do not measure them as if they were the same thing.
| Trigger | Example |
|---|---|
| Designing a pilot or rollout scorecard | "We're rolling out Copilot to 200 engineers — what do we measure?" |
| Diagnosing usage-up / outcomes-flat | "Seat utilization is 80% but PR throughput is unchanged" |
| Comparing assistant vs. agent workflows | "Should we instrument these separately?" |
| Building an ROI model or leadership report | "Finance wants a renewal decision by Q3" |
| Designing an experiment better than vendor benchmarks | "We can't trust the vendor's numbers — how do we run our own study?" |
| Rule | Rationale |
|---|---|
| Start from the decision, not the telemetry available | Prevents instrument-what-is-easy bias |
| Separate assistant and agent funnels | Mixing hides which workflow drives results |
| Pair every speed metric with quality + experience | Speed alone is misleading |
| Aggregate at team level | Individual dashboards become surveillance |
| Treat benchmarks as capability signals, not business KPIs | Benchmark gaps do not equal production gaps |
AI coding metrics request
-> decision to support: buy, renew, improve, prove, or diagnose
-> split mode: assistant, agent, or mixed
-> select scorecard families: adoption, delivery, quality, economics, experience
-> choose study design and baseline window
-> collect team-level and task-level evidence
-> report confidence, sample size, and confounds
-> deliver ROI model, dashboard, experiment plan, or executive report
| Decision | Default Output |
|---|---|
| buy, renew, or cut a tool | ROI model plus executive report |
| improve adoption | adoption metrics plus survey |
| prove delivery impact | productivity metrics plus experiment plan |
| check quality drift | quality metrics plus dashboard |
| understand trust or friction | developer-experience metrics plus survey |
| evaluate coding agents | agent-execution metrics plus experiment plan |
| Mode | Unit of Analysis | Primary Emphasis |
|---|---|---|
| assistant | developer-day, team-week, repo-month | adoption, delivery, quality, experience |
| agent | task, PR, workflow run | task success, merge, revert, review burden, cost per accepted change |
| mixed | team-week plus task-level samples | separate the two funnels before combining results |
Use the smallest scorecard that can answer the decision:
| Family | What It Tells You |
|---|---|
| adoption | whether usage is real and sustained |
| delivery | whether software flow is faster where AI actually touches the path |
| quality | whether speed gains are offset by defects, rework, review burden, or declining extension robustness |
| economics | whether the value justifies tool and operating cost |
| experience | whether developers trust the tool and want to keep using it |
| agent execution | whether autonomous workflows succeed in production, not just in demos |
Minimum baseline: 8 weeks of pre-intervention data. Two-week baselines produce noisy causal inference — week-to-week variance in PR throughput, review lag, and defect escape routinely exceeds the signal size of AI tooling effects.
| Situation | Design |
|---|---|
| new pilot, no control group | before/after with ≥8 weeks baseline |
| enough comparable teams | matched A/B or stratified assignment |
| teams resist permanent denial of tools | crossover design |
| agent workflow change on one task family | task-level shadow comparison or reviewer-blind evaluation |
| leadership wants a fast answer | balanced scorecard with explicit caveats, not a causal claim |
Use before publishing any AI coding report:
| Claim | Evidence | Caveat |
|---|---|---|
| AI amplifies existing strengths and weaknesses | DORA 2025 AI report; conditional-impact model confirmed | Not a universal accelerant |
| Experienced developers ~19% slower with early-2025 tools (RCT) | METR July 2025 RCT, realistic open-source tasks | Specific to early-2025 tooling generation |
| METR believes developers more sped-up in 2026 than 2025 | METR Feb 2026 update | 30-50% of participants declined no-AI tasks (selection bias); unreliable signal |
| Self-reported: median 1.4-2x value of work from AI (2026) | METR May 2026 survey, n=349 | Self-report; METR found 40pp gap between perceived and actual gains in 2025 study |
| Throughput +66%, PR review time +441%, incidents per PR +243% | Faros AI 2026 telemetry, 22k devs / 4k teams | Organizational telemetry, not RCT; PRs merged without review up +31% |
| DORA 2025: 90% of developers use AI daily | DORA 2025 AI report | Adoption does not equal delivery impact |
| Modeled first-year AI ROI ~39% (500-person org); adoption raises change-failure rate (5%->6%), an "instability tax" | DORA 2026 ROI of AI-Assisted Software Development report (Apr 2026) | Vendor-modeled scenario, not a cross-org RCT; treat the 39% figure as an illustrative scenario, not a universal benchmark |
| AI yields 35-40% gains on simple tasks but ~10% on complex legacy code | DORA 2026 ROI report | Reinforces task-complexity segmentation already required by this skill's study design defaults |
| DX Core 4 unifies DORA + SPACE + DevEx into 4 dimensions (Speed, Effectiveness, Quality, Business Impact) | DX Core 4, formalized publicly Apr 2026 | Vendor framework; specific benchmarks need independent replication |
| One-shot pass rates can miss degradation across repeated agent edits | SlopCodeBench v1, Mar 2026 preprint | Python experiments only; trajectory signals are not correctness proofs or universal targets |
Reject a scorecard or report if any of the following apply:
References
Assets and data
Scripts
Before applying this skill on a non-trivial task, read learnings.consolidated.md in this directory (and learnings.md if present).
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.
Frequently asked questions
Measures coding assistants and coding agents without collapsing results into vanity metrics or one blended score.
The source record exposes this install command: npx skills add https://github.com/vasilyu1983/AI-Agents-public --skill "frameworks/shared-skills/skills/dev-ai-coding-metrics". Inspect the command and pinned source before running it.
The pinned source record declares support for: codex, claude code.
Alternatives
vasilyu1983/AI-Agents-public
Guides iOS testing with XCTest, XCUITest, Swift Testing, simctl, and xcresult. Use when choosing destinations, controlling flakes, or parsing test artifacts for native apps.
vasilyu1983/AI-Agents-public
Designs and audits UI/UX systems with usability and accessibility requirements. Use when shaping flows, design systems, interaction patterns, or WCAG-aware product behavior.
vasilyu1983/AI-Agents-public
Designs session lifecycle for coding-agent runtimes. Use when implementing resume, transcript restoration, checkpoint rewind, cross-worktree recovery, or session-state persistence.
vasilyu1983/AI-Agents-public
Designs and audits native Android interfaces. Use when reviewing Compose layout, typography, color, motion, or adaptive patterns on a verified emulator build.