Source profileQuality 96/100

magnus919/agent-skills/ai-operating-economics/SKILL.md

ai-operating-economics

Use when deciding whether an AI-enabled workflow should be adopted, scaled, constrained, redesigned, or retired, and the decision must connect business outcomes, worker or user effects, quality guardrails, full operating cost, telemetry, uncertainty, and accountable governance. Do not use for a standalone financial model, infrastructure cost calculation, agent evaluation design, runtime operations, or general AI governance; route those details to the neighboring specialist skills.

Source repository stars
61
Declared platforms
0
Static risk flags
1
Last source update
2026-08-26
Source checked
2026-08-28

Decision brief

What it does: where it fits

Use when deciding whether an AI-enabled workflow should be adopted, scaled, constrained, redesigned, or retired, and the decision must connect business outcomes, worker or user effects, quality guardrails, full operating cost, telemetry, uncertainty, and accountable governance. Do not use for a standalone financial model, infrastructure cost calculation, ag…

Best for

  • Build an evidence-backed business case for an AI use case or agentic workflow.
  • Decide whether an AI pilot should scale, remain bounded, be redesigned, or stop.
  • Review claimed AI productivity, savings, adoption, or transformation results.

Not for

  • Treating an AI benchmark, speed increase, or demo as evidence of business value.
  • Treating a vendor survey as an audited financial result or causal estimate.

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/magnus919/agent-skills --skill "ai-operating-economics"
Safe inspection promptEditorial

Inspect the Agent Skill "ai-operating-economics" from https://github.com/magnus919/agent-skills/blob/531ff6753784823c878c92b988c6e55266ce09a9/ai-operating-economics/SKILL.md at commit 531ff6753784823c878c92b988c6e55266ce09a9. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Core Workflow

    Use this sequence for an AI initiative review. Load the detailed method and the evidence-record template when the task requires a durable artifact.

    Primary outcome: the result the initiative exists to improve.Leading indicators: early evidence that the mechanism is operating.Countermetrics: quality, safety, customer, worker, privacy, reliability, or equity measures that could worsen.
  2. 02

    Quick Start by Need

    Review the “Quick Start by Need” section in the pinned source before continuing.

    Review and apply the “Quick Start by Need” source section.
  3. 03

    Choose Review Depth

    Review the “Choose Review Depth” section in the pinned source before continuing.

    Review and apply the “Choose Review Depth” source section.
  4. 04

    Verification Checklist

    Before delivering an AI operating economics decision, verify:

    [ ] The workflow, intervention, population, baseline, decision owner, and human-control boundary are explicit.[ ] The value hypothesis is falsifiable and tied to an observable outcome.[ ] At least one countermetric is defined for each benefit claim.
  5. 05

    Entry Points

    Review the “Entry Points” section in the pinned source before continuing.

    Review and apply the “Entry Points” source section.

Permission review

Static risk signals and limitations

Reads files

low · line 206

The documentation asks the agent to read local files, directories, or repositories.

| Need | Load when | File |

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score96/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars61SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
magnus919/agent-skills
Skill path
ai-operating-economics/SKILL.md
Commit
531ff6753784823c878c92b988c6e55266ce09a9
License
MIT
Collected
2026-08-28
Default branch
main
View the original SKILL.md

AI Operating Economics

Overview

AI initiatives are operating interventions, not merely model purchases or ROI spreadsheets. Their value depends on what work changes, who benefits, what quality or risk changes with it, what the complete intervention costs, and whether the organization can observe and govern those changes.

This skill provides the cross-domain decision spine for evaluating an AI-enabled workflow. It does not replace financial modeling, product measurement, statistical inference, agent evaluation, runtime operations, or AI governance. It makes those inputs meet in one accountable decision record.

The core question is not “Did the model make people faster?” It is: “What changed in this workflow, for whom, at what full cost, with what outcome and countermetric evidence, and what authority should the organization grant next?”

Entry Points

Starting stateStart withPrimary artifact or route
Idea or proposed AI workflowSteps 1–2templates/ai-initiative-evidence-record.md
Existing pilot or outcome dataSteps 3–7references/evidence-method.md plus the evidence record
Request for broader population or side-effect authoritySteps 7–8; load references/evidence-method.md section 7a for the governance packetGovernance evidence packet plus the evidence record
Executive, portfolio, launch, or lifecycle reviewSteps 8–9templates/ai-economics-review.md; route launch/runtime details onward
Standalone financial, statistical, telemetry, runtime, or governance implementation taskWhen Not to UseNamed adjacent specialist skill

When to Use

Load this skill when the user needs to:

  • Build an evidence-backed business case for an AI use case or agentic workflow.
  • Decide whether an AI pilot should scale, remain bounded, be redesigned, or stop.
  • Review claimed AI productivity, savings, adoption, or transformation results.
  • Design an AI value-realization or post-launch outcome review.
  • Connect model and tool spend to workflow outcomes and worker or customer effects.
  • Compare AI options while accounting for measurement uncertainty and non-comparable evidence.
  • Prepare an executive, product, portfolio, or lifecycle decision about an AI-enabled intervention.

When Not to Use

If the task is primarily...Route toThis skill still contributes...
Financial statements, pricing, CAC/LTV, runway, or SaaS metricsfinancial-modelingThe AI workflow's outcome and cost evidence can feed the model
Token, infrastructure, quota, capacity, or SLO-cost modelingcapacity-and-cost-engineeringThe economic decision can consume the resulting cost boundary
Metric trees, event schemas, instrumentation QA, or product dashboardsproduct-analytics-and-measurementThe decision defines which outcome and countermetric evidence matters
Experimental design, causal inference, statistical testing, or power analysisdata-scientistThe decision specifies the claim and comparison it must support
Agent datasets, graders, traces, regression analysis, or telemetry implementationagent-evals-and-observabilityThe decision consumes verified evaluation and telemetry evidence
Production rollout, runtime budgets, authority, fallback, escalation, or disablementagent-production-operationsThe decision sets the evidence and authority boundary
Organization-wide AI risk, policy, compliance, or governance operating modelsai-governanceThe initiative record supplies an operating case and unresolved gaps
Launch-readiness packet or production go/no-go decisionproduction-readinessThe initiative disposition becomes one readiness input
General product governance cadence without an AI-specific value questionproduct-operations-and-governanceUse this skill only for the AI-specific value and operating-economics question

Non-Negotiable Reasoning Rules

  1. Workflow evidence beats model evidence. A benchmark, demo, or vendor claim does not establish value in the target workflow.
  2. Speed is not value. Time saved can be spent on lower-value work, offset by review and exception handling, or enable higher-value work. Measure the business or user outcome directly.
  3. Averages are not enough. Inspect worker, user, task, geography, tenure, risk, and quality slices. An aggregate gain can hide a subgroup loss.
  4. Every benefit metric needs a countermetric. Pair throughput or cost with quality, safety, customer, worker, privacy, or reliability measures appropriate to the workflow.
  5. Token cost is not total cost. Include model calls, tools, retrieval, storage, networking, observability, engineering, human review, change management, governance, and unused committed capacity when material.
  6. Evidence classes must stay separate. Label observed results, causal estimates, inferences, vendor-reported findings, stakeholder assertions, and normative requirements distinctly.
  7. Missing evidence is a decision input. Do not turn an unknown into a favorable assumption. Record the gap, owner, consequence, and next evidence needed.
  8. Authority follows evidence. A positive pilot does not justify unrestricted autonomy. Scale capability and authority in bounded slices with explicit reversal conditions.
  9. Do not manufacture precision. Use ranges, scenarios, sensitivity, and confidence where inputs are uncertain. Do not rank non-comparable studies or vendors.
  10. The decision is reversible only if the artifact says how. Record the stop trigger, rollback or containment path, decision owner, and review date.

Core Workflow

Use this sequence for an AI initiative review. Load the detailed method and the evidence-record template when the task requires a durable artifact.

Quick Start by Need

NeedFirst actionLoad next
Triage a claimName the workflow, decision, and evidence classSteps 1–3; evidence classes are defined in Step 7
Build a durable recordCopy the initiative evidence record and complete the header firsttemplates/ai-initiative-evidence-record.md
Investigate uncertain evidenceFreeze the claim table before drafting conclusionsreferences/evidence-method.md
Prepare a reviewAssemble evidence, slices, cost, gaps, and dispositiontemplates/ai-economics-review.md

Choose Review Depth

ModeUse whenMinimum evidenceOutput
TriageA claim or opportunity needs a bounded first decisionWorkflow, value hypothesis, one outcome, one countermetric, known gapsHold, with a routing/evidence plan
StandardA pilot or workflow decision can change population or investmentComparison, outcome/countermetrics, slices, cost boundary, owner, reversal pathScale, constrain, redesign, or hold
High-assuranceAuthority, sensitive data, material user impact, or irreversible change is involvedStandard evidence plus governance packet, human oversight, incident/revalidation, and decommissioning evidenceScale only within an explicit authority boundary, or Hold

1. Define the intervention and decision

Name the workflow, population, task boundary, intervention mode, baseline, decision sought, and decision owner. State whether the AI assists, recommends, routes, executes, or replaces/removes work. Define what remains human-controlled.

Do not begin with the model name or a claimed percentage. Begin with the work that changes and the decision the evidence must support.

2. State the value hypothesis

Write a falsifiable hypothesis:

For [population] doing [workflow], [intervention] will change [outcome] by [direction/range] without exceeding [countermetric boundary], at [full operating cost boundary], compared with [baseline], over [period].

If the proposed outcome is only “productivity,” decompose it into the actual customer, employee, operational, financial, or mission outcome. If the outcome cannot be observed or credibly proxied, mark the initiative measurement-incomplete rather than inventing a proxy.

3. Build the outcome and countermetric map

Define:

  • Primary outcome: the result the initiative exists to improve.
  • Leading indicators: early evidence that the mechanism is operating.
  • Countermetrics: quality, safety, customer, worker, privacy, reliability, or equity measures that could worsen.
  • Adoption and substitution measures: who uses the system, what work changes, and what work is displaced or added.
  • Guardrail thresholds: contextual limits with an owner and response.

Route metric definitions and instrumentation plans to product analytics. Route statistical or causal design to data science. This skill owns the connection between the evidence and the decision, not the detailed statistical method.

4. Establish the full economic boundary

Record both:

  • Marginal economics: what changes when one more task, user, or workflow unit is served.
  • Fully loaded economics: the costs required to make the intervention available and govern it.

At minimum consider inference, tool use, retrieval, storage, data transfer, observability, engineering, evaluation, human review, training, support, change management, governance, security, and committed capacity. Separate fixed, variable, step-function, and avoided costs. Define the denominator precisely: task, resolved case, completed workflow, active user, customer outcome, or another meaningful unit.

Route the detailed model to capacity-and-cost-engineering or financial-modeling. Never divide total spend by an undifferentiated request count when requests have materially different resource or outcome profiles.

5. Design the evidence comparison

Choose the strongest feasible comparison before interpreting results:

  • Randomized or staggered rollout when feasible.
  • Matched or difference-in-differences comparison when appropriate.
  • Within-workflow baseline with explicit pre-period and seasonality limits.
  • Controlled pilot with a documented task and population boundary.
  • Descriptive before/after evidence only when stronger designs are infeasible, labeled accordingly.

Record selection effects, learning effects, concurrent initiatives, task-mix changes, worker self-selection, quality measurement gaps, and changes in pay or incentives. If the comparison cannot support the requested claim, narrow the claim rather than upgrading the method rhetorically.

6. Segment before aggregating

Report the overall result and inspect slices that could change the decision:

  • Worker experience, skill, role, and training status.
  • Task complexity, risk, volume, and exception rate.
  • Customer or user segment.
  • Geography, language, accessibility, and relevant demographic groups when lawful and appropriate.
  • Human-review burden and escalation path.
  • Quality, safety, and error severity.

Treat heterogeneous effects as a finding, not noise to average away. A tool that helps novices while harming expert quality may need differentiated assistance modes, not universal rollout.

7. Classify the evidence

For every material claim, label it:

ClassMeaningPermitted use
ObservedDirectly measured in the target workflow with a stated methodDescribe what happened within the stated scope
Causal estimateSupported by a credible comparison or experimentAttribute an effect only within the design's limits
InferredReasoned from observed evidence and explicit assumptionsGuide a bounded hypothesis or scenario
Vendor-reportedProvider survey, case study, or product documentationEstablish reported adoption or available capability, not realized ROI
AssertedStakeholder or proposal claim not yet verifiedTrack as an assumption and evidence gap
NormativeStandard or framework recommendationDefine a control expectation, not an outcome claim

Keep the source, access date, scope, version, caveat, and permitted interpretation with each claim. Load references/source-index.md for the research basis and evidence boundaries.

Minimum Claim Ledger

For each material claim, record: claim, evidence class, source and scope, what it supports, what it does not support, open challenge, and permitted language. Keep unknown claims visible; do not let a source URL or vendor report stand in for direct workflow evidence.

Minimum Decision Record

Every completed review must expose, in one durable artifact: the intervention and population, value hypothesis, primary outcome, countermetrics, comparison and limitations, cost boundary, relevant slices, evidence classes, missing evidence with owner, disposition, authority limit, reversal path, and review trigger.

Disposition Quick Pick

Evidence stateDefault dispositionNext control
Outcome and countermetrics support a bounded expansion; cost and slices are understoodScaleName the next population and authority slice
Value is plausible but a cost, quality, subgroup, or authority boundary remains unresolvedConstrainLimit population, task, quota, or human review
The mechanism creates avoidable failure or burdenRedesignChange the workflow or control and rerun the comparison
Required evidence is missing or conflictingHoldAssign the evidence owner and review trigger
Value is absent or countermetrics exceed boundsRetireProtect affected people, migrate, and record learning
A material gap is accepted temporarily by a named humanExceptionSet expiry, containment, approver, and revisit condition

8. Produce a bounded decision

Choose exactly one primary disposition:

  • Scale: evidence supports expansion within a named scope and authority boundary.
  • Constrain: value is plausible, but cost, quality, risk, or distributional effects require limits.
  • Redesign: the mechanism or workflow needs modification before another test.
  • Hold: evidence is insufficient for the requested decision; specify the missing evidence.
  • Retire: observed value is absent or countermetrics exceed acceptable bounds, with a transition path.
  • Exception: proceed despite a named gap only with an accountable human approver, expiry or revisit trigger, and containment plan.

Closure Conditions

  • Scale: next population, authority slice, owner, and review trigger are recorded.
  • Constrain: the boundary, quota, human-review rule, and condition for expansion are recorded.
  • Redesign: the changed mechanism, rerun comparison, and new acceptance boundary are recorded.
  • Hold: the missing evidence, owner, method, and due trigger are recorded.
  • Retire: transition, affected-person protection, decommissioning, and retained learning are recorded.
  • Exception: named human approver, scope, expiry, containment, and revisit condition are recorded.

A decision is incomplete without an owner, review date or trigger, evidence gaps, and reversal path. Route launch or runtime consequences to the appropriate specialist skill.

9. Close the learning loop

At the review date, compare expected versus observed outcomes, cost, quality, worker or user effects, adoption, and incidents. Preserve the updated evidence record and state whether the prior hypothesis was supported, weakened, refuted, or still unresolved. Feed verified incidents and near misses into evaluation and governance work rather than treating them as anecdotal follow-up.

Load-on-Demand References

NeedLoad whenFile
Apply the full research and decision method, including comparison design and uncertaintyEvidence is incomplete, contested, or consequentialreferences/evidence-method.md
Review sources and permitted interpretationsA claim needs provenance or a source boundaryreferences/source-index.md
Fill a durable initiative recordStarting a new workflow review or pilot assessmenttemplates/ai-initiative-evidence-record.md
Prepare an executive or lifecycle reviewCombining one or more initiative records for a decisiontemplates/ai-economics-review.md

Common Pitfalls

  • Treating an AI benchmark, speed increase, or demo as evidence of business value.
  • Treating a vendor survey as an audited financial result or causal estimate.
  • Reporting one average while omitting worker, task, quality, or customer slices.
  • Calling token spend “AI cost” while omitting review, tooling, retrieval, infrastructure, or change costs.
  • Choosing a denominator that makes the economics look favorable, such as all requests instead of completed or resolved workflows.
  • Treating a missing baseline as zero or assuming adoption means benefit.
  • Using a normative framework as proof that an intervention is safe or effective.
  • Granting broader authority because a pilot had a positive mean result.
  • Reusing a prior decision after the workflow, model, population, cost boundary, or evidence source changed.
  • Writing a sophisticated recommendation without preserving the source-level evidence that supports it.

Verification Checklist

Before delivering an AI operating economics decision, verify:

  • The workflow, intervention, population, baseline, decision owner, and human-control boundary are explicit.
  • The value hypothesis is falsifiable and tied to an observable outcome.
  • At least one countermetric is defined for each benefit claim.
  • Fixed, variable, step-function, and fully loaded costs are separated where material.
  • The denominator represents meaningful work or value, not merely requests or tokens.
  • The comparison design and its limitations are stated.
  • Relevant worker, user, task, quality, and risk slices are inspected or explicitly unavailable.
  • Claims are labeled by evidence class and traced to sources.
  • Missing evidence is visible with an owner and next step.
  • The disposition, authority boundary, reversal path, and review trigger are recorded.
  • Detailed statistical, financial, instrumentation, governance, runtime, and launch checks were routed to their owning skills.

Exit Criteria

Stop when the requested decision is supported by a durable evidence record, or when a bounded hold/escalation is the honest result. Do not continue refining prose to conceal missing evidence.

Frequently asked questions

What to verify before installation and use

What does the ai-operating-economics source document cover?

Use when deciding whether an AI-enabled workflow should be adopted, scaled, constrained, redesigned, or retired, and the decision must connect business outcomes, worker or user effects, quality guardrails, full operating cost, telemetry, uncertainty, and accountable governance. Do not use for a standalone financial model, infrastructure cost calculation, ag…

How do I install ai-operating-economics?

The source record exposes this install command: npx skills add https://github.com/magnus919/agent-skills --skill "ai-operating-economics". Inspect the command and pinned source before running it.

Which permission-related actions were detected?

Static rules flagged read-files in the source; the page lists the matching lines and excerpts.

Alternatives

Compare before choosing

Computed 10014,706

prowler-cloud/prowler

postgresql-indexing

PostgreSQL indexing best practices for Prowler: index design, partial indexes, partitioned table indexing, EXPLAIN ANALYZE validation, concurrent operations, monitoring, and maintenance. Trigger: When creating or modifying PostgreSQL indexes, analyzing query performance with EXPLAIN, debugging slow queries, reviewing index usage statistics, reindexing, dropping indexes, or working with partitioned table indexes. Also trigger when discussing index strategies, partial indexes, or index maintenance

Computed 100147

oaustegard/claude-skills

featuring

Generate hierarchical _FEATURES.md files that describe what a codebase DOES from a user/consumer perspective, anchored to source symbols via tree-sitting. Supports large complex codebases through feature-driven decomposition into sub-feature files. Uses a multi-pass synthesis: orientation → detail → overview rewrite. Use when someone says "what does this do", "document features", "feature inventory", "_FEATURES.md", or needs to understand a codebase's purpose before modifying it. Complements tre

Computed 9931,947

HKUDS/Vibe-Trading

strategy-generate

Create, modify, and optimize quantitative trading strategies, then backtest and evaluate them.

Computed 9982

vasilyu1983/AI-Agents-public

qa-testing-ios

Guides iOS testing with XCTest, XCUITest, Swift Testing, simctl, and xcresult. Use when choosing destinations, controlling flakes, or parsing test artifacts for native apps.