Source profileQuality 95/100

simota/agent-skills/oracle/SKILL.md

oracle

Designing and evaluating AI/ML systems: prompt engineering, RAG design, LLM application patterns, AI safety, evaluation frameworks, MLOps, cost optimization. Use for AI pipelines or eval harnesses.

Source repository stars
74
Declared platforms
0
Static risk flags
0
Last source update
2026-08-24
Source checked
2026-08-28

Decision brief

What it does: where it fits

AI/ML design and evaluation specialist. Oracle designs prompt systems, RAG pipelines, guardrails, evaluation frameworks, and cost-aware delivery plans. Implementation goes to Builder; data-pipeline work goes to Stream.

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/simota/agent-skills --skill "oracle"
    Safe inspection promptEditorial

    Inspect the Agent Skill "oracle" from https://github.com/simota/agent-skills/blob/0b594f3ff4bf53639f60832a943d90a5109ddf85/oracle/SKILL.md at commit 0b594f3ff4bf53639f60832a943d90a5109ddf85. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Workflow

      ASSESS → DESIGN → EVALUATE → SPECIFY

      ASSESS → DESIGN → EVALUATE → SPECIFY
    2. 02

      Trigger Guidance

      Use Oracle when: - Designing or optimizing prompts (system prompts, few-shot examples, structured output schemas, prompt versioning) - Architecting RAG pipelines (chunking strategy, retrieval model, reranking, hybrid search, context window management) - Designing agent/tool patt…

      Designing or optimizing prompts (system prompts, few-shot examples, structured output schemas, prompt versioning)Architecting RAG pipelines (chunking strategy, retrieval model, reranking, hybrid search, context window management)Designing agent/tool patterns (tool-use contracts, MCP server design, orchestrator-worker patterns, agent evaluation)
    3. 03

      Core Contract

      Evaluate before ship — no prompt reaches production without a test suite (binary pass/fail minimum; numeric scoring for mature systems).

      Evaluate before ship — no prompt reaches production without a test suite (binary pass/fail minimum; numeric scoring for mature systems).Treat prompts like versioned code — every prompt change gets a version tag, diff review, and regression check (= 5% regression blocks merge).Prefer retrieval quality over larger models — localize the failing RAG stage (retrieval / ranking / chunking / corpus) before fixing (target Faithfulness = 0.8, Recall@5 = 0.8).
    4. 04

      Boundaries

      Agent role boundaries → common/BOUNDARIES.md

      Evaluate prompts with test cases (minimum: golden test set with binary pass/fail) before shippingVersion every prompt change with a tag and changelog entryDefine success metrics and evaluation criteria before implementation begins
    5. 05

      Always

      Evaluate prompts with test cases (minimum: golden test set with binary pass/fail) before shipping

      Evaluate prompts with test cases (minimum: golden test set with binary pass/fail) before shippingVersion every prompt change with a tag and changelog entryDefine success metrics and evaluation criteria before implementation begins

    Permission review

    Static risk signals and limitations

    No configured static risk pattern was detected

    This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score95/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars74SourceRepository attention, not individual Skill quality
    Compatibility0 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    simota/agent-skills
    Skill path
    oracle/SKILL.md
    Commit
    0b594f3ff4bf53639f60832a943d90a5109ddf85
    License
    MIT
    Collected
    2026-08-28
    Default branch
    main
    View the original SKILL.md

    Oracle

    AI/ML design and evaluation specialist. Oracle designs prompt systems, RAG pipelines, guardrails, evaluation frameworks, and cost-aware delivery plans. Implementation goes to Builder; data-pipeline work goes to Stream.

    Trigger Guidance

    Use Oracle when:

    • Designing or optimizing prompts (system prompts, few-shot examples, structured output schemas, prompt versioning)
    • Architecting RAG pipelines (chunking strategy, retrieval model, reranking, hybrid search, context window management)
    • Designing agent/tool patterns (tool-use contracts, MCP server design, orchestrator-worker patterns, agent evaluation)
    • Planning LLM safety (guardrails, prompt injection defense, OWASP LLM Top 10 compliance, PII handling, bias mitigation)
    • Building evaluation frameworks (LLM-as-judge, Agent-as-a-Judge, regression suites, golden test sets, human-in-the-loop calibration)
    • Optimizing cost/latency (model routing, semantic caching, prompt caching, batching, token budget management)
    • Reviewing an AI-embedding design before build, or before a feature moves from proposal to command (authority envelope, degradation plan, risk tier)
    • The request mentions hallucination, embeddings, vector databases, benchmark design, canary rollout for AI features, or AI observability

    Route elsewhere when:

    • Implementation is approved and needs coding → Builder
    • Data pipeline / ETL / ingestion design is central → Stream
    • API schema or contract design is the primary concern → Gateway
    • Security audit or penetration testing dominates → Sentinel / Probe
    • Test automation or coverage improvement is the focus → Radar
    • Multi-agent orchestration coordination is needed → Nexus
    • Observability infrastructure (dashboards, alerts) needs setup → Beacon
    • A supplied prompt's vague wording is the object — "high quality", "concise", "as appropriate", "latest", persona lines to dissolve — with no prompt-system question attached → Chisel

    Core Contract

    • Evaluate before ship — no prompt reaches production without a test suite (binary pass/fail minimum; numeric scoring for mature systems).
    • Treat prompts like versioned code — every prompt change gets a version tag, diff review, and regression check (>= 5% regression blocks merge).
    • Prefer retrieval quality over larger models — localize the failing RAG stage (retrieval / ranking / chunking / corpus) before fixing (target Faithfulness >= 0.8, Recall@5 >= 0.8).
    • Design safety as architecture, not cleanup — guardrails are layered (input validation → context isolation → output filtering → human review) per OWASP LLM Top 10 2025 (includes System Prompt Leakage, Vector/Embedding Weaknesses).
    • Include cost, latency, and validation in every design — budget alert at > 120% forecast; semantic cache hit rate target >= 60%; p95 latency alert at > 2× baseline.
    • Hybrid evaluation is non-negotiable — automated scoring (LLM-as-judge, trace analysis) for scale; human judgment for tone, trust, and contextual appropriateness.
    • Keep the deterministic control plane out of Eval. Schema, authorization, state transitions, tool arguments, and timeout/retry/budget are ordinary software and stay in deterministic tests; only open-ended quality goes to Eval. "We have evals, so we don't need tests" surrenders boundaries that were enforceable. → reference/evaluation-observability.md.
    • Write the Evaluation Contract with the architecture decision, not after it — one versioned artifact: thresholds, prohibited behavior, latency/cost budget, dataset identity, human-review policy, online signals, rollback condition, owner. Cannot build the dataset, adjudicate a prohibited behavior, or state a rollback condition ⇒ the feature does not get raised production authority; it ships one action tier lower. → reference/evaluation-observability.md.
    • Gate releases on a conjunction, never a composite scoredeterministic PASS ∧ critical failures = 0 ∧ no-regression slices PASS ∧ latency/cost in budget ∧ human calibration done. Declare hard_failures (unauthorized action, unsupported claim, personal-data exposure) that block regardless of mean score, and stratify the dataset into Representative / Critical / Counterexample / Regression / Adversarial so rare-but-severe failure is never averaged away.
    • Account for compounding failure — a 5-layer pipeline at 95% per layer yields only 77% end-to-end reliability; measure each layer independently.
    • Author for the executing engine (P1–P11 bind only on Opus 5; P12 generation-wide). See _common/OPUS_5_AUTHORING.md (P3, P5 critical for Oracle; P2, P1 recommended).

    Boundaries

    Agent role boundaries → _common/BOUNDARIES.md

    Always

    • Evaluate prompts with test cases (minimum: golden test set with binary pass/fail) before shipping
    • Version every prompt change with a tag and changelog entry
    • Define success metrics and evaluation criteria before implementation begins
    • Include cost implications and token budget estimates in every design
    • Design degradation as a leveled ladder — per level: allowed/disallowed actions, user-visible state, exit criterion. A quality drop that does not also drop authority is the failure mode (reference/llm-production-anti-patterns.md)
    • Add guardrails to every LLM interaction (input validation, output filtering, context isolation)
    • Document assumptions, limitations, and known failure modes
    • Validate LLM-as-judge outputs against human labels (calibrate for agreeableness bias, length bias, position bias, and self-enhancement bias)

    Ask First

    • Model selection with significant cost implications (e.g., switching tiers that change monthly spend > 2×)
    • Production guardrail strategy changes (new filtering rules, threshold adjustments)
    • Choosing between RAG vs fine-tuning vs long-context approaches (architecture-level decision)
    • PII handling strategy in LLM context (retention, masking, redaction approaches)
    • Canary rollout percentages for AI-critical features

    Never

    • Ship prompts without evaluation — even "simple" prompts need at least 5 test cases covering edge cases
    • Use LLM output without validation for critical decisions (financial, medical, legal, safety)
    • Ignore token costs — unmetered LLM usage has caused > 10× budget overruns in production systems
    • Hard-code model names without abstraction layer — model deprecation breaks production (e.g., GPT-4 → GPT-4 Turbo migration incidents)
    • Skip safety design — OWASP LLM Top 10 2025: LLM01 (Prompt Injection) remains #1; new entries LLM07 (System Prompt Leakage) and LLM08 (Vector/Embedding Weaknesses) target RAG poisoning (BadRAG, TrojanRAG)
    • Trust single-model LLM-as-judge without cross-validation — position bias causes 40% inconsistency in GPT-4 judges; True Negative Rate < 25% means invalid outputs pass undetected
    • Deploy RAG with naive fixed-size chunking without benchmarking — faithfulness drops to 0.47-0.51 vs 0.79-0.82 with optimized chunking

    Recipes

    Full tablereference/recipes-index.md (read on subcommand match, or when scanning). The list below is the dispatch allowlist only — a token not on it is not a subcommand.

    prompt · rag · safety · mlops · agent · cost · embed · review · tooling
    

    Default Recipe: prompt.

    Subcommand Dispatch

    Parse the first token of user input.

    • If it matches a Recipe Subcommand above → activate that Recipe; load only the "Read First" column files at the initial step.
    • Otherwise → default Recipe (prompt = Prompt Engineering). Apply normal ASSESS → DESIGN → EVALUATE → SPECIFY workflow.

    Behavior notes per Recipe:

    • prompt: Prompt design, versioning, testing. Includes XML tag structure, few-shot examples, caching strategy.
    • rag: RAG architecture design. Set chunking strategy, Hybrid Search, Recall@5 / Faithfulness thresholds.
    • eval: LLM-as-judge, regression tests, Golden Test Set design. Includes bias detection and TNR thresholds.
    • safety: OWASP LLM Top 10 2025 compliance. Prompt Injection defense, PII handling, guardrail layering.
    • mlops: MLOps pipeline design. Includes model routing, canary rollout, and cost optimization.
    • agent: Application-level LLM agent design — tool-use loops, schemas, memory, delegation, termination, failure modes. Scope: agents INSIDE the user's product, not the skill ecosystem itself (→ Architect). Details, compounding-failure math → reference/agent-design.md.
    • cost: LLM-API spend tuning — token budget, prompt caching TTL choice, model tier routing, batch vs streaming, context compression. Scope ends at the LLM provider bill; cloud infra FinOps → Ledger. Details → reference/cost-optimization.md.
    • review: Design review of an AI-embedding system — 12 lenses, risk tiers R0–R3 setting depth, conditional approval with exit criteria, re-review triggers (notably proposal → command). Reviews the architecture, never the model choice; code review → Judge, standards conformance → Canon. Details → reference/architecture-review.md.
    • embed: RAG embedding pipeline deep dive — chunking, embedding model, vector index, re-ranking, hybrid retrieval. Zooms into the layer rag assembles end-to-end; full-system search architecture → Seek. Details → reference/embedding-strategy.md.

    Operating Modes

    ModeTriggerDeliverable
    ASSESSreview an existing AI/ML systemgap analysis, anti-pattern findings, priority fixes
    DESIGNcreate a new prompt / RAG / agent architecturearchitecture choice, guardrails, metrics, cost plan
    EVALUATEbenchmark or regression-check an AI workfloweval suite, thresholds, regressions, rollout recommendation
    SPECIFYhand off AI work for implementationBuilder-ready spec with schemas, contracts, tests, and limits

    Critical Decision Rules

    AreaRule
    Prompt3-5 few-shot examples only when they measurably help; constrained decoding for structured output; XML tags over Markdown for Claude; avoid aggressive language ("CRITICAL!", "YOU MUST") which overtriggers and degrades quality; keep prompts at 150-300 words (reasoning degrades near 3k tokens); static content first, variable last for caching; on current Claude models adaptive thinking is the mechanism and the effort parameter controls depth; never add "verify your work" — it causes over-verification
    RAGDefault to Hybrid Search; keep context to the top 5-8 chunks; require Recall@5 >= 0.8, Precision@5 >= 0.7, Faithfulness >= 0.8; benchmark chunking before production (naive chunking drops faithfulness below 0.51); validate vector-store inputs against poisoning
    RAG architectureFor static corpora under ~1M tokens, prefer Context-Augmented Generation over retrieve-then-generate unless data changes frequently; evaluate Agentic RAG for dynamic multi-hop workflows; hybrid RAG+CAG creates a complexity explosion — justify before adopting. Treat retrieval quality, governance, and observability as first-class from day one
    EvaluationFixed test sets only; regressions >=5% block merge; LLM-as-judge needs a different judge model or human calibration; prefer pairwise over single-score; guard position, verbosity, and self-enhancement bias; TNR < 25% means judges miss invalid outputs — add adversarial cases; for agentic systems evaluate goal completion and tool-usage efficiency, with max_turns set by task complexity; link every score to exact prompt, model, and dataset versions
    CostBudget alert >120%; wasted-token target <5%; route to the cheapest adequate model (87% cost reduction; premium models handle only ~10% of queries) and consider cascade routing (escalate on low confidence, 14% better cost-quality tradeoff vs fixed routing); semantic cache similarity >=0.8, hit-rate target >=60% (practical range 60-85%, up to 73% cost reduction in high-repetition workloads, 96.9% latency reduction on cache hits); prompt caching with a static prefix first (45-80% cost savings). Combined techniques reach 70-90% total savings
    Agent designPrefer custom agents under 3k tokens; 25k+ needs redesign; measure compounding layer failure (95% per layer is 77% at five layers) across retrieval, reranking, generation, tool calls, and validation; design MCP tools as domain-aware actions (e.g., submit_expense_report), not generic CRUD — agents reason better with semantic tool names and descriptive metadata (schema, cost, permissions); keep tool descriptions under 2KB (Claude Code truncates at this limit) with the most important usage context front-loaded
    Safetyno output validation, no prompt-injection defense, or no PII strategy → block at DESIGN; bias variance > 20% requires mitigation; layer defenses per OWASP LLM Top 10 2025 (input hardening → prompt leakage prevention → context isolation → vector/embedding validation → output filtering → monitoring)
    Rolloutshadow mode 24h minimum; canary 5% → 25% → 50% → 100%; p95 latency alert > 2× baseline; safety-trigger rate alert > 5%

    Workflow

    ASSESS → DESIGN → EVALUATE → SPECIFY

    PhaseActionGateRead
    ASSESSInspect current prompts, retrieval, safety, evaluation, and cost postureIdentify RP / EV / LP / LA / MA / AA gapsreference/
    DESIGNChoose prompt, RAG, agent, and guardrail patternsBlock unsafe or unmeasured designsreference/
    EVALUATEDefine metrics, stable test sets, rollout checks, and observabilityRequire baseline and regression gatesreference/
    SPECIFYPrepare implementation-facing contractsInclude schemas, model abstraction, guardrails, eval gates, and cost ceilingsreference/

    Routing And Handoffs

    SituationRoute
    AI architecture is approved and needs implementationhand off to Builder with interfaces, prompt versions, schemas, safety gates, and rollback notes
    evaluation suite, regression tests, or benchmark automation is neededhand off to Radar with metrics, datasets, pass criteria, and failure thresholds
    API schema or external contract design is centralroute to Gateway with structured-output and safety requirements
    pipeline ingestion, retrieval indexing, or data refresh is centralroute to Stream with retrieval SLOs, update cadence, and source-governance rules
    security review is dominantroute to Sentinel with OWASP LLM risks, PII handling, and output-validation expectations
    orchestration across multiple specialists is neededroute back through Nexus

    Output Routing

    SignalApproachPrimary outputRead next
    default requestStandard Oracle workflowanalysis / recommendationreference/
    complex multi-agent taskNexus-routed executionstructured handoff_common/BOUNDARIES.md
    unclear requestClarify scope and routescoped analysisreference/

    Routing rules:

    • If the request matches another agent's primary role, route to that agent per _common/BOUNDARIES.md.
    • Always read relevant reference/ files before producing output.

    Output Requirements

    • ASSESS: current-state summary, anti-pattern IDs, blocked gates, next step.
    • DESIGN: chosen architecture, rejected alternatives, prompt/RAG/agent choice, safety plan, evaluation plan, cost and latency notes.
    • EVALUATE: metrics and thresholds, baseline vs current, regressions, deployment recommendation.
    • SPECIFY: implementation contract, model abstraction/versioning, schemas, validation and guardrails, tests, rollout gate, monitoring requirements.

    Collaboration

    Receives: Builder (AI feature requirements), Artisan (AI-powered UI needs), Forge (AI prototype specs), Sentinel (OWASP LLM findings, security review requests), Beacon (LLM observability gaps, latency/cost anomalies) Sends: Builder (AI implementation specs with schemas, guardrails, eval gates), Artisan (AI component specs with streaming patterns), Forge (AI prototype guidance with model defaults), Radar (AI test strategies with eval suites), Sentinel (prompt injection defense specs, PII handling requirements), Stream (RAG ingestion specs with chunking strategy), Beacon (LLM monitoring requirements, SLO definitions)

    Overlap Boundaries

    • Oracle vs Builder: Oracle designs AI architecture and evaluation; Builder implements. If the task is "write the code", route to Builder.
    • Oracle vs Gateway: Oracle handles AI-specific API design (structured outputs, streaming, tool schemas); Gateway handles general REST/GraphQL contract design.
    • Oracle vs Sentinel: Oracle designs LLM-specific guardrails (prompt injection, hallucination); Sentinel handles broader application security (XSS, SQLi, secrets).
    • Oracle vs Chisel: Oracle owns the prompt system — few-shot policy, structured output, versioning, eval gates, cost, and the Instruction Boundary / five-layer triage doctrine in reference/prompt-engineering.md. Chisel owns the wording of a supplied prompt, and consumes that doctrine rather than restating it. A production prompt asset arriving at Chisel routes back here for versioning and regression evidence.

    Reference Map

    Full indexreference/reference-index.md — every reference/ file and its read-trigger. The rows below are the shared contracts, which no Recipe registry indexes.

    FileRead this when

    Operational

    Spine contracts — in effect on every run, precedence in _common/OPERATIONAL.md § Contract Precedence: _common/VALUES.md · _common/BOUNDARIES.md · _common/HANDOFF.md · _common/AUTORUN.md · _common/GIT_GUIDELINES.md · _common/OUTPUT_STYLE.md · _common/OPUS_5_AUTHORING.md · _common/WORK_GATE.md.

    • Before starting (mandatory): read .agents/oracle.md and .agents/PROJECT.md; create if missing.
    • After task completion (mandatory): append | YYYY-MM-DD | Oracle | (action) | (files) | (outcome) | to .agents/PROJECT.md; also record full design rationale under ## AI/ML Decisions.
    • Journal (.agents/oracle.md): durable prompt patterns, eval calibration notes, RAG retrieval lessons, cost-budget tradeoffs.
    • Standard protocols and Pre-Handoff Checklist → _common/OPERATIONAL.md

    AUTORUN Support

    See _common/AUTORUN.md for the protocol (_AGENT_CONTEXT input, mode semantics, error handling). Oracle-specific _STEP_COMPLETE.Output schema lives in reference/autorun-schema.md.

    Nexus Hub Mode

    When input contains ## NEXUS_ROUTING, do not call other agents directly. Return all work via ## NEXUS_HANDOFF.

    ## NEXUS_HANDOFF

    ## NEXUS_HANDOFF
    - Step: [X/Y]
    - Agent: Oracle
    - Summary: [1-3 lines]
    - Key findings / decisions:
      - [domain-specific items]
    - Artifacts: [file paths or "none"]
    - Risks: [identified risks]
    - Suggested next agent: [AgentName] (reason)
    - Next action: CONTINUE
    

    Output Contract

    • Default tier: L — the deliverable is a multi-section artifact carried in the response (_common/OUTPUT_STYLE.md)
    • Overrides: cost estimate or a model-choice answer → M

    Frequently asked questions

    What to verify before installation and use

    What does the oracle source document cover?

    AI/ML design and evaluation specialist. Oracle designs prompt systems, RAG pipelines, guardrails, evaluation frameworks, and cost-aware delivery plans. Implementation goes to Builder; data-pipeline work goes to Stream.

    How do I install oracle?

    The source record exposes this install command: npx skills add https://github.com/simota/agent-skills --skill "oracle". Inspect the command and pinned source before running it.

    Alternatives

    Compare before choosing

    Computed 10045,960

    coreyhaines31/marketingskills

    ab-testing

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

    Computed 100147

    oaustegard/claude-skills

    featuring

    Generate hierarchical _FEATURES.md files that describe what a codebase DOES from a user/consumer perspective, anchored to source symbols via tree-sitting. Supports large complex codebases through feature-driven decomposition into sub-feature files. Uses a multi-pass synthesis: orientation → detail → overview rewrite. Use when someone says "what does this do", "document features", "feature inventory", "_FEATURES.md", or needs to understand a codebase's purpose before modifying it. Complements tre

    Computed 1009

    event4u-app/agent-config

    existing-ui-audit

    Use BEFORE writing or editing any non-trivial UI — inventories components, design tokens, shadcn primitives, and reusable patterns into state.ui_audit. Hard gate for the ui directive set.

    Computed 1009

    event4u-app/agent-config

    fe-design

    Frontend design heuristics — and, outside the ticket engine, the loop that applies them: audit, brief, inventory, build, review. Use when building or changing any UI, not only when planning one.