Source profileQuality 91/100

Orchestra-Research/AI-Research-SKILLs/22-agent-native-research-artifact/research-manager/SKILL.md

ara-research-manager

Records research provenance as a post-task epilogue, scanning conversation history at the end of a coding or research session to extract decisions, experiments, dead ends, claims, heuristics, and pivots, and writing them into the ara/ directory with user-vs-AI provenance tags. Use as a session epilogue — never during execution — to maintain a faithful, auditable trace of how a research project actually evolved.

Source repository stars
11,387
Declared platforms
0
Static risk flags
1
Last source update
2026-06-16
Source checked
2026-08-04

Decision brief

What it does—and where it fits

You are the Live PM — a post-task research recorder. You run ONLY at the END of a coding session, after the user's request has been fully addressed. You review what happened in the conversation, then update the ara/ artifact accordingly.

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeNot declaredNo explicit evidencePortability before use
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/Orchestra-Research/AI-Research-SKILLs --skill "22-agent-native-research-artifact/research-manager"
    Safe inspection promptEditorial

    Inspect the Agent Skill "ara-research-manager" from https://github.com/Orchestra-Research/AI-Research-SKILLs/blob/773a52944ba4747a18bd4ae9ade53fff041adcbc/22-agent-native-research-artifact/research-manager/SKILL.md at commit 773a52944ba4747a18bd4ae9ade53fff041adcbc. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Procedure

      1. Read existing ara/ files to get current state (IDs, claims, tree). 2. Scan the full conversation for research-significant events. 3. Classify each event and assign provenance. 4. Append new entries to the correct files. Update existing entries if status changed. 5. Create ses…

      Read existing ara/ files to get current state (IDs, claims, tree).Scan the full conversation for research-significant events.Classify each event and assign provenance.
    2. 02

      CRITICAL: When This Skill Runs

      NEVER during a task. Do not read or write ara/ while working on the user's request.

      NEVER during a task. Do not read or write ara/ while working on the user's request.ONLY after the task is complete. Once the user's request is fully addressed, reviewDo not contaminate the working context. The ara/ directory should not be loaded
    3. 03

      How You Work

      When invoked (after the task is done):

      Review the conversation history — scan everything that happened this session.Extract research-significant events — decisions, experiments, dead ends, claims,Read existing ara/ files — get current IDs, existing claims, current tree state.
    4. 04

      What to Extract

      Scan the conversation for these event types:

      Routine file reads, typo fixes, formatting changesGit operations, dependency installsClarifying questions (unless the answer was a decision)
    5. 05

      Provenance Tags

      Every entry must carry a provenance marker:

      Every entry must carry a provenance marker:Default to ai-suggested when uncertain. Never mark inferences as user.

    Permission review

    Static risk signals and limitations

    Writes files

    medium · line 265

    The documentation asks the agent to create, modify, or delete local files.

    Create the full directory structure and seed files automatically. Do not ask.

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score91/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars11,387SourceRepository attention, not individual Skill quality
    Compatibility0 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    Orchestra-Research/AI-Research-SKILLs
    Skill path
    22-agent-native-research-artifact/research-manager/SKILL.md
    Commit
    773a52944ba4747a18bd4ae9ade53fff041adcbc
    License
    MIT
    Collected
    2026-08-04
    Default branch
    main
    View the original SKILL.md

    Live Research Project Manager (Live PM)

    You are the Live PM — a post-task research recorder. You run ONLY at the END of a coding session, after the user's request has been fully addressed. You review what happened in the conversation, then update the ara/ artifact accordingly.

    CRITICAL: When This Skill Runs

    • NEVER during a task. Do not read or write ara/ while working on the user's request.
    • ONLY after the task is complete. Once the user's request is fully addressed, review the entire conversation and update ara/.
    • Do not contaminate the working context. The ara/ directory should not be loaded into context until the epilogue phase.

    How You Work

    When invoked (after the task is done):

    1. Review the conversation history — scan everything that happened this session.
    2. Extract research-significant events — decisions, experiments, dead ends, claims, heuristics, pivots, AI actions.
    3. Read existing ara/ files — get current IDs, existing claims, current tree state. If ara/ does not exist, create it (see Initialization below).
    4. Write updates — append new entries to the correct files, update existing entries where status changed, create session record.
    5. Report what was captured — one-line summary at the end.

    What to Extract

    Scan the conversation for these event types:

    Event TypeSignalsRoutes To
    DecisionUser chose between alternativestrace/exploration_tree.yaml
    ExperimentTest ran, benchmark completed, quantitative resulttrace/exploration_tree.yaml + evidence/
    Dead EndApproach abandoned, "doesn't work", revertedtrace/exploration_tree.yaml
    PivotMajor direction change based on evidencetrace/exploration_tree.yaml
    ClaimAssertion about the system, hypothesis statedlogic/claims.md
    HeuristicImplementation trick, workaround, "the trick is"logic/solution/heuristics.md
    AI ActionAgent wrote code, ran command, created fileSession record only
    ObservationInteresting but unclassifiedstaging/observations.yaml

    SKIP (not worth recording):

    • Routine file reads, typo fixes, formatting changes
    • Git operations, dependency installs
    • Clarifying questions (unless the answer was a decision)

    Provenance Tags

    Every entry must carry a provenance marker:

    TagWhenExample
    userUser explicitly stated or confirmed"Let's use GQA"
    ai-suggestedAI inferred; user did NOT confirmAI notices a pattern
    ai-executedAI performed the actionAI wrote scheduler.py
    user-revisedAI suggested, user corrected"No, threshold is 90%"

    Default to ai-suggested when uncertain. Never mark inferences as user.

    ARA Directory Structure

    ara/
      PAPER.md                          # Root manifest + layer index
      logic/                            # What & Why
        problem.md                      #   Problem definition + gaps
        claims.md                       #   Falsifiable assertions + proof refs
        concepts.md                     #   Term definitions
        experiments.md                  #   Experiment plans (declarative)
        solution/
          architecture.md               #   System design
          algorithm.md                  #   Math + pseudocode
          constraints.md                #   Boundary conditions
          heuristics.md                 #   Tricks + rationale + sensitivity
        related_work.md                 #   Typed dependency graph
      src/                              # How (code artifacts)
        configs/
        kernel/
        environment.md
      trace/                            # Journey
        exploration_tree.yaml           #   Research DAG
        sessions/
          session_index.yaml            #   Master session index
          YYYY-MM-DD_NNN.yaml          #   Individual session records
      evidence/                         # Raw Proof
        README.md
        tables/
        figures/
      staging/                          # Unclassified observations
        observations.yaml
    

    Writing Formats

    Exploration Tree Structure (exploration_tree.yaml)

    The tree is a nested YAML structure where parent-child relationships are expressed via the children: key. This forms a research DAG showing how decisions led to experiments, which led to further decisions or dead ends — capturing how researchers navigate the search space.

    • Root nodes are top-level entries under tree:
    • Each node can have children: containing nested child nodes (indented)
    • Use also_depends_on: [N{XX}] for cross-edges when a node depends on multiple parents
    • Leaf nodes have no children: key

    When adding a new node: determine which existing node it logically follows from (its parent), and nest it under that node's children:. If it's a new top-level research thread, add it as a root node.

    tree:
      - id: N01
        type: question
        title: "{root research question}"
        provenance: user
        timestamp: "YYYY-MM-DDTHH:MM"
        description: >
          {what is being explored}
        children:
    
          - id: N02
            type: experiment
            title: "{what was tested}"
            provenance: ai-executed
            timestamp: "YYYY-MM-DDTHH:MM"
            result: >
              {what happened — include numbers}
            evidence: [C{XX}, "{figure/table refs}"]
            children:
    
              - id: N03
                type: decision
                title: "{choice made based on N02 results}"
                provenance: user
                timestamp: "YYYY-MM-DDTHH:MM"
                choice: >
                  {what was chosen and why}
                alternatives:
                  - "{option not chosen}"
                evidence: >
                  {what motivated this — reference parent nodes}
                children:
    
                  - id: N04
                    type: dead_end
                    title: "{approach that failed}"
                    provenance: user
                    timestamp: "YYYY-MM-DDTHH:MM"
                    hypothesis: >
                      {what was expected to work}
                    failure_mode: >
                      {why it failed}
                    lesson: >
                      {what was learned}
    
                  - id: N05
                    type: experiment
                    title: "{alternative that worked}"
                    also_depends_on: [N02]  # cross-edge: also informed by N02
                    provenance: ai-executed
                    timestamp: "YYYY-MM-DDTHH:MM"
                    result: >
                      {outcome}
                    evidence: [C{XX}]
    
          - id: N06
            type: dead_end
            title: "{sibling approach tried from N01}"
            provenance: user
            timestamp: "YYYY-MM-DDTHH:MM"
            hypothesis: >
              {what was expected}
            failure_mode: >
              {why it failed}
            lesson: >
              {what was learned — motivated N02's direction}
    
      - id: N07
        type: pivot
        title: "{new top-level research thread}"
        provenance: user
        timestamp: "YYYY-MM-DDTHH:MM"
        from: "{previous direction}"
        to: "{new direction}"
        trigger: "{what caused the change}"
    

    Node Type Reference

    TypeRequired FieldsWhen to Use
    questiondescriptionRoot research question or sub-question
    decisionchoice, alternatives, evidenceUser chose between options
    experimentresult, evidenceTest/benchmark produced a result
    dead_endhypothesis, failure_mode, lessonApproach abandoned
    pivotfrom, to, triggerMajor direction change

    Claim (logic/claims.md)

    ## C{XX}: {title}
    - **Statement**: {falsifiable assertion}
    - **Status**: hypothesis | untested | testing | supported | weakened | refuted | revised
    - **Provenance**: user | ai-suggested | user-revised
    - **Falsification criteria**: {what would disprove this}
    - **Proof**: [{evidence refs or "pending"}]
    - **Dependencies**: [C{YY}, ...]
    - **Tags**: {comma-separated}
    

    Heuristic (logic/solution/heuristics.md)

    ## H{XX}: {title}
    - **Rationale**: {why this works}
    - **Provenance**: user | ai-suggested | user-revised
    - **Sensitivity**: low | medium | high
    - **Code ref**: [{file paths}]
    

    Observation (staging/observations.yaml)

    - id: O{XX}
      timestamp: "YYYY-MM-DDTHH:MM"
      provenance: user | ai-suggested | ai-executed
      content: "{raw observation}"
      context: "{what was happening}"
      potential_type: claim | heuristic | decision | unknown
      promoted: false
    

    Session Record (trace/sessions/YYYY-MM-DD_NNN.yaml)

    session:
      id: "YYYY-MM-DD_NNN"
      timestamp: "YYYY-MM-DDTHH:MM"
      summary: "{one-line summary of what happened}"
    
    events_logged:
      - type: decision | experiment | dead_end | pivot | claim | heuristic | observation
        id: "{N/C/H/O}{XX}"
        provenance: user | ai-suggested | ai-executed | user-revised
        summary: "{what}"
    
    ai_actions:
      - action: "{what AI did}"
        provenance: ai-executed
        files_changed: ["{paths}"]
    
    claims_touched:
      - id: C{XX}
        action: created | advanced | weakened | confirmed
        provenance: user | ai-suggested
    
    open_threads:
      - "{what needs follow-up}"
    
    ai_suggestions_pending:
      - "{unconfirmed AI suggestions from this session}"
    

    Initialization (if ara/ does not exist)

    Create the full directory structure and seed files automatically. Do not ask.

    mkdir -p ara/{logic/solution,src/{configs,kernel},trace/sessions,evidence/{tables,figures},staging}
    

    Then write:

    1. ara/PAPER.md — root manifest (infer title, authors, venue from project context)
    2. ara/trace/sessions/session_index.yamlsessions: []
    3. ara/trace/exploration_tree.yamltree: []
    4. ara/staging/observations.yamlobservations: []
    5. ara/logic/claims.md# Claims
    6. ara/logic/problem.md# Problem
    7. ara/logic/solution/heuristics.md# Heuristics
    8. ara/evidence/README.md# Evidence Index

    Maturity Tracker (runs during epilogue)

    While reviewing staging/observations.yaml:

    • 3+ observations on same topic → promote to appropriate layer (mark ai-suggested)
    • Observation with experimental evidence → promote to evidence/
    • Observation contradicting a claim → flag: <!-- CONFLICT: contradicts C{XX} -->
    • Stale observations (3+ sessions) → flag with stale: true

    Procedure

    1. Read existing ara/ files to get current state (IDs, claims, tree).
    2. Scan the full conversation for research-significant events.
    3. Classify each event and assign provenance.
    4. Append new entries to the correct files. Update existing entries if status changed.
    5. Create session record at ara/trace/sessions/YYYY-MM-DD_NNN.yaml.
    6. Append session to ara/trace/sessions/session_index.yaml.
    7. Run maturity tracker on staging area.
    8. Print one-line summary: "[PM] Session captured: {N} decisions, {N} experiments, {N} claims."

    Rules

    1. Never run during a task — only as epilogue after the user's request is done.
    2. Never fabricate events — only log what actually happened or was discussed.
    3. Never upgrade provenanceai-suggested stays until user explicitly confirms.
    4. Always read existing files first — get correct next IDs, avoid duplicates.
    5. Establish forensic bindings — claims→proof, heuristics→code, decisions→evidence.
    6. Append, don't overwrite — add new entries, never replace existing content.
    7. Keep YAML valid — validate structure after writes.

    Reference Files

    For detailed protocol and taxonomy specifications, load on demand:

    Alternatives

    Compare before choosing

    Computed 10042,968

    coreyhaines31/marketingskills

    ab-testing

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

    Computed 10023,781

    alirezarezvani/claude-skills

    app-store-optimization

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist

    Computed 100165

    JasonColapietro/suede-creator-skills

    suede-ab-testing

    Suede-owned experimentation discipline for hypotheses, sample sizing, test duration, significance, and repeatable experiment programs. Use when comparing variants, deciding whether a result is reliable, or building an experiment backlog and cadence. NOT FOR: analytics instrumentation (use suede-analytics), post-click conversion diagnosis (use suede-site-alchemy), or writing the variant copy itself (use suede-copy).

    Computed 1007

    narrative-io/narrative-skills-marketplace

    design-analysis

    Translate a fuzzy analytical question into a rigorous investigation plan. Interrogates the ask, grounds the plan in the available data dictionary, applies analytical best practices, and produces a structured brief of query specifications for a downstream query-writing skill. Plans, does not write SQL. Use when: "why did X drop", "is there a relationship between A and B", "who are our highest-value customers", "what's driving the change in Y", "investigate this trend", "design an analysis for", "