Best for
- User asks to create a new disorder/disease entry
- User names a disorder that doesn't exist in kb/disorders/
monarch-initiative/dismech/.claude/skills/initiate-new-disorder-creation/SKILL.md
Skill for initiating new disorder YAML files in the dismech knowledge base. Use this skill when the user asks to create a new disorder entry. Also useful for enhancing existing entries.
Decision brief
Skill for initiating new disorder YAML files in the dismech knowledge base. Also useful for enhancing existing entries.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/monarch-initiative/dismech --skill ".claude/skills/initiate-new-disorder-creation"Inspect the Agent Skill "initiate-new-disorder-creation" from https://github.com/monarch-initiative/dismech/blob/8fd58adcb26220902c524da7be9d8aa7fc215e18/.claude/skills/initiate-new-disorder-creation/SKILL.md at commit 8fd58adcb26220902c524da7be9d8aa7fc215e18. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Choose the clinically preferred name for the disorder, use title case (e.g. Foo Bar Syndrome). For file names, spaces will. be replaced by underscores, and characters such as apostrophes removed.
Choose the clinically preferred name for the disorder, use title case (e.g. Foo Bar Syndrome). For file names, spaces will. be replaced by underscores, and characters such as apostrophes removed.
The preferred mode of working is to use git worktrees, unless the user has expressed a preference not to do this in advance.
Create an initial yaml file using the underscore form of the disease, e.g.
Execute at least one deep research query. Always do this via the just command, do not perform your own deep research.
Permission review
The documentation asks the agent to run terminal commands or scripts.
git fetch origin mainThe documentation asks the agent to run terminal commands or scripts.
git grep -n -i -e "<MONDO_ID>" -e "<preferred disorder name>" origin/main -- kb/disorders || trueThe documentation asks the agent to create, modify, or delete local files.
### Step 2b: Create initial YAML fileThe documentation asks the agent to create, modify, or delete local files.
Create an initial yaml file using the underscore form of the disease, e.g.The documentation includes network, browsing, or remote request actions.
curl -sG "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi" \The documentation includes network, browsing, or remote request actions.
##### Search the PubMed APIEvidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 93/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 56 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Guide the creation of new disorder YAML files in the dismech knowledge base. This skill emphasizes a research-first approach to ensure scientific accuracy and prevent AI hallucinations by requiring deep research queries before file creation.
kb/disorders/This skill can also be consulted for ongoing curation of existing disorders.
Choose the clinically preferred name for the disorder, use title case (e.g. Foo Bar Syndrome).
For file names, spaces will. be replaced by underscores, and characters such as apostrophes removed.
Before creating a new disorder file, check all three duplicate surfaces:
origin/main and search the current
upstream disorder YAMLs by MONDO ID, preferred label, and important synonyms.Use specific identifiers first, then human-readable labels and synonyms:
git fetch origin main
# Knowledgebase on the latest origin/main, not just the local working tree.
git grep -n -i -e "<MONDO_ID>" -e "<preferred disorder name>" origin/main -- kb/disorders || true
git grep -n -i -e "<important synonym>" origin/main -- kb/disorders || true
# PRs and issues across all states.
gh pr list --repo monarch-initiative/dismech --state all \
--search "\"<MONDO_ID>\" OR \"<preferred disorder name>\"" \
--json number,title,state,url,headRefName --limit 100
gh issue list --repo monarch-initiative/dismech --state all \
--search "\"<MONDO_ID>\" OR \"<preferred disorder name>\"" \
--json number,title,state,url,labels --limit 100
Repeat the PR and issue searches for important synonyms if the first search is empty. If the disorder already exists in the knowledgebase, edit the existing file instead of creating a new one. If an open PR or issue already covers the same disorder, continue there rather than starting duplicate work. If a closed PR or issue appears relevant, inspect it before deciding whether new curation is still needed.
The preferred mode of working is to use git worktrees, unless the user has expressed a preference not to do this in advance.
Create an initial yaml file using the underscore form of the disease, e.g.
kb/disorders/Foo_Bar.yaml:
name: Foo Bar
creation_date: "2025-06-12T20:16:27Z"
updated_date: "2025-06-12T20:16:27Z"
category: Complex
disease_term:
term:
id: MONDO:nnnnnnn
label: foo bar ## mondo name will follow OBO case conventions
parents:
<yaml list of strings>
has_subtypes:
<optional yaml list of Subtype objects>
pathophysiology:
<yaml list of Pathophysiology objects>
phenotypes:
<yaml list of Phenotype objects>
biochemical:
<optional yaml list of Biochemical objects>
genetic:
<optional yaml list of Genetic objects>
environmental:
<optional yaml list of Environmental objects>
treatments:
<optional yaml list of Treatment objects>
datasets:
creation_date and updated_date must be ISO 8601/RFC 3339 datetime strings.
When editing an existing file, preserve creation_date and bump updated_date.
The objects must follow the LinkML schema in src/dismech/schema.
It can be validated with just validate kb/disorders/Foo_Bar.yaml
This first pass should use textbook knowledge about the disease: you will later refine this.
Execute at least one deep research query. Always do this via the just command, do
not perform your own deep research.
Depending on user preference, use one or more of the following commands
just research-disorder asta DISORDER_NAMEjust research-disorder perplexity DISORDER_NAMEjust research-disorder falcon DISORDER_NAMEjust research-disorder openai DISORDER_NAMEjust research-disorder cyberian DISORDER_NAMEjust research-disorder openscientist DISORDER_NAMEjust research-disorder claude_code DISORDER_NAMEUse the filesystem-friendly name here.
claude_code needs no separate API key. It wraps the local claude CLI as
a subprocess (claude --print --output-format json), reusing the Claude Code
credential that is already present — CLAUDE_CODE_OAUTH_TOKEN is exported in
both dragon-ai.yml and curation-scanner.yml, so it works in CI with no new
secret. The provider auto-enables whenever claude is on PATH (disable it
with DISABLE_CLAUDE_CODE_PROVIDER=true). For security it restricts the
subprocess to read-only research tools (WebSearch, WebFetch) — no
filesystem mutations from the research prompt. It captures run provenance (the
model used, cost, turn count, web-search count) in a run_metadata field and
forces the report inline rather than deferring to a background workflow
artifact. This makes it the natural default when you are already inside a Claude
Code agentic session and just want a web-grounded report without standing up
extra credentials. Requires deep-research-client >= 0.2.7.
falcon requires EDISON_API_KEY or FUTUREHOUSE_API_KEY to be exported
in the environment — both names refer to the same key and are accepted
interchangeably by deep-research-client. The provider was originally named
"FutureHouse Falcon" and later rebranded as "Edison Scientific"; the falcon
provider slug in just research-disorder is unchanged. Use whichever key name
your environment/secrets manager provides; if you have FUTUREHOUSE_API_KEY
that is sufficient. Edison is a large-scale literature agent that performs
deep bibliographic research. falcon runs may take 20 minutes or longer.
In addition to the narrative report, Edison runs frequently produce artifacts
— structured tables, figures, or supplementary documents — that summarise key
findings in machine-readable form. When EDISON_API_KEY is set, artifact
retrieval happens automatically at the end of just research-disorder falcon ….
Artifacts are written to a sub-directory alongside the report:
research/DISORDER_NAME-deep-research-falcon_artifacts/
The report's YAML frontmatter records the trajectory_id used to retrieve them
and lists each artifact under the artifacts key. An ## Artifacts section
is inserted into the report body for any image artifacts so they render in
Markdown viewers. If artifact retrieval was skipped (e.g. EDISON_API_KEY was
not set at the time), you can run it later with:
just fetch-research-artifacts <trajectory_id> research/DISORDER_NAME-deep-research-falcon.md
asta requires ASTA_API_KEY to be exported in the environment. Asta behaves
more like a literature search agent than a full narrative deep-research agent:
its outputs are primarily lists of relevant papers, usually with summaries,
evidence snippets, and relevance scores. The just research-disorder asta ...
command automatically uses an Asta-specific template tailored for this output
style.
openscientist requires OPENSCIENTIST_API_KEY to be exported in the
environment. OpenScientist (https://www.openscientist.io) is an autonomous AI
research agent from Berkeley Lab that runs iterative hypothesis-driven research
using PubMed search and code execution. It produces markdown reports with PMID
citations. To set up:
name:secret format)export OPENSCIENTIST_API_KEY="name:secret"OpenScientist jobs are asynchronous — the provider submits a job, then polls
until completion. Jobs are queued server-side and processed sequentially, so
wait times depend on queue depth. The API's /report endpoint returns PDF;
the provider automatically extracts the markdown final_report.md from the
/artifacts ZIP. The artifacts ZIP also contains provenance data (iteration
transcripts, generated plots as PNG/JSON) and agent logs.
Timing varies by provider. As a rule of thumb:
asta usually completes in secondsopenai and perplexity usually complete within a few minutesfalcon may take 20 minutes or longercyberian runtime varies with workflow complexity and can also be long-runningopenscientist typically takes 10–30 minutes depending on queue depth and iteration countclaude_code typically completes in a few minutes — it runs a single bounded
agentic session of web searches/fetches rather than a long iterative pipeline,
so it is usually faster than falcon/openscientist but with web-grounded
(not exhaustive bibliographic) coverage. A representative run (Sarcoidosis,
#4761) took ~4m50s, did 11 web searches over 13 turns, returned 24 citations,
and cost ~$2 in Claude Code usage. The report's YAML frontmatter records this
provenance under run_metadata (models used, web_search_requests,
num_turns, total_cost_usd, session_id).On completion, this will create a file here:
./research/DISORDER_NAME-deep-research-PROVIDER.md
and a separate citations file here:
./research/DISORDER_NAME-deep-research-PROVIDER.md.citations.md
For Edison (falcon) runs, artifacts (figures, structured tables, etc.) are also saved in:
./research/DISORDER_NAME-deep-research-falcon_artifacts/
and referenced in the report's YAML frontmatter under the artifacts key.
For example:
research/Urticaria-deep-research-openai.mdresearch/Urticaria-deep-research-openai.md.citations.mdresearch/Urticaria-deep-research-falcon_artifacts/ (falcon only)You MUST read this before progressing.
Every just research-* recipe now resolves the report's citations as part of
generating it (deep-research-client >= 0.2.10, backed by the same
linkml-reference-validator the KB validators use). The answer is already in
the report — read it before you cite anything from that report.
Two places to look:
The frontmatter carries a machine-readable summary:
reference_validation:
total_references: 24
verified: 22
not_found: 2
confabulation_rate: 0.083
quotes_checked: 9
quotes_valid: 8
relevance_assessed: 22
on_topic: 19
off_topic: 1
off_topic_references:
- PMID:28123456
unresolved_references:
- PMID:99999999
needs_review: true
A ## Reference Validation section at the end of the body, with a counts
table, an ### Unresolved references list naming each failing identifier, and
a ### References that may not be about this subject list.
What to do with it:
needs_review first. It is the one key that cannot give you a false
all-clear: it is set when any identifier failed to resolve, or any quote
failed to match, or any reference looks off topic. Do not read
confabulation_rate as the whole-report signal — it measures identifier
resolution and nothing else, so a report whose every PMID exists but whose
quotes do not match still reports 0.0.unresolved_references — do not cite it. Either find a
different source for the claim or drop the claim. Do not "verify it yourself"
by fetching it again and moving on if it happens to work the second time
without saying so; if you do re-check one, say in the history record which
identifiers you re-checked and what you found.confabulation_rate (say, above ~0.1) is a signal about the whole
report's identifiers, not just the listed ones. Treat the rest of it with
extra suspicion and prefer claims you can independently anchor. A low one
clears nothing else.quotes_valid < quotes_checked means the report attributed a quote to a
paper that does not contain it. Read which one before reusing any quoted
material from that report.off_topic_references resolved, so it is not a fabrication —
it just shares almost none of the report's vocabulary. That is evidence, not
a verdict: read the paper before citing or dropping it, since a paper can be
relevant in ways its title and abstract do not spell out. Note also that
off_topic: 0 is not "all cleared" — a record with no abstract can never be
called off topic, so some references are simply undecided.Reports generated before this existed (most of research/) have no
validation section. Add one:
just validate-research-reference research/DISORDER_NAME-deep-research-PROVIDER.md
That rewrites the report in place with a ## Reference Validation section (it
does not add a frontmatter summary — on a retro-fitted report, read the
section at the bottom). Re-running is safe.
This does not replace anything downstream. It checks the report's
citations. The snippet you paste into the KB entry is a different quote in a
different file and still needs the normal checks (Step 4 onwards), and none of
it catches Named Entity Confusion — run just preflight-dr as usual. The
relevance check is not a substitute for that: references are scored against
the report's own vocabulary, so a report built around the wrong disease has
wrong-disease vocabulary too and scores all of its wrong-disease citations as
on topic. See
docs/deep-research-reference-validation.md.
The same recipes also resolve every ontology CURIE the report suggests
(deep-research-client >= 0.2.11, backed by the same linkml-term-validator
just validate-terms runs). This is a separate check from the one above, and it
catches a different failure: the CMTX report in
#9729 had 26/26
citations verified and still offered MONDO:0010674 — Hunter syndrome — as the
Charcot-Marie-Tooth X-linked term.
Two places to look, as before: a term_validation: frontmatter block and a
## Term Validation section at the end of the body.
What to do with it:
unresolved_terms. It does not exist. Find
the right term with the dismech-terms skill instead.needs_review, not confabulation_rate. The rate measures identifier
resolution only, so a report whose every CURIE resolves but whose labels name
different terms still shows 0.0.mislabelled_terms is where the wrong bindings surface. Each entry gives
the report's name and the ontology's. Some are harmless paraphrase ("distal
weakness" for HP:0002460, Distal muscle weakness). Look for the ones where
the ontology label names a different disease, or a different term in the
same ontology — a sibling, a parent, a near-miss. The second kind is easy to
skim past: the CMTX report writes "areflexia" beside HP:0001265, which HPO
calls Hyporeflexia (Areflexia is HP:0001284), and those are clinically
distinct.unresolvable_prefixes means nothing was checked for that prefix — not that
anything is wrong. HGNC is skipped by default; verify gene CURIEs the usual
way.Reports generated before this existed have no term section. Add one:
just validate-research-terms research/DISORDER_NAME-deep-research-PROVIDER.md
Like the reference retro-fit, this adds the markdown section but not a
frontmatter summary, and re-running is safe. See
docs/deep-research-term-validation.md.
GeneReviews (https://www.ncbi.nlm.nih.gov/books/NBK1116/) is the authoritative expert-curated clinical reference for Mendelian disorders. Before curating phenotypes, you MUST check whether a GeneReviews article exists for the disease. If one exists, it is the mandatory phenotype baseline — not just a convenient source.
Scope: This step applies primarily to Mendelian (single-gene) disorders. For complex, multifactorial, infectious, or cancer entries where GeneReviews coverage is unlikely, skip directly to Step 4 — the PubMed search below will confirm either way.
curl -sG "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi" \
--data-urlencode "db=pubmed" \
--data-urlencode "retmode=json" \
--data-urlencode "term=<DISEASE_NAME>[TI] GeneReviews[TI]"
If no results, try a broader search: <DISEASE_NAME> GeneReviews[All Fields]
just fetch-reference PMID:XXXXXXXX
references: blockreferences:
- reference: PMID:XXXXXXXX
title: "<GeneReviews article title>"
tags:
- GeneReviews
You can also run just tag-references after adding the PMID to inline
evidence items — the script detects GeneReviews PMIDs from the cached
abstract and writes the top-level tag automatically.
references_cache/PMID_XXXXXXXX.mdphenotypes: sectionNote: The cached abstract captures only the structured PubMed abstract, which is a condensed summary of the full GeneReviews chapter. The full Clinical Characteristics section in the chapter body often lists additional phenotypes not in the abstract. Cross-reference the deep-research artifact (from Step 3) for comprehensive coverage — treat the abstract as the minimum baseline, not the ceiling.
GeneReviews often has an Agents/Circumstances to Avoid section.
If the abstract mentions any, add a note in the relevant treatment entry's
description: and include a GeneReviews evidence item quoting it exactly.
GeneReviews uses narrative frequency language. Map to the enum as follows:
| GeneReviews phrase | FrequencyEnum | HPO range |
|---|---|---|
| "virtually all", "most individuals", ">80%" | VERY_FREQUENT | 80–100% |
| "many", "majority", "common", "~50%–79%", ">30%" | FREQUENT | 30–79% |
| "some", "occasional", "uncommon", "~5%–29%" | OCCASIONAL | 5–29% |
| "rare", "few", "<5%", "infrequently reported" | VERY_RARE | 1–4% |
| "isolated reports", "single case" | (omit frequency) | <1% |
When frequency is ambiguous, omit frequency: rather than guessing.
If no GeneReviews article exists for the disease, proceed to Step 4 without this baseline. No action needed — the absence itself is not a problem.
Use the results of deep research to enhance the yaml file, providing evidence for as many assertions as possible.
Find the pubmed IDs or DOIs for the papers in the deep research and retrieve these:
just fetch-reference PMID:nnnnnnnjust fetch-reference DOI:...The just fetch-reference command can accept multiple identifiers of different types, such as:
just fetch-reference PMID:nnnnnnn DOI:nn.nnnnYou can also find additional references relevant to individual assertions, on top of what is in the deep research.
Note that a validated report (Step 3a) has already fetched most of these — its
lookups are cached into the same references_cache/ — so just fetch-reference
on a reference the report resolved is a cache hit and returns immediately. Run it
anyway rather than assuming; it costs nothing when the file is already there, and
it is still required for any reference you found outside the report.
When an Edison (falcon) run produces artifacts, check whether any images in the
artifact directory directly support a specific evidence claim you are curating.
If so, include the image path in the images slot on the evidence item.
CRITICAL relevance rule: Only include an image if it directly illustrates the specific claim made in that evidence item. Do NOT include images for general background, unrelated figures, or mere "this might be interesting" reasons. Every listed image must be clearly connected to the snippet or explanation it accompanies.
To check available artifacts for a falcon report:
ls research/DISORDER_NAME-deep-research-falcon_artifacts/
The images slot is a list of paths relative to the research/ directory:
evidence:
- reference: PMID:35533128
supports: SUPPORT
evidence_source: HUMAN_CLINICAL
snippet: "Exactly quoted text from the abstract..."
explanation: "Why this supports the claim."
images:
- Dimethylglycine_Dehydrogenase_Deficiency-deep-research-falcon_artifacts/figure-01.png
Multiple images per evidence item are allowed when each is distinctly relevant:
evidence:
- reference: PMID:35533128
supports: SUPPORT
snippet: "..."
images:
- MyDisorder-deep-research-falcon_artifacts/pathway-diagram.png
- MyDisorder-deep-research-falcon_artifacts/clinical-data-table.png
Do not invent image paths. Only reference files that actually exist in the
artifact directory and have been committed to the repository. Non-image
artifacts (e.g., .md tables, .json data) should generally not be listed
under images; they are already linked in the report's ## Artifacts section.
Use PubMed first whenever possible. Use Semantic Scholar as a backup discovery tool if PubMed search is not finding the paper you want.
Use the NCBI E-utilities API to search PubMed directly. A typical workflow is:
esearchesummaryefetch if neededjust fetch-referenceExample search:
curl -sG "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi" \
--data-urlencode "db=pubmed" \
--data-urlencode "retmode=json" \
--data-urlencode "term=<SEARCH_TERMS>"
Example summary lookup once you have one or more PMIDs:
curl -sG "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi" \
--data-urlencode "db=pubmed" \
--data-urlencode "retmode=json" \
--data-urlencode "id=<PMID1>,<PMID2>"
Example abstract fetch:
curl -sG "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi" \
--data-urlencode "db=pubmed" \
--data-urlencode "id=<PMID>" \
--data-urlencode "rettype=abstract" \
--data-urlencode "retmode=text"
Once you identify the paper you want, cache it with:
just fetch-reference PMID:nnnnnnn
If PubMed search is sparse or the deep research output only gives you a title, DOI, or Semantic Scholar paper ID, use the Semantic Scholar API to find the paper and recover identifiers.
Important: the Semantic Scholar search endpoint may return 429 Too Many Requests without an API key, even for simple queries. If you have a
Semantic Scholar API key, include it as an x-api-key header. If you do not,
use Semantic Scholar as a secondary/manual fallback rather than your primary
search method.
Example paper search (often requires an API key):
curl -sG "https://api.semanticscholar.org/graph/v1/paper/search" \
-H "x-api-key: <SEMANTIC_SCHOLAR_API_KEY>" \
--data-urlencode "query=<SEARCH_TERMS>" \
--data-urlencode "limit=10" \
--data-urlencode "fields=title,year,externalIds,url"
Example lookup by Semantic Scholar paper ID (this should work without an API key):
curl -s "https://api.semanticscholar.org/graph/v1/paper/<S2_PAPER_ID>?fields=title,year,externalIds,url"
Prefer records that expose externalIds such as DOI, PubMed, or PMC.
If you have a DOI, use the PMC ID Converter API to recover PubMed/PMC IDs when available:
curl -sG "https://pmc.ncbi.nlm.nih.gov/tools/idconv/api/v1/articles/" \
--data-urlencode "ids=<DOI>" \
--data-urlencode "format=json" \
--data-urlencode "tool=dismech" \
--data-urlencode "email=<YOUR_EMAIL>"
This can return:
pmidpmciddoiIf you start from a Semantic Scholar paper ID:
externalIdsPubMed is present, use that PMID directlyPMC is present, keep the PMCID as supporting metadataDOI is present, run the DOI through the PMC ID Converter API aboveThen cache the article locally with whichever identifier you recovered:
just fetch-reference PMID:nnnnnnn
just fetch-reference DOI:10.xxxx/xxxxx
Use PMID-based references in YAML evidence whenever possible. Keep PMCID as useful supporting metadata, but DisMech evidence validation is centered on PMID abstracts.
Then use this to provide snippets/excerpts and explanations for assertions. For example, for a phenotype assertion:
phenotypes:
- name: <Phenotype Name>
description: <Description from research>
evidence:
- reference: PMID:XXXXXXXX
supports: <SUPPORT | REFUTE | PARTIAL>
evidence_source: <HUMAN_CLINICAL | MODEL_ORGANISM | IN_VITRO | COMPUTATIONAL>
snippet: "<Exact quote from abstract>"
explanation: "<Why this supports the phenotype>"
IMPORTANT: The evidence_source field classifies the type of evidence in the cited publication (human study, animal model, cell culture, computational simulation), NOT whether the curation was performed by an AI agent. Always classify based on what the paper reports, regardless of who or what is doing the curation.
The same generic evidence list schema is used for most types.
Add term objects using ontology term IDs; for example, for a pathophsyiology object, it might look like this:
pathophysiology:
- name: <Mechanism Name>
description: >
<Detailed mechanism description from research>
cell_types:
- preferred_term: <Cell Type>
term:
id: CL:XXXXXXX
label: <exact CL label>
biological_processes:
- preferred_term: <Process Name>
term:
id: GO:XXXXXXX
label: <exact GO label>
Consult the LinkML schema to see what terms are appropriate for any given object type. These will be validated.
You can use OAK commands to find relevant terms.
General term search (use mondo for diseases)
uv run runoak -i sqlite:obo:mondo info "l~<disease name>"
starts-with queries (use hp for phenotypes)
uv run runoak -i sqlite:obo:hp info "l^<phenotype>"
exact:
uv run runoak -i sqlite:obo:cl info CL:nnnnnnn
relationships (up and down):
uv run runoak -i sqlite:obo:go relationships --direction both GO:nnnnnnn
Strict validation check (adherence to schema, term and reference checks):
just validate kb/disorders/<Disease_Name>.yaml
Compliance report (completeness, term and evidence coverage):
just compliance kb/disorders/<Disease_Name>.yaml
Add an append-only curation history record so provenance keeps pace with the KB entry. CI posts an advisory warning when a KB entry changes without one. Scaffold it with the helper (never hand-write the path, timestamp, or session id):
just new-history --kind disorder --slug <Disease_Name> \
--event CREATE --outcome changed \
--summary "Create: <Disease_Name>" \
--agent-tool claude-code --model <model-id> \
--sections phenotypes,pathophysiology,evidence,treatments \
--details "What was curated, which deep-research provider(s) were used, and how it was validated."
Use --event EDIT when augmenting an existing entry. Then validate and stage it:
just validate-history <path-printed-by-new-history>
git add history/
See docs/history.md for the full format and event/outcome vocabularies.
Use the dismech-pr-review/ to do an initial round of review. Use a subagent for fresh context
(note that we haven't made the PR yet, but we want to do our own "red team" before making the actual PR
IF THE USER asks, then go ahead and make a PR on behalf of the user
Some time (~5 mins) after making the PR, a review will appear. You should prioritize in order:
Follow these priorities but use judgment. If something doesn't sit right, ask for clarification on the PR. Be proactive. If the review says "moderate" go ahead and fix it as you are fixing things anyway. Ignore things that seem super-minor but if there is no cost in making a fix and you agree, do it.
Once you have made the changes:
Stage ONLY disorder-relevant files — never use git add -A or git add .:
git add kb/disorders/ references_cache/ research/
This prevents committing unrelated generated files (HTML, schema docs, cache CSVs) that cause merge conflicts.
Commit and push:
git commit --no-verify -m "feat: Add <Disease Name> (<gene>) with deep research and validated evidence"
git push
Post a PR comment summarizing what you did:
Convert the disease name to a file-safe format:
Examples:
Type_2_Diabetes.yamlAlzheimers_Disease.yamlCOVID-19.yamlA new disorder file MUST include at minimum:
| Field | Source | Notes |
|---|---|---|
name | - | Human-readable disease name |
category | Research | Mendelian, Complex, Infectious, etc. |
disease_term | OAK lookup | MONDO term binding |
phenotypes (1+) | Research | At least one phenotype with HPO term |
pathophysiology (1+) | Research | At least one mechanism |
evidence (1+) | Research | At least one PMID reference |
All evidence items MUST:
evidence_source based on the publication's evidence type (human clinical, animal model, in vitro, computational), NOT based on whether an AI agent performed the curationNEVER fabricate PMIDs or paraphrase snippets.
Evidence Source Classification: When adding evidence_source, ask "What kind of study does this paper report?" not "How was this entry curated?" A computational fluid dynamics study gets COMPUTATIONAL, a mouse model study gets MODEL_ORGANISM, a human clinical trial gets HUMAN_CLINICAL - regardless of whether the curation was done by a human or an AI agent.
info "l~<term>"just count-verified-snippets <file> —
seconds, offline, and it names each snippet it could not findjust validate-references <file> is the slow full check; use
--fix-threshold 0.80 there to auto-repair minor mismatchesname, category, and at least one pathophysiology entryUse all loaded skills, including:
When asked to address review comments on an existing PR:
supports: PARTIAL when evidence is indirect — don't overstate evidence strength# ONLY stage disorder-relevant files
git add kb/disorders/ references_cache/ research/ history/
# NEVER do this — picks up generated files from other disorders
# git add -A
# git add .
git commit --no-verify -m "fix: Address PR review comments"
git push
Before finalizing a new disorder file, verify:
GeneReviews in top-level references:just validate passesjust validate-terms passesjust count-verified-snippets <file> reports N/N verified (fast, per-edit)just validate-disorders <every changed file> passes — one batched run at
the end, the same check CI runs. Tick this only after reading its output;
naming a check that was killed partway is what #8119 was filed about.just new-history) and just validate-history passesFrequently asked questions
Skill for initiating new disorder YAML files in the dismech knowledge base. Also useful for enhancing existing entries.
The source record exposes this install command: npx skills add https://github.com/monarch-initiative/dismech --skill ".claude/skills/initiate-new-disorder-creation". Inspect the command and pinned source before running it.
Static rules flagged exec-script, write-files, network in the source; the page lists the matching lines and excerpts.
Alternatives
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
wanshuiyin/Auto-claude-code-research-in-sleep
Use it for operations and research tasks; the detail page covers purpose, installation, and practical steps.
prowler-cloud/prowler
PostgreSQL indexing best practices for Prowler: index design, partial indexes, partitioned table indexing, EXPLAIN ANALYZE validation, concurrent operations, monitoring, and maintenance. Trigger: When creating or modifying PostgreSQL indexes, analyzing query performance with EXPLAIN, debugging slow queries, reviewing index usage statistics, reindexing, dropping indexes, or working with partitioned table indexes. Also trigger when discussing index strategies, partial indexes, or index maintenance
brucesongs/kali-claw
Insecure Design (OWASP A06:2025) focuses on security flaws in system architecture and design phases, rather than code implementation-level bugs.