K-Dense-AI/scientific-agent-skills

scholar-evaluation

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

88Collecting
See how to use itView GitHub source
npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/scholar-evaluation"
Automated source guide

Source checked Jul 28, 2026·Refresh due Oct 26, 2026

Reorganized from the pinned upstream SKILL.md

Turn scholar-evaluation's source instructions into a guide you can follow

According to the pinned SKILL.md from K-Dense-AI/scientific-agent-skills: Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/scholar-evaluation"
Check the pinned source

Best fit

  • Provide developmental, evidence-traceable feedback on a scholarly work: paper, draft, protocol, literature synthesis, or research idea. Use qualitative judgment first. Optional scores only describe how submitted evidenc…
  • This skill also audits whether a low-stakes assessment process documents its construct, provenance, rater quality, uncertainty, traceability, sensitivity, fairness, accessibility, privacy, and human governance.
  • Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

Bring this context

  • A concrete task that matches the documented purpose of scholar-evaluation.
  • The files, examples, or context the task depends on.
  • Your constraints, target environment, and definition of done.

Expected outputs

  • A result that follows the pinned scholar-evaluation instructions.
  • A concise record of assumptions, inputs used, and unresolved questions.
  • A final check against the source workflow and relevant permission signals.

Key source sections

Read scholar-evaluation through these 5 source sections

Sections are extracted automatically from the pinned SKILL.md and link back to the source.

01

Workflow

Stop on a prohibited decision context or unnecessary private data.

SKILL.md · Workflow
developmental purpose;unit of assessment: scholarlywork;work type, stage, discipline, language, and audience;
02

8. Human review and release

Before releasing an organizational report, a qualified accountable human committee must verify:

SKILL.md · 8. Human review and release
construct and rubric provenance;content-validity evidence and limits;rater training, agreement, inter-rater reliability evidence, and drift;
03

Purpose

Provide developmental, evidence-traceable feedback on a scholarly work: paper, draft, protocol, literature synthesis, or research idea. Use qualitative judgment first. Optional scores only describe how submitted evidence maps to a predeclared bounded rubric.

SKILL.md · Purpose
Provide developmental, evidence-traceable feedback on a scholarly work: paper, draft, protocol, literature synthesis, or research idea. Use qualitative judgment first. Optional scores only describe how submitted evidenc…This skill also audits whether a low-stakes assessment process documents its construct, provenance, rater quality, uncertainty, traceability, sensitivity, fairness, accessibility, privacy, and human governance.
04

Hard safety boundary

Never use this skill to automate, recommend, materially influence, or score:

SKILL.md · Hard safety boundary
hiring, promotion, or tenure;admissions;grants or other funding;
05

ScholarEval status

The referenced ScholarEval project is an experimental literature-grounded research-idea evaluation framework, not validated psychometrics.

SKILL.md · ScholarEval status
The referenced ScholarEval project is an experimental literature-grounded research-idea evaluation framework, not validated psychometrics.The verified primary record is Moussa et al., ScholarEval: Research Idea Evaluation Grounded in Literature, arXiv:2510.16234v2, revised 2026-02-28. It reports a retrieval-augmented soundness/contribution framework, a 11…Do not generalize those results to person assessment, consequential decisions, all disciplines, or this skill's rubric. No peer-reviewed publication status was verified during the dated review. See references/sourceledg…

SkillSignal prompt templates

Provide the task, context, and acceptance criteria

These prompts were written by SkillSignal from the source structure; they are not upstream text.

Task-start prompt

Confirm source fit, inputs, and outputs before acting.

Use scholar-evaluation to help me with: [specific task]. Context: [files, data, or background]. Constraints: [environment, scope, and prohibited actions]. Before acting, check the pinned SKILL.md and explain which sections apply, what inputs are still missing, and what you will deliver.

Source-guided execution

Make the Agent explicitly follow the key extracted sections.

Apply the pinned scholar-evaluation source to [task]. Pay particular attention to these source sections: “Workflow”, “8. Human review and release”, “Purpose”, “Hard safety boundary”, “ScholarEval status”. Preserve the important decision at each step. Mark facts not covered by the source as “needs confirmation” instead of inventing them. Then verify the result against my acceptance criteria: [criteria].

Result-review prompt

Check omissions, permissions, and source drift before delivery.

Review the current scholar-evaluation result: (1) does it satisfy the original task; (2) were any applicable steps or limits in the pinned SKILL.md missed; (3) did it perform any unauthorized file, command, network, or data action; and (4) which conclusions remain unverified? List issues first, then fix only what the source or user authorization supports.

Output checklist

Verify each item before delivery

The task matches the purpose documented in the SKILL.md.

The source section “Workflow” has been checked.

The source section “8. Human review and release” has been checked.

The source section “Purpose” has been checked.

The source section “Hard safety boundary” has been checked.

Inputs, constraints, and acceptance criteria are explicit.

Unverified facts, compatibility, and outcome claims are clearly marked.

Any file, command, network, or data action has been reviewed.

Choose a different workflow

When another Skill is the better fit

dask

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

A separate implementation from K-Dense-AI/scientific-agent-skills; compare its source, maintenance signals, and permission requirements.

Open source detail

neuropixels-analysis

Analyze Neuropixels extracellular recordings end-to-end with SpikeInterface. Covers loading SpikeGLX/Open Ephys/NWB data, preprocessing, drift/motion correction, Kilosort4 (and CPU) spike sorting, quality metrics, and unit curation (threshold-based, model-based UnitRefine, and AI-assisted visual review). Use when working with Neuropixels 1.0/2.0 recordings, spike sorting, or extracellular electrophysiology analysis.

A separate implementation from K-Dense-AI/scientific-agent-skills; compare its source, maintenance signals, and permission requirements.

Open source detail

scanpy

Standard single-cell RNA-seq analysis pipeline. Use for QC, normalization, dimensionality reduction (PCA/UMAP/t-SNE), clustering, differential expression, visualization, and converting R-friendly single-cell formats such as Seurat or SingleCellExperiment RDS files into h5ad for Scanpy. Best for exploratory scRNA-seq analysis with established workflows. For deep learning models use scvi-tools; for data format questions use anndata.

A separate implementation from K-Dense-AI/scientific-agent-skills; compare its source, maintenance signals, and permission requirements.

Open source detail

FAQ

What does scholar-evaluation do?

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

How do I start using scholar-evaluation?

The catalog detected this source-specific install command: npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill "skills/scholar-evaluation". Inspect the command and pinned source before running it.

Which Agent platforms does it declare?

No dedicated Agent platform is declared in the pinned source record.

Repository stars
31,966
Repository forks
3,175
Quality
88/100
Source repository last pushed

Quality breakdown

Based on traceable docs and repository signals; stars are not treated as quality.

88/100
Documentation30/30
Specificity16/25
Maintenance20/20
Trust signals22/25

Compare before choosing

Related Agent Skills and source variants

These links are selected from shared tasks, functions, stacks, platforms, and same-name variants. Compare the source owner, documentation, permissions, and maintenance signals.

dask by k-dense-ai

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

neuropixels-analysis by k-dense-ai

Analyze Neuropixels extracellular recordings end-to-end with SpikeInterface. Covers loading SpikeGLX/Open Ephys/NWB data, preprocessing, drift/motion correction, Kilosort4 (and CPU) spike sorting, quality metrics, and unit curation (threshold-based, model-based UnitRefine, and AI-assisted visual review). Use when working with Neuropixels 1.0/2.0 recordings, spike sorting, or extracellular electrophysiology analysis.

scanpy by k-dense-ai

Standard single-cell RNA-seq analysis pipeline. Use for QC, normalization, dimensionality reduction (PCA/UMAP/t-SNE), clustering, differential expression, visualization, and converting R-friendly single-cell formats such as Seurat or SingleCellExperiment RDS files into h5ad for Scanpy. Best for exploratory scRNA-seq analysis with established workflows. For deep learning models use scvi-tools; for data format questions use anndata.

lead-generation by MoizIbnYousaf

Generate enriched ICP-based lead lists with Exa Agent, including structured scoring and CSV output. Use when generating leads, building prospect lists, finding companies to sell to, outbound research, or ICP-based company discovery. Triggers on leads, lead gen, prospect list, find companies, ICP, outbound list. Distinct from lead-magnet (content asset that captures emails).

astropy by k-dense-ai

Core Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.

View original Skill.mdThis page is parsed directly from the repository SKILL.md without editorial rewriting. Collected: Jul 28, 2026 · about 6 min

Scholar Evaluation

Purpose

Provide developmental, evidence-traceable feedback on a scholarly work: paper, draft, protocol, literature synthesis, or research idea. Use qualitative judgment first. Optional scores only describe how submitted evidence maps to a predeclared bounded rubric.

This skill also audits whether a low-stakes assessment process documents its construct, provenance, rater quality, uncertainty, traceability, sensitivity, fairness, accessibility, privacy, and human governance.

Hard safety boundary

Never use this skill to automate, recommend, materially influence, or score:

  • hiring, promotion, or tenure;
  • admissions;
  • grants or other funding;
  • prizes, honors, or awards;
  • discipline, dismissal, or sanctions; or
  • any other high-impact personnel decision.

Never rank people. Never reduce a person to a composite score. Never infer ability, character, integrity, protected traits, future performance, or worth. A nominal human-in-the-loop does not remove this boundary.

If asked for a prohibited use, stop. Offer developmental comments on a scholarly work or a process-only audit that does not process applications, compare people, recommend an outcome, or advise a decision.

Do not issue publication-readiness, accept/reject, or “top-tier” judgments.

Read references/responsible_assessment.md before any organizational use.

ScholarEval status

The referenced ScholarEval project is an experimental literature-grounded research-idea evaluation framework, not validated psychometrics.

The verified primary record is Moussa et al., ScholarEval: Research Idea Evaluation Grounded in Literature, arXiv:2510.16234v2, revised 2026-02-28. It reports a retrieval-augmented soundness/contribution framework, a 117-idea four-discipline dataset, coverage experiments, and a user study.

Do not generalize those results to person assessment, consequential decisions, all disciplines, or this skill's rubric. No peer-reviewed publication status was verified during the dated review. See references/source_ledger.md.

Metric and prestige policy

Do not score or infer quality from:

  • Journal Impact Factor or other journal measures;
  • h-index, publication counts, or citation counts;
  • altmetrics or attention;
  • journal, conference, venue, institution, employer, or geographic prestige;
  • author affiliation, reputation, network, or career path.

The rubric validator rejects common proxy-measure criteria.

If a qualified reviewer mentions an indicator descriptively outside the scoring tools, record its exact purpose, source, coverage, field and time effects, uncertainty, missingness, biases, gaming risk, and why it does not directly measure quality. Never hide indicators inside an opaque composite.

Data boundary

Bundled scripts accept only strict local JSON/CSV containing pseudonymous IDs, bounded ratings, statuses, uncertainty, and local references.

Do not put raw private applications, CVs, letters, reviewer identities, contact details, protected attributes, or source-document text in inputs, outputs, logs, examples, or prompts. Keep source content in the authorized records system and use opaque local references.

Allowed classifications are:

  • synthetic
  • public_scholarly_work
  • deidentified_low_stakes

No script searches the web, loads environment files, reads credentials, calls a model, executes supplied text, deserializes executable objects, or launches a process.

Use Bash only to invoke the documented local python3 commands.

Workflow

1. Confirm allowed use and authorization

Record:

  • developmental purpose;
  • unit of assessment: scholarly_work;
  • work type, stage, discipline, language, and audience;
  • authorized source location and data classification;
  • accountable committee owner;
  • conflicts and recusals;
  • accessibility and accommodation process;
  • appeal or correction route; and
  • data purpose, access, retention, and deletion.

Stop on a prohibited decision context or unnecessary private data.

2. Define the construct before criteria

State:

  • what quality or support is being examined;
  • excluded constructs;
  • intended interpretation;
  • contexts where the interpretation does not travel;
  • evidence requirements; and
  • known limitations.

Start with values and disciplinary context, not available metrics.

3. Adapt and validate the rubric

Begin with assets/rubric_template.json, then obtain qualified disciplinary, assessment-methods, stakeholder, accessibility, privacy, and fairness review.

The template deliberately records content validity as not_established. Do not change that status without documented evidence for the exact intended use.

Validate structure:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/validate_rubric.py \
  --rubric assets/rubric_template.json

Read references/evaluation_framework.md for construct, anchor, validity, and rater guidance.

4. Build traceable evidence records

Reviewers may read an authorized work outside the scripts. Record only stable local locators and claim references in assets/evidence_manifest_template.json.

For every criterion, distinguish:

  • observed evidence from interpretation;
  • supporting from contrary evidence;
  • available from unavailable evidence;
  • missing from not_applicable; and
  • uncertainty from absence.

Failure to find prior work does not prove novelty.

5. Rate independently

Use assets/evaluation_template.json. Each criterion must be:

  • rated with an anchor score, bounded uncertainty, evidence IDs, and a local rationale reference;
  • missing with null score/uncertainty and a rationale reference; or
  • not_applicable with null score/uncertainty and a rationale reference.

Do not encode missing or not-applicable as zero. Raters should train, calibrate, disclose conflicts, rate independently, and document disagreement.

6. Run local quality checks

Bounded scoring, without labels or recommendation:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/calculate_scores.py \
  --rubric assets/rubric_template.json \
  --evaluation assets/evaluation_template.json

Evidence traceability:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_traceability.py \
  --rubric assets/rubric_template.json \
  --evaluation assets/evaluation_template.json \
  --evidence assets/evidence_manifest_template.json

Inter-rater agreement:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/summarize_agreement.py \
  --rubric assets/rubric_template.json \
  --ratings assets/ratings_template.csv

Weight sensitivity requires two or more distinct scholarly-work evaluation files:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/weight_sensitivity.py \
  --rubric assets/rubric_template.json \
  --evaluation /tmp/work-a-evaluation.json \
  --evaluation /tmp/work-b-evaluation.json

Process controls:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_process.py \
  --process assets/process_checklist_template.json

The checklist template is intentionally unconfirmed and fails closed. Instructions and exact schemas are in references/local_tooling.md.

7. Synthesize qualitative findings

Lead with criterion-level evidence, not the composite. For each criterion:

  1. cite evidence references;
  2. state rated, missing, or not_applicable;
  3. explain the anchor interpretation;
  4. report score and uncertainty only if rated;
  5. note disagreements and context;
  6. identify strengths and limitations; and
  7. offer non-prescriptive improvement options.

Generate an empty-reference scaffold if useful:

PYTHONDONTWRITEBYTECODE=1 python3 scripts/generate_report_scaffold.py \
  --rubric assets/rubric_template.json \
  --evaluation assets/evaluation_template.json \
  --output /tmp/developmental-report-scaffold.json

The scaffold does not read source documents or draft findings.

8. Human review and release

Before releasing an organizational report, a qualified accountable human committee must verify:

  • construct and rubric provenance;
  • content-validity evidence and limits;
  • rater training, agreement, inter-rater reliability evidence, and drift;
  • evidence traceability and source access;
  • missingness, not-applicable rationales, and uncertainty;
  • weight sensitivity and order instability;
  • disciplinary and subgroup bias review;
  • conflicts and recusals;
  • accessibility and accommodations;
  • privacy, minimization, retention, and output controls; and
  • correction or appeal information.

Document dissent. Do not imply consensus, validity, or precision beyond the evidence. Periodically evaluate the evaluation and retire harmful criteria.

Interpretation rules

  • A score is an ordinal rubric summary, not a natural measurement.
  • Normalization does not repair incomplete evidence.
  • The bundled uncertainty range is not a confidence interval.
  • Agreement does not establish reliability, validity, fairness, or correctness.
  • Stable results under tested weights do not establish validity.
  • The overall score never overrides criterion evidence or qualified judgment.
  • No output is a decision recommendation.

Bundled resources

  • references/responsible_assessment.md — safety, metrics, governance, accessibility, privacy, and bias.
  • references/evaluation_framework.md — ScholarEval boundary, construct, criteria, anchors, validity, and interpretation.
  • references/local_tooling.md — strict schemas, formulas, commands, and output behavior.
  • references/source_ledger.md — authoritative sources and publication-status verification dated 2026-07-23.
  • references/security_validation.md — baseline remediation, validation, and residual security-scan record.
  • assets/rubric_template.json — bounded rubric template.
  • assets/evaluation_template.json — rating template.
  • assets/evidence_manifest_template.json — traceability template.
  • assets/process_checklist_template.json — fail-closed process checklist.
  • assets/ratings_template.csv — synthetic agreement data.
Skill path
skills/scholar-evaluation/SKILL.md
Commit SHA
e7ac42510774
Repository license
MIT
Data collected