jimezsa/opencolab/projects/SKILLS/pdf-figure-extract/SKILL.md
pdf-figure-extract
Extract and return figures from already-downloaded local PDFs with PyMuPDF, optionally reusing PageIndex artifacts for page selection and verifying shortlisted candidates multimodally before delivery.
- Source repository stars
- 11
- Declared platforms
- 0
- Static risk flags
- 1
- Last source update
- 2026-08-04
- Source checked
- 2026-08-04
Decision brief
What it does—and where it fits
Use this skill when the user wants a figure, architecture image, pipeline diagram, qualitative result panel, or other visual artifact extracted from a paper that already exists locally under the current project.
Not for
- Tasks that require unconfirmed production actions or broad system permissions.
- Environments where the pinned source and install steps cannot be inspected.
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/jimezsa/opencolab --skill "projects/SKILLS/pdf-figure-extract"Inspect the Agent Skill "pdf-figure-extract" from https://github.com/jimezsa/opencolab/blob/f647b8e4c37a18b4bd3443bd4a8f5470ea1b9d09/projects/SKILLS/pdf-figure-extract/SKILL.md at commit f647b8e4c37a18b4bd3443bd4a8f5470ea1b9d09. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
Workflow
Use the request plus whatever local artifacts already exist:
research/INDEX.md/RUN.md/meta/.json - 02
Mission
Given a request for a figure in an already-downloaded local PDF:
Select a bounded local paper set.Reuse PageIndex artifacts when available to narrow likely pages, otherwise shortlist pages directly from the PDF.Extract or render 1-3 figure candidates with PyMuPDF. - 03
Prerequisites
Optional but recommended:
Local PDFs already exist under an active research run folder, normally research/-/pdf/, or under the legacy flat research/pdf/ layout.python3 is installed and available in PATH.PyMuPDF is installed: - 04
Hard Requirements
Operate only on already-downloaded local PDFs. Do not use this skill for paper discovery.
Operate only on already-downloaded local PDFs. Do not use this skill for paper discovery.Prefer selecting the active research run folder from research/INDEX.md when it exists. If there is no index, infer the best run folder from the user's topic and existing research//RUN.md files; fall back to legacy resea…Use projects/SKILLS/pdf-figure-extract/scripts/pdffigureextract.py as the canonical local extractor. - 05
OpenColab Progress Helper
OpenColab exposes this progress channel by default during provider runs. When OPENCOLABPROGRESSFILE is available, use this helper:
selected paper set knownPageIndex artifacts found or missingcandidate pages shortlisted
Permission review
Static risk signals and limitations
Runs scripts
The documentation asks the agent to run terminal commands or scripts.
python3 -m pip install PyMuPDFRuns scripts
The documentation asks the agent to run terminal commands or scripts.
python3 projects/SKILLS/pdf-figure-extract/scripts/pdf_figure_extract.py \Evidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 91/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 11 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- jimezsa/opencolab
- Skill path
- projects/SKILLS/pdf-figure-extract/SKILL.md
- Commit
- f647b8e4c37a18b4bd3443bd4a8f5470ea1b9d09
- License
- MIT
- Collected
- 2026-08-04
- Default branch
- main
View the original SKILL.md
PDF Figure Extract Skill
Use this skill when the user wants a figure, architecture image, pipeline diagram, qualitative result panel, or other visual artifact extracted from a paper that already exists locally under the current project.
Typical use cases:
- "send me the architecture figure from this paper"
- "extract Figure 3 from the local PDF"
- "return the pipeline overview image"
- "find the model diagram from the paper and send it back"
- "pull out the qualitative comparison panel on page 7"
This skill complements pageindex-grounded, not replaces it.
When cached PageIndex artifacts already exist, reuse them to narrow candidate pages before extraction.
When PageIndex is missing or inactive, continue in standalone PyMuPDF mode instead of blocking the run.
Mission
Given a request for a figure in an already-downloaded local PDF:
- Select a bounded local paper set.
- Reuse PageIndex artifacts when available to narrow likely pages, otherwise shortlist pages directly from the PDF.
- Extract or render 1-3 figure candidates with PyMuPDF.
- Inspect those candidates with the agent's multimodal capability when available to confirm which image best matches the user's request.
- Save the exported artifacts and manifest data under the active research run folder, normally
<RUN_ROOT>/figures/. - Return the best figure with paper and page context, plus explicit limitations when confidence is reduced.
Prerequisites
- Local PDFs already exist under an active research run folder, normally
research/<YYYY-MM-DD>-<topic-slug>/pdf/, or under the legacy flatresearch/pdf/layout. python3is installed and available inPATH.- PyMuPDF is installed:
python3 -m pip install PyMuPDF
Optional but recommended:
<RUN_ROOT>/pageindex/manifest.jsonand cached PageIndex trees already exist.- paper summaries exist under
<RUN_ROOT>/pdf/*.md.
If PyMuPDF is missing, only install it when the user explicitly asks for installation or setup work.
Hard Requirements
- Operate only on already-downloaded local PDFs. Do not use this skill for paper discovery.
- Prefer selecting the active research run folder from
research/INDEX.mdwhen it exists. If there is no index, infer the best run folder from the user's topic and existingresearch/*/RUN.mdfiles; fall back to legacyresearch/pdf/only for older projects. - Use
projects/SKILLS/pdf-figure-extract/scripts/pdf_figure_extract.pyas the canonical local extractor. - Treat PageIndex as optional acceleration:
- reuse it when available and relevant
- do not fail just because PageIndex artifacts are missing
- Keep selection bounded:
- normally 1 paper for a single-paper request
- normally 2-3 papers for a cross-paper figure request unless the user explicitly asks for broader coverage
- normally 1-8 candidate pages per selected paper before exporting images
- Persist artifacts under
<RUN_ROOT>/figures/, not under temporary ad hoc folders. - Export a user-deliverable PNG even when the figure is vector-heavy or mixed-content.
- Before returning the figure, inspect the shortlisted candidate images directly with the active agent's multimodal capability when the provider runtime supports local image inspection.
- If multimodal inspection is unavailable, say so explicitly and fall back to caption, page, and layout heuristics instead of overstating certainty.
- If confidence is low, prefer returning the best candidate or top candidates with limitations rather than pretending the match is exact.
- If the chosen figure should be sent back through Telegram, emit a raw
@telegram-file {"kind":"photo","file":"<path>","caption":"optional"}line on its own line with no backticks or code fences. Keep the JSON on one line, keepkindasphoto(neverimage/png/jpg), and on Windows write the path with forward slashes. - The final reply must identify the paper and page, and include nearby caption text or figure number when available.
OpenColab Progress Helper
OpenColab exposes this progress channel by default during provider runs. When OPENCOLAB_PROGRESS_FILE is available, use this helper:
emit_progress() {
if [ -z "${OPENCOLAB_PROGRESS_FILE:-}" ]; then
return 0
fi
printf '%s\n' "$1" >> "$OPENCOLAB_PROGRESS_FILE"
}
Useful update categories for this skill:
- selected paper set known
- PageIndex artifacts found or missing
- candidate pages shortlisted
- figure candidates exported
- multimodal verification started or skipped
- degraded standalone fallback
- final figure delivered
Workflow
1. Select the paper set
Use the request plus whatever local artifacts already exist:
research/INDEX.md<RUN_ROOT>/RUN.md<RUN_ROOT>/meta/*.json<RUN_ROOT>/pdf/*.md<RUN_ROOT>/pageindex/manifest.json- prior
<RUN_ROOT>/findings.md
Selection guidance:
- exact single-paper request: 1 paper
- "compare the architecture figures in these two papers": 2 papers
- broader but still bounded figure request: 2-3 papers
2. Prepare the figure workspace
RUN_ROOT="research/<YYYY-MM-DD>-<topic-slug>"
mkdir -p "$RUN_ROOT/figures"/{exports,manifests,notes}
3. Run the extractor
Standalone or auto mode:
python3 projects/SKILLS/pdf-figure-extract/scripts/pdf_figure_extract.py \
--pdf-path "$RUN_ROOT/pdf/<safe_id>.pdf" \
--query "architecture figure" \
--output-root "$RUN_ROOT/figures" \
--top-k 3
PageIndex-assisted mode when the cached tree is already known:
python3 projects/SKILLS/pdf-figure-extract/scripts/pdf_figure_extract.py \
--pdf-path "$RUN_ROOT/pdf/<safe_id>.pdf" \
--query "architecture figure" \
--pageindex-tree "$RUN_ROOT/pageindex/trees/<safe_id>.json" \
--output-root "$RUN_ROOT/figures" \
--top-k 3
Direct figure or page hint mode:
python3 projects/SKILLS/pdf-figure-extract/scripts/pdf_figure_extract.py \
--pdf-path "$RUN_ROOT/pdf/<safe_id>.pdf" \
--query "Figure 3" \
--figure-number 3 \
--page-hint 5 \
--output-root "$RUN_ROOT/figures" \
--top-k 2
The script writes a per-run manifest under $RUN_ROOT/figures/manifests/ and updates $RUN_ROOT/figures/manifest.json with the latest run summary.
4. Verify the candidates multimodally
After extraction:
- Read the manifest and note the top 1-3 candidate image paths.
- Inspect those local images directly with the active provider's multimodal capability.
- Check whether the image actually matches the request:
- architecture or pipeline overview,
- the requested figure number,
- the nearby caption or page context,
- the expected visual content.
- Choose the best candidate only after that inspection step.
If the provider runtime cannot inspect local images, say so explicitly and rely on the manifest, page context, caption text, and extraction score as a degraded fallback.
5. Write an optional note
For non-trivial figure retrieval, write:
$RUN_ROOT/figures/notes/<date>-<topic-slug>.md
Recommended structure:
# Figure Extraction Note: <topic>
## Request
...
## Selected Paper
...
## Extraction Mode
- `pageindex-assisted` or `standalone`
## Chosen Figure
- file: `$RUN_ROOT/figures/exports/...`
- page: ...
- caption: ...
## Limitations
...
6. Return the result
The user-facing reply should:
- answer directly
- identify the selected paper and page
- say whether the result came from
pageindex-assistedorstandalone - mention when the returned artifact is a clipped page render rather than a direct embedded-image extraction if that matters
- surface low-confidence matching, missing multimodal verification, missing PageIndex artifacts, or other limitations when they affect confidence
- point to the saved note when one was written
If returning the figure through Telegram, emit the raw @telegram-file line after the short reply.
Output Contract
<RUN_ROOT>/figures/manifest.jsonfor the latest run summary<RUN_ROOT>/figures/manifests/<slug>.jsonfor the per-run manifest<RUN_ROOT>/figures/exports/<slug>__p<page>__cand<rank>.pngfor shortlisted figure candidates- optional raw extracted image files when the source figure was directly embedded
- optional
<RUN_ROOT>/figures/notes/<date>-<topic-slug>.md - a concise final reply with paper, page, and limitations
Canonical Assets
- Skill doc:
projects/SKILLS/pdf-figure-extract/SKILL.md - Python extractor:
projects/SKILLS/pdf-figure-extract/scripts/pdf_figure_extract.py