Best for
- Use when the user wants a last30days-style pre-research pass in Deepline: discover the critical public, private, CRM, workflow, social, and web data sources for a research/enrichment job; compare provider coverage; esti…
getaero-io/gtm-eng-skills/skills/deepline-pre-research/SKILL.md
Use when the user wants a last30days-style pre-research pass in Deepline: discover the critical public, private, CRM, workflow, social, and web data sources for a research/enrichment job; compare provider coverage; estimate Deepline credit cost; recommend the source plan before building or running the workflow; or build custom language/messaging from buyer, competitor, community, and CRM evidence. Triggers: pre-research, source discovery, provider strategy, research data sources, ScrapeCreators,
Decision brief
Triggers: pre-research, source discovery, provider strategy, research data sources, ScrapeCreators,
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/getaero-io/gtm-eng-skills --skill "skills/deepline-pre-research"Inspect the Agent Skill "deepline-pre-research" from https://github.com/getaero-io/gtm-eng-skills/blob/3075ea09fbcaca9cf114c0d02e4fc7c62181d093/skills/deepline-pre-research/SKILL.md at commit 3075ea09fbcaca9cf114c0d02e4fc7c62181d093. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
bash npm install -g deepline
Review the “Recommended Workflow” section in the pinned source before continuing.
deepline auth register --wait auto deepline auth wait --timeout 120 completes Cowork/browser approval; no-op if already connected deepline auth status deepline -h bash git ls-remote --tags https://github.com/mvanhorn/last30days-skill.git | tail -20 bash curl -s "$DEEPLINEAPIBASE…
Run deepline when it is available. If the shell reports that command is missing, use /.deepline/runtime/bin/deepline (or the npm-created .cmd shim on Windows). If neither exists, follow https://code.deepline.com/SKILL.md to set up Deepline.
Portions of the query-design, public-source fanout, and consolidation approach are adapted from mvanhorn/last30days-skill, MIT licensed, copyright (c) 2026 Matt Van Horn. Keep THIRDPARTYNOTICES.md with this skill when packaging or distributing it.
Permission review
The documentation asks the agent to run terminal commands or scripts.
npm install -g deeplineThe documentation includes network, browsing, or remote request actions.
# Fallback for secure sandboxes: mkdir -p "$HOME/.local" && npm config set prefix "$HOME/.local" && export PATH="$HOME/.local/bin:$PATH" && npm install -g deepline --registry https://code.deepline.com/api/v2/npm/The documentation includes network, browsing, or remote request actions.
**Resolve every materializable dataset to its exact artifact, not its family.** For each public dataset/registry you recommend, the fanout must return and record: (1) the **exact file/endpoint name** (e.g. `IA_FIRM_SEC_Feed_YYYY_MM_DD.xml.gThe documentation asks the agent to read local files, directories, or repositories.
**Resolve every materializable dataset to its exact artifact, not its family.** For each public dataset/registry you recommend, the fanout must return and record: (1) the **exact file/endpoint name** (e.g. `IA_FIRM_SEC_Feed_YYYY_MM_DD.xml.gThe documentation asks the agent to run terminal commands or scripts.
git ls-remote --tags https://github.com/mvanhorn/last30days-skill.git | tail -20The documentation asks the agent to read local files, directories, or repositories.
**exact artifact**: the specific file/endpoint name, canonical URL, mirror, and any open-source parser/repo (not just "the X bulk feed" or "search Google") — see Non-Negotiables and the Artifact Resolution Gate (§4.55)The documentation asks the agent to create, modify, or delete local files.
For standalone eval harness runs, write reports into the harness output directory:Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 95/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 48 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
npm install -g deepline
# Fallback for secure sandboxes: mkdir -p "$HOME/.local" && npm config set prefix "$HOME/.local" && export PATH="$HOME/.local/bin:$PATH" && npm install -g deepline --registry https://code.deepline.com/api/v2/npm/
deepline auth register --wait auto
deepline auth wait --timeout 120 # completes Cowork/browser approval; no-op if already connected
deepline auth status
deepline -h
Run deepline when it is available. If the shell reports that command is missing, use <workspace-root>/.deepline/runtime/bin/deepline (or the npm-created .cmd shim on Windows). If neither exists, follow https://code.deepline.com/SKILL.md to set up Deepline.
Find the highest-signal GTM data sources, public evidence, and market language for a research or enrichment job before building the pipeline. This is a standalone Deepline skill that should behave like last30days with a GTM data lens: broad source coverage, recency, community signals, citations, source stats, and a grounded "What I learned" synthesis. In Deepline, the report first explains what the research found; only after that does it translate the findings into Deepline tool contracts, private/proprietary joins, and Deepline-facing cost.
Portions of the query-design, public-source fanout, and consolidation approach are adapted from mvanhorn/last30days-skill, MIT licensed, copyright (c) 2026 Matt Van Horn. Keep THIRD_PARTY_NOTICES.md with this skill when packaging or distributing it.
serper/exa, or the equivalent web-search tool) during the public-source fanout. If a run names public datasets without having searched for them this session, it has failed the fanout — no exceptions for "obvious" verticals. Naming a source family from memory is a draft, not a finding; the finding is the exact artifact the search returns. (Eval evidence: runs that skipped web search lost or tied on exactly the prompts where a competitor searched and surfaced concrete artifacts.)IA_FIRM_SEC_Feed_YYYY_MM_DD.xml.gz, not "the ADV bulk feed"); (2) the canonical download/API URL; (3) any mirror (e.g. data.gov catalog copy) that is easier to pull; (4) an existing open-source parser or GitHub repo that already structures it, when one exists (search "<dataset> parser github"); (5) the government statistical registry for the vertical when one exists (BLS QCEW + NAICS codes, Census County Business Patterns, etc.) for free establishment counts and sizing. "Source family named" is not done. "Exact file + URL + mirror + parser + NAICS code recorded" is done. See the Artifact Resolution Gate (§4.55)./last30days at runtime. Reference it only as a design benchmark for source breadth and synthesis discipline.npi tool exists. Classify this as available through generic route, not as an unusable gap.native, available through generic route, private connector, or missing provider to add.references/source-map.md.references/last30days-gtm-corpus.md; it summarizes the relevant saved last30days runs and the source patterns they proved useful.last30days parity, also read references/query-design.md and references/fanout-consolidation.md.../deepline-gtm/SKILL.md and the sub-doc it routes to.../deepline-analytics/SKILL.md.last30days release when using it as a design reference:git ls-remote --tags https://github.com/mvanhorn/last30days-skill.git | tail -20
Do not copy last30days local scripts into Deepline and do not call it as a required sub-skill. Use it for source taxonomy and output discipline, then route execution through Deepline tools, plays, workflows, and private-data connectors.
This skill owns pre-research and source planning. It may hand off after the source plan is clear:
deepline-gtm: execute GTM sourcing, enrichment, waterfalls, personalization, contact discovery, or CRM activation.deepline-analytics: query warehouse/semantic-layer metrics and RevOps/customer datasets.deepline-plays: build, wrap, fork, run, inspect, or automate the final repeatable workflow/play.last30days: design reference only, never a dependency for Deepline execution.Extract:
OBJECTIVE: what decision or output the research should supportENTITY_SCOPE: companies, people, markets, accounts, customers, workflows, competitors, or topicsTIME_WINDOW: default to last 30 days for trend/community research; preserve user-provided windowsPRIVATE_SOURCES: CRM, warehouse, product events, workflow runs, support/calls, sheets, user CSVsPUBLIC_SOURCES: web, social, jobs, ads, technographics, funding, news, communities, app stores, directoriesDATASET_LEADS: public registries, open datasets, papers, GitHub repos, government records, niche directories, or platform posts that point to materializable dataCUSTOM_LANGUAGE_OUTPUTS: outbound snippets, ad hooks, landing-page copy, objection language, category positioning, call scripts, lead magnets, or CRM personalization fieldsOUTPUT: source plan, workflow design, CSV schema, play spec, or final research briefTell the user the parsed scope before provider calls.
Do not default to browser scraping, raw x.com scraping, or unvetted actors. Prefer managed external APIs and private connectors:
| Source family | Preferred provider | Credential needed | Notes |
|---|---|---|---|
| Reddit threads/comments | scrapecreators | Deepline integration API key | Required for managed Reddit comments/thread coverage. |
| TikTok / Instagram / YouTube fallback | scrapecreators | Deepline integration API key | Preferred managed API for captions, engagement, and transcript fallback. Search ScrapeCreators unfiltered when profile/contact tools matter; some profile tools are categorized as admin, not research. |
| X/Twitter | twitterapi | Deepline integration API key | Managed X/Twitter search; avoids scraping x.com. |
| Web/news/docs | serper | Deepline integration API key or SERPER_API_KEY for server-owned dev/prod use | Search API, not search-page scraping. |
| Web extraction fallback | firecrawl or exa | Deepline integration API key | Use after URL discovery; do not use for blind scrape-at-scale. |
| Hacker News | hackernews | none | Public Algolia API. |
| Bluesky | bluesky | none | Public AppView API. |
| CRM/private | Salesforce, HubSpot, Attio | Deepline OAuth/private connector | Query customer-authorized APIs, not UI scraping. |
| Warehouse/product/workflow | Snowflake/customer DB/workflow runtime | Deepline private connector | Use scoped warehouse/runtime access. |
| Last-resort social actor fallback | apify | Deepline integration API key | Optional fallback only after explicit approval and provider gap status. |
Treat any route not covered by a described Deepline tool contract, key/auth plan, test endpoint, and Deepline-facing pricing as unapproved until after the public research synthesis is complete.
Before provider routing, run a last30days-style public discovery pass. The goal is to produce the research synthesis first: what public evidence, communities, registries, directories, reviews, posts, docs, and datasets actually teach us about the GTM problem.
Search across:
This pass is search-driven, not recall-driven. Issue real queries. For any vertical dataset, run at minimum these query shapes and read the top results before writing the source line:
"<vertical> <entity> registry OR bulk data OR dataset download" — find the authoritative file"<dataset name> download" / site:data.gov <vertical> — find the canonical URL and any mirror"<dataset name> parser github" / "<dataset name> python" — find an existing parser/repo so you don't reinvent extraction"<vertical> NAICS code" + BLS QCEW <NAICS> / Census County Business Patterns — free establishment counts and sizingFor each useful source lead, record:
Do not stop at "Serper can search this." Search tools are routes; the deliverable is the public source or dataset discovered through them. If you find yourself writing a dataset name you did not just see in a search result, stop and search for it.
When running inside Deepline or the V2 SDK, use the native planning endpoint only after the public-source fanout has produced a first synthesis and source map. The API is for translating the research into Deepline routes, gaps, approval gates, and cost, not for replacing the research pass:
curl -s "$DEEPLINE_API_BASE_URL/api/v2/pre-research/plan" \
-H "Authorization: Bearer $DEEPLINE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"objective":"best GTM data sources for SMB consumer services companies","depth":"deep"}'
Agent validation endpoint:
curl -s "$DEEPLINE_API_BASE_URL/api/v2/pre-research/test" \
-H "Authorization: Bearer $DEEPLINE_API_KEY"
The endpoint returns the query plan, source coverage, current Deepline tool candidates, sanitized Deepline-facing pricing metadata, explicit provider gaps, and approval gate. It does not execute paid provider calls.
For SDK V2 users, ask the installed Deepline GTM skill to build the source plan:
/deepline-gtm Build a pre-research source plan for my GTM problem before running paid provider calls.
The skill prompts the agent to call /api/v2/pre-research/plan, inspect the returned providerRequirements, and only then decide whether to run probes or build a play.
Run several focused searches, usually in parallel. deepline tools search accepts an optional intent query, but requires either that query or one of --categories / --search_terms; those filters accept comma-separated values. Use --json for machine-readable output. There is no --prefix flag, so put a provider name in the query instead.
deepline tools search "web search news source discovery" --categories research --search_terms "web search,news,recency,source discovery"
deepline tools search "social posts reddit x twitter youtube tiktok instagram" --categories research --search_terms "social posts,reddit,x twitter,youtube,tiktok,instagram"
deepline tools search scrapecreators --json
deepline tools search "facebook profile email scrapecreators" --json
deepline tools search "instagram profile bio links scrapecreators" --json
deepline tools search "company dataset firmographics funding technographics jobs" --categories company_search --search_terms "company dataset,firmographics,funding,technographics,jobs"
deepline tools search "crm warehouse workflow session usage" --categories admin --search_terms "crm,warehouse,workflow,session,usage"
For CRM/private data, also search by provider name when relevant:
deepline tools search salesforce
deepline tools search hubspot
deepline tools search attio
deepline tools search snowflake
Do not send the raw user prompt to every provider. Generate a source-specific query plan first:
for dir in \
"$PWD/.skills/deepline-pre-research" \
"$HOME/.claude/skills/deepline-pre-research" \
"$HOME/.agents/skills/deepline-pre-research"; do
[ -f "$dir/scripts/query_design.py" ] && SKILL_ROOT="$dir" && break
done
[ -n "${SKILL_ROOT:-}" ] || { echo "Could not find deepline-pre-research skill root" >&2; exit 1; }
python3 "$SKILL_ROOT/scripts/query_design.py" "$OBJECTIVE" --depth default
Use the plan to decide:
For production Deepline implementation, port this helper to the runtime language or call equivalent logic before deepline tools execute/deepline enrich.
Before recommending a plan, check every required source family in references/source-map.md:
For each family, mark one of:
native: a Deepline tool/play exists and can be describedgeneric route: use apify, serper, exa, parallel, firecrawl, or deeplineagentprivate connector: use CRM, warehouse, workflow, or customer dataset accessgap: provider/action should be added before Deepline can replicate that part wellDo not skip a family just because it is inconvenient. Missing coverage is part of the pre-research output.
Before writing the report, verify every materializable public dataset has been resolved from a family to an artifact. For each recommended public dataset/registry, confirm you can fill this row from something you actually searched or fetched this session:
| Dataset | Exact file/endpoint | Canonical URL | Easier mirror | OSS parser/repo | Vertical stat registry (NAICS/QCEW/CBP) |
|---|
Rules:
n/a is allowed only when the artifact genuinely does not exist (e.g. no public bulk file, no known parser) — and you must have searched to establish that, not assumed it.A report that names "the SEC ADV bulk feed" fails this gate; a report that names IA_FIRM_SEC_Feed_YYYY_MM_DD.xml.gz at its sec.gov URL, notes the data.gov mirror, and links an existing ADV parser repo passes it.
When you cite a parser/repo alongside a dataset, name the exact input file that repo actually consumes — check its README. A vertical often has multiple real distributions (e.g. SEC publishes both the IA_FIRM_SEC_Feed_*.xml.gz bulk feed and the ia<MMDDYY>.zip IAPD compilation); pairing a parser with the wrong-but-real file is a precision error, not a fabrication, but it still breaks the handoff. Match the parser to its file.
For last30days parity, preserve the public-source fanout mechanics described in references/fanout-consolidation.md:
If implementation coverage is being audited, run the local evaluator:
for dir in \
"$PWD/.skills/deepline-pre-research" \
"$HOME/.claude/skills/deepline-pre-research" \
"$HOME/.agents/skills/deepline-pre-research"; do
[ -f "$dir/scripts/evaluate_examples.py" ] && SKILL_ROOT="$dir" && break
done
[ -n "${SKILL_ROOT:-}" ] || { echo "Could not find deepline-pre-research skill root" >&2; exit 1; }
python3 "$SKILL_ROOT/scripts/evaluate_examples.py"
Use evals/side-by-side.md to compare saved last30days examples against this skill's coverage contract.
For standalone eval harness runs, write reports into the harness output directory:
python3 "$SKILL_ROOT/scripts/evaluate_examples.py" \
--out-md "${OUTPUT_DIR:-$SKILL_ROOT/evals}/side_by_side_pre_research.md" \
--out-json "${OUTPUT_DIR:-$SKILL_ROOT/evals}/side_by_side_pre_research.json"
python3 "$SKILL_ROOT/scripts/evaluate_public_private_corpus.py" \
--out-md "${OUTPUT_DIR:-$SKILL_ROOT/evals}/public_private_corpus_results.md" \
--out-json "${OUTPUT_DIR:-$SKILL_ROOT/evals}/public_private_corpus_results.json"
For every candidate, inspect the live contract:
deepline tools describe <tool-id> --json
Record:
If a source family is not present in Deepline, mark it as a provider gap and propose the integration path.
Use free/low-cost probes first. For paid probes, ask before running unless the user already approved a specific pilot budget.
Good probe shapes:
limit: 1 or numResults: 3At scale, use deepline enrich rather than ad hoc loops so results are inspectable in Playground.
Use the same output discipline as last30days agent mode. The report should feel like a real research result, not a bespoke segment-analysis template or provider catalog.
## Research Report: <topic>
Generated: <date> | Sources: Reddit, X, YouTube, TikTok, Instagram, HN, Polymarket, Web, Public registries, Deepline catalog
### Key Findings
- <3-5 highest-signal findings grounded in public/source evidence>
### What I learned
**<Theme 1>** - <1-2 sentences about the segment, source, buyer language, or dataset reality. Cite sparingly using source names, handles, communities, registries, or domains.>
**<Theme 2>** - <1-2 sentences.>
**<Theme 3>** - <1-2 sentences.>
KEY PATTERNS from the research:
1. <Pattern, source family, and why it matters for GTM>
2. <Pattern, source family, and why it matters for GTM>
3. <Pattern, source family, and why it matters for GTM>
### GTM Data Sources Found
| Source family | Best source / route | Status | Rows / join keys | Caveat |
| ------------- | ------------------- | ------ | ---------------- | ------ |
### Materializable Datasets (exact artifacts)
| Dataset | Exact file/endpoint | URL | Mirror | OSS parser/repo | Stat registry (NAICS/QCEW/CBP) |
| ------- | ------------------- | --- | ------ | --------------- | ------------------------------ |
<!-- One row per public dataset you can turn into rows. No "family" placeholders — exact filename, canonical URL, mirror, and an existing parser repo where one exists. This table is the answer when the ask is "give me the dataset/github". -->
### Market Language
| Source | What to extract | How to use it | Guardrail |
| ------ | --------------- | ------------- | --------- |
### Proprietary Data To Join Later
| Dataset | Owner / connector | Join key | Why it changes the answer |
| ------- | ----------------- | -------- | ------------------------- |
### Deepline Route
| Public source family | Route status | Deepline route | Probe needed | Cost basis |
| -------------------- | ------------ | -------------- | ------------ | ---------- |
### Gaps To Add
| Missing source | Why it matters | Suggested provider/integration | Required fields |
| -------------- | -------------- | ------------------------------ | --------------- |
### Recommended Workflow
1. <first retrieval/probe>
2. <second retrieval/probe>
3. <synthesis/scoring/join step>
### Stats
---
✅ All agents reported back!
├─ <source family>: <count / coverage / engagement when available>
├─ <source family>: <count / coverage / engagement when available>
├─ 🌐 Web/Public data: <pages, registries, directories, datasets>
└─ 🗣️ Top sources: <handles, communities, registries, domains, datasets>
---
### Cost Estimate
- Pilot: <credits or range>
- Full run: <credits or range and assumptions>
- Unknowns: <what must be described/probed before quoting>
For --agent-style or non-interactive use, stop after the report. For interactive use, end with a short, specific invitation based on the actual findings, mirroring last30days:
---
I'm now an expert on <topic>. Some things I can help with:
- <specific follow-up grounded in the findings>
- <specific workflow/probe to run next>
- <specific segment or source to go deeper on>
If the user wants execution, hand off to deepline-gtm after this output.
Match last30days cleanup rules:
NPI Registry, Google Maps, r/SaaS, G2, Meta docs, or Hike Medical site.last30days stats block style exactly: start with ✅ All agents reported back!, use ├─, └─, and │ separators where useful, and include source emojis when they clarify the source family.Before any full paid run, include:
If the user approves, execute through deepline enrich, deepline tools execute, or a Deepline play/workflow as appropriate.
The task is done only when the user has:
last30days-style GTM research synthesis grounded in public sourcesAlternatives
alirezarezvani/claude-skills
Analyzes RFP/RFI responses for coverage gaps, builds competitive feature comparison matrices, and plans proof-of-concept (POC) engagements for pre-sales engineering. Use when responding to RFPs, bids, or proposal requests; comparing product features against competitors; planning or scoring a customer POC or sales demo; preparing a technical proposal; or performing win/loss competitor analysis. Handles tasks described as 'RFP response', 'bid response', 'proposal response', 'competitor comparison'
equinor/neqsim
Subsea production systems, DNV-RP-F109 on-bottom stability screening, DNV-RP-F105 free-span screening, DNV-RP-F101 corroded-pipeline screening, well design, SURF cost estimation, and tieback analysis with NeqSim. USE WHEN: designing subsea fields, screening pipeline/cable/umbilical seabed stability or inspected metal loss, sizing flowlines and umbilicals, estimating well costs, performing casing design, running tieback comparisons, or configuring subsea equipment (trees, manifolds, boosters, ris
mgiovani/cc-arsenal
Multi-agent review team: architecture, security, performance, testing, style, docs/UX, plus an adversary that cross-examines the other 6, for security-sensitive, architectural, or large PRs (15+ files) where a single-agent pass risks missing cross-cutting issues. Use for auth/payments/PII changes, schema/pattern changes, compliance sign-off, or when asked to 'get the review team on this' / 'multi-agent review' / 'thorough review before merge'. For a standard PR or a quick pre-merge check, use /r
alirezarezvani/claude-skills
Discover, find, compare, audit, repair, adapt, and design repeatable AI-agent loops with explicit triggers, actions, verification, stopping conditions, guardrails, and handoffs. Use when a user asks to analyze a codebase for potential loops, mine coding-thread history for work done more than once, turn repeated engineering work into a loop, find or recommend a published loop, create a recurring agent workflow or automation cadence, turn an outcome into a bounded copy-ready loop, or review an exi