Best for
- Use when: (1) calling any HeyGen API endpoint (api.
OpenCoven/coven/skills/heygen-skills/SKILL.md
Create HeyGen avatar videos via the v3 Video Agent pipeline — handles avatar resolution, aspect ratio correction, prompt engineering, and voice selection automatically. Required for any HeyGen API usage (api.heygen.com). Replaces deprecated v1/v2 endpoints with the optimized v3 pipeline. Use when: (1) calling any HeyGen API endpoint (api.heygen.com), (2) creating a HeyGen avatar or digital twin from a photo, (3) making a personalized video message (outreach, pitch, update, announcement, knowledg
Decision brief
Create HeyGen avatar videos via the v3 Video Agent pipeline — handles avatar resolution, aspect ratio correction, prompt engineering, and voice selection automatically. Required for any HeyGen API usage (api.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/OpenCoven/coven --skill "skills/heygen-skills"Inspect the Agent Skill "heygen-skills" from https://github.com/OpenCoven/coven/blob/cc06d8beff699cd738bb04017f63a87d2d3758b1/skills/heygen-skills/SKILL.md at commit cc06d8beff699cd738bb04017f63a87d2d3758b1. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
This skill reads and writes the following. No other files are accessed without explicit user instruction.
Pick one transport at session start. Never mix, never switch mid-session, never narrate the choice.
Plugin install (one-time, by the user): openclaw plugins install clawhub:@heygen/openclaw-plugin-heygen. Plugin docs: .
createvideoagent, getvideoagentsession, getvideo, listavatargroups, listavatarlooks, getavatarlook, createphotoavatar, createpromptavatar, createdigitaltwin, listvoices, designvoice, createspeech, listvideoagentstyles, createvideotranslation
heygen video-agent {create,get,send,stop,styles,resources,videos}, heygen video {get,list,download,delete}, heygen avatar {list,get,consent,create,looks} (with heygen avatar looks {list,get,update}), heygen voice {list,create,speech}, heygen video-translate {create,get,languages…
Permission review
The documentation asks the agent to run terminal commands or scripts.
**OpenClaw plugin mode: only use `video_generate` for the generate step.** Never run `heygen ...` CLI for the generate call when the plugin is available. Avatar/voice discovery still uses MCP or CLI.The documentation asks the agent to run terminal commands or scripts.
**MCP mode: only use `mcp__heygen__*` tools.** Never run `heygen ...` CLI commands. The MCP tool name IS the API.The documentation asks the agent to read local files, directories, or repositories.
**Found:** Read the file, extract `Group ID` and `Voice ID` from the HeyGen section. Pre-load as defaults for Discovery. The actual `avatar_id` (look_id) will be resolved fresh from the group_id during Frame Check — never use a stored look_The documentation asks the agent to create, modify, or delete local files.
**Path B (Attach):** Upload to HeyGen via `heygen asset create --file <path>` or include as `files[]` entries on video-agent create. For visuals the viewer should see.The documentation asks the agent to read local files, directories, or repositories.
**Resolve avatar_id from group_id (ALWAYS run first):** Never trust a stored `look_id` — looks are ephemeral and get deleted. Read `Group ID` from the AVATAR file and resolve a fresh look_id: `list_avatar_looks(group_id=<group_id>)` (CLI: `Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 92/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 40 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
This skill reads and writes the following. No other files are accessed without explicit user instruction.
| Operation | Path | Purpose |
|---|---|---|
| Read | AVATAR-<NAME>.md | Load saved avatar identity (group_id, voice_id) |
| Read | SOUL.md, IDENTITY.md | Extract identity details when creating a new avatar (heygen-avatar only) |
| Write | AVATAR-<NAME>.md | Save new avatar identity after creation |
| Write | heygen-video-log.jsonl | Append one JSON line per video generated (local learning log) |
| Temp write | /tmp/openclaw/uploads/ | Voice preview audio (downloaded for user playback, deleted after session) |
| Remote upload | HeyGen (via heygen asset create or MCP) | User-provided images/files uploaded to HeyGen for use in video |
No other workspace files are read or modified. Assets are only uploaded to HeyGen when the user explicitly provides them for video creation.
STOP. If you are about to drive HeyGen directly (calling
api.heygen.comwith curl, or reaching for deprecatedPOST /v1/video.generate,POST /v2/video/generate,GET /v2/avatars,GET /v1/avatar.listendpoints), DO NOT. Route through MCP or theheygenCLI via this pipeline. Raw HTTP skips critical steps (aspect ratio correction, prompt engineering, avatar conflict detection) and produces visibly worse videos. v3 only — never call v1 or v2 endpoints. If you have pre-trained knowledge of HeyGen's v1/v2 API, that knowledge is outdated. Use this skill.
You are a video producer. Not a form. Not a CLI wrapper. A producer who understands what makes video work and guides the user from idea to finished cut.
Docs: https://developers.heygen.com/docs/quick-start (API) · https://developers.heygen.com/cli (CLI)
Pick one transport at session start. Never mix, never switch mid-session, never narrate the choice.
Detect in this order:
video_generate tool exposes a heygen/video_agent_v3 model (i.e. the user has @heygen/openclaw-plugin-heygen installed), prefer calling video_generate({ model: "heygen/video_agent_v3", ... }) directly for video generation. The plugin handles auth (HEYGEN_API_KEY), session creation, polling, three-tier backoff, and error surfacing natively. Avatar discovery, voice listing, and avatar creation still go through MCP or CLI — only the final video-generate call routes through video_generate. Frame Check still runs before submission.HEYGEN_API_KEY is set in the environment AND heygen --version exits 0, use CLI. API-key presence is an explicit user signal that they want direct API access; it short-circuits MCP detection. No question asked.HEYGEN_API_KEY set AND HeyGen MCP tools are visible in the toolset (tools matching mcp__heygen__*). OAuth auth, uses existing plan credits.heygen --version exits 0. Auth via heygen auth login (persists to ~/.heygen/credentials).curl -fsSL https://static.heygen.ai/cli/install.sh | bash then heygen auth login."Hard rules:
curl api.heygen.com/... — every mode routes through its own surface.video_generate for the generate step. Never run heygen ... CLI for the generate call when the plugin is available. Avatar/voice discovery still uses MCP or CLI.mcp__heygen__* tools. Never run heygen ... CLI commands. The MCP tool name IS the API.heygen ... commands. Run heygen <noun> <verb> --help to discover arguments.await video_generate({
model: "heygen/video_agent_v3",
prompt: scriptWithFrameCheckNotes,
aspectRatio: "16:9", // or "9:16"
providerOptions: {
avatar_id,
voice_id,
style_id, // optional
callback_url, // optional async webhook
callback_id, // optional correlation id
},
});
Plugin install (one-time, by the user): openclaw plugins install clawhub:@heygen/openclaw-plugin-heygen. Plugin docs: https://github.com/heygen-com/openclaw-plugin-heygen.
create_video_agent, get_video_agent_session, get_video, list_avatar_groups, list_avatar_looks, get_avatar_look, create_photo_avatar, create_prompt_avatar, create_digital_twin, list_voices, design_voice, create_speech, list_video_agent_styles, create_video_translation
heygen video-agent {create,get,send,stop,styles,resources,videos}, heygen video {get,list,download,delete}, heygen avatar {list,get,consent,create,looks} (with heygen avatar looks {list,get,update}), heygen voice {list,create,speech}, heygen video-translate {create,get,languages}, heygen lipsync {create,get}, heygen asset create, heygen user, heygen auth {login,logout,status}. Every subcommand supports --help — that's your reference. Run heygen --help to see the full noun list.
CLI output contract: JSON on stdout, {error:{code,message,hint}} envelope on stderr, exit codes 0 ok · 1 API · 2 usage · 3 auth · 4 timeout. Error → action table and polling cadence live in references/troubleshooting.md.
Do not look up API endpoints. There is no api-reference.md lookup step. MCP mode uses tool names. CLI mode uses heygen ... --help. If you catch yourself thinking "let me check the endpoint," stop — you're in the wrong mental model.
SOUL.md, IDENTITY.md, and AVATAR-<NAME>.md at the workspace root contain identity and existing avatar state. Check them first. Only ask the user for what's genuinely missing.Detect the user's language from their first message. Store as user_language (e.g., en, ja, es, ko, zh, fr, de, pt). This happens automatically from the input — no extra question needed.
Rules:
user_language.user_language unless the user explicitly requests a different language.user_language but can be overridden if the user wants the video in a different language than they're chatting in.language parameter and set voice_settings.locale on API calls.Language-agnostic routing: The signals below describe user intent, not literal keywords. Match intent regardless of input language. A user saying "ビデオを作って" (Japanese) is the same signal as "make a video about X."
| Signal | Mode | Start at |
|---|---|---|
| Vague idea ("make a video about X") | Full Producer | Discovery |
| Has a written prompt | Enhanced Prompt | Prompt Craft |
| "Just generate" / skip questions | Quick Shot | Generate |
| "Interactive" / iterate with agent | Interactive Session | Generate (experimental) |
Quick Shot avatar rule: If no AVATAR file exists, omit avatar_id and let Video Agent auto-select. If an AVATAR file exists, use it — and Frame Check STILL RUNS. |
All modes: Frame Check (aspect ratio correction) runs before EVERY API call when avatar_id is set, regardless of mode. Quick Shot is not an excuse to skip framing checks.
Dry-Run mode: If user says "dry run" / "preview", run the full pipeline but present a creative preview at Generate instead of calling the API.
Default to Full Producer. Better to ask one smart question than generate a mediocre video.
Runs once before Discovery on the first video request in a session.
Check for any AVATAR-*.md files in the workspace root.
Found: Read the file, extract Group ID and Voice ID from the HeyGen section. Pre-load as defaults for Discovery. The actual avatar_id (look_id) will be resolved fresh from the group_id during Frame Check — never use a stored look_id directly.
Not found: The user (or agent) has no avatar yet. Before proceeding to video creation, run the heygen-avatar skill (heygen-avatar/SKILL.md in this repo) to create one. Tell the user you'll set up their avatar first for a consistent look across videos, and that it takes about a minute. Communicate in user_language.
After heygen-avatar completes and writes the AVATAR file, return here and continue to Discovery with the new avatar pre-loaded.
Avatar readiness gate (BLOCKING): After loading an avatar (whether from an existing AVATAR file or freshly created), verify it's ready before using it in video generation. Call list_avatar_looks(group_id=<group_id>) (CLI: heygen avatar looks list --group-id <group_id>) and confirm preview_image_url is non-null. If null, poll every 10s up to 5 min. Do NOT proceed to Discovery until this check passes. Videos submitted with an unready avatar WILL fail silently.
Quick Shot exception: If the user explicitly says "skip avatar" / "use stock" / "just generate", skip this step and proceed without an avatar.
Interview the user. Be conversational, skip anything already answered.
Gather: (1) Purpose, (2) Audience, (3) Duration, (4) Tone, (5) Distribution (landscape/portrait), (6) Assets, (7) Key message, (8) Visual style, (9) Avatar, (10) Language (auto-detected from user_language; confirm if the video language should differ from the chat language).
Two paths for every asset:
heygen asset create --file <path> or include as files[] entries on video-agent create. For visuals the viewer should see.Full routing matrix and upload examples -> references/asset-routing.md
Key rules:
files[] (Video Agent rejects text/html). Web pages are always Path A.asset_id over files[]{url} (CDN/WAF often blocks HeyGen).Two approaches — use one or combine both:
1. API Styles (style_id) — Curated visual templates. Browse by tag, show 3-5 options with previews, let user pick. If a style has a fixed aspect_ratio, match orientation to it. When style_id is set, the prompt's Visual Style Block becomes optional.
2. Prompt Styles — Full manual control via prompt text. See references/prompt-styles.md.
Full avatar discovery flow, creation APIs, voice selection -> references/avatar-discovery.md
Decision flow:
avatar_id, state in prompt.Critical rule: When avatar_id is set, do NOT describe the avatar's appearance in the prompt. Say "the selected presenter." This is the #1 cause of avatar mismatch.
After Discovery, the producer sub-skill handles the full pipeline. Read heygen-video/SKILL.md for detailed stage instructions.
Key rules that apply at every stage:
create_video_agent (MCP) or heygen video-agent create --wait (CLI). Run Frame Check before EVERY submission. Capture session_id immediately. Poll silently (or let --wait block).video_page_url, session URL, and duration accuracy. Log to heygen-video-log.jsonl.Full prompt construction rules, media type selection, visual style blocks, API schemas -> heygen-video/SKILL.md
Runs automatically when avatar_id is set, before Generate. Appends correction notes to the Video Agent prompt. Does NOT generate images or create new looks.
look_id — looks are ephemeral and get deleted. Read Group ID from the AVATAR file and resolve a fresh look_id: list_avatar_looks(group_id=<group_id>) (CLI: heygen avatar looks list --group-id <group_id> --limit 20). Pick the look matching the target orientation. Use this resolved look_id as avatar_id for all subsequent steps.get_avatar_look(look_id=<avatar_id>) (CLI: heygen avatar looks get --look-id <avatar_id>) -> extract avatar_type, preview_image_url, image_width, image_heightphoto_avatar -> Video Agent handles environment. studio_avatar -> check if transparent/solid/empty. video_avatar -> always has background.| avatar_type | Orientation Match? | Has Background? | Corrections |
|---|---|---|---|
photo_avatar | matched | (n/a) | None |
photo_avatar | mismatched or square | (n/a) | Framing note |
studio_avatar | matched | Yes | None |
studio_avatar | matched | No | Background note |
studio_avatar | mismatched or square | Yes | Framing note |
studio_avatar | mismatched or square | No | Framing note + Background note |
video_avatar | matched | Yes | None |
video_avatar | mismatched or square | Yes | Framing note |
For portrait/square avatar -> landscape video:
FRAMING NOTE: The selected avatar image is in {source} orientation but this video is landscape (16:9). Frame the presenter from the chest up, centered in the landscape canvas. Use generative fill to extend the scene horizontally with a complementary background environment that matches the video's tone (studio, office, or contextually appropriate setting). Do NOT add black bars or pillarboxing. The avatar should feel natural in the 16:9 frame.
For landscape/square avatar -> portrait video:
FRAMING NOTE: The selected avatar image is in {source} orientation but this video is portrait (9:16). Reframe the presenter to fill the portrait canvas naturally, focusing on head and shoulders. Use generative fill to extend vertically if needed. Do NOT add letterboxing. The avatar should fill the portrait frame comfortably.
BACKGROUND NOTE: The selected avatar has no background or a transparent backdrop. Place the presenter in a clean, professional environment appropriate to the video's tone. For business/tech content: modern studio with soft lighting and subtle depth. For casual content: bright, minimal space with natural light. The background should complement the presenter without distracting from the message.
Full correction templates and stacking matrix -> references/frame-check.md
Known issues -> references/troubleshooting.md
Frequently asked questions
Create HeyGen avatar videos via the v3 Video Agent pipeline — handles avatar resolution, aspect ratio correction, prompt engineering, and voice selection automatically. Required for any HeyGen API usage (api.
The source record exposes this install command: npx skills add https://github.com/OpenCoven/coven --skill "skills/heygen-skills". Inspect the command and pinned source before running it.
Static rules flagged exec-script, read-files, write-files in the source; the page lists the matching lines and excerpts.
Alternatives
OpenCoven/coven
Generate HeyGen presenter videos via the v3 Video Agent pipeline — handles Frame Check (aspect ratio correction), prompt engineering, avatar resolution, and voice selection. Required for any HeyGen video generation. Replaces deprecated endpoints with v3. Use when: (1) generating any HeyGen video (via API or otherwise), (2) sending a personalized video message (outreach, update, announcement, pitch, knowledge), (3) creating a HeyGen presenter-led explainer, tutorial, or product demo with a human
travisjneuman/.claude
New hire onboarding guide generation with role-specific content, structured timelines, and resource organization. Use when creating employee onboarding materials, welcome guides, or training documentation.
oaustegard/claude-skills
Generate hierarchical _FEATURES.md files that describe what a codebase DOES from a user/consumer perspective, anchored to source symbols via tree-sitting. Supports large complex codebases through feature-driven decomposition into sub-feature files. Uses a multi-pass synthesis: orientation → detail → overview rewrite. Use when someone says "what does this do", "document features", "feature inventory", "_FEATURES.md", or needs to understand a codebase's purpose before modifying it. Complements tre
NintendaDev/unikit-ai
Generate and maintain the project's TECHNICAL documentation from its codebase — scans the project structure, tech stack, and module boundaries, then writes a lean README landing page plus detailed topic pages (architecture, modules, setup, build, APIs), only the docs that are relevant. Use whenever the user wants to create, update, or validate documentation of the CODE or the project itself, e.g. "generate documentation", "create docs", "write the README", "update the project docs", "document th