Best for
- Generate voiceover/narration from text
- Create text-to-speech audio for videos
- Add, replace, or redo narration/voiceover for an existing video, timeline,
0xsline/OpenChatCut/src/agent/skills/voice/SKILL.md
Text-to-Speech (TTS), voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.
Decision brief
Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete provider and voice before calling submitvoice.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/0xsline/OpenChatCut --skill "src/agent/skills/voice"Inspect the Agent Skill "voice" from https://github.com/0xsline/OpenChatCut/blob/ce56a9392ce46349c02d87ddfccc79f54e87fe10/src/agent/skills/voice/SKILL.md at commit ce56a9392ce46349c02d87ddfccc79f54e87fe10. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Generate voiceover/narration from text
If the current request has an existing visual target and the user wants narration, voiceover, dubbing, or replacement speech for that target, read references/video-sync.md before drafting new narration, using existing narration text to generate TTS, or placing audio. Do this eve…
When the user needs TTS and has not already chosen a concrete voice, first separate providers with curated OpenChatCut choices from providers that require an account-specific voice ID.
For ordinary editing sound effects (SFX), do not generate first. Use the built-in Sound Effects library before generating:
Review the “Parameters” section in the pinned source before continuing.
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 95/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 1,360 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete
provider and voice before calling submit_voice.
If the current request has an existing visual target and the user wants narration, voiceover, dubbing, or replacement speech for that target, read references/video-sync.md before drafting new narration, using existing narration text to generate TTS, or placing audio. Do this even when the user did not explicitly say "sync" or "match the visuals"; the existence of a visual target means narration timing and meaning may need to follow on-screen content. Use the normal standalone TTS path only when there is no visual target or the user just wants an audio asset from text.
Also read references/video-sync.md when the timeline already has narration/voiceover and the user asks to change the visuals while keeping that voiceover aligned. This is a sync maintenance task even if no new TTS is needed.
Use submit_voice to create a TTS audio asset. The current MCP tool contract is:
provider is required. Configured choices may be doubao, elevenlabs,
minimax, inworld, fishaudio, speechify, openai, gemini,
mistral, or cartesia. All providers are opt-in; use only providers shown
as configured in the capabilities prompt.voiceId is required, concrete, and provider-specific. The only exception is
deliberate MiniMax timbreWeights mixing, where voiceId must be empty. Do
not mix catalogs./voice-samples/... URL.modelId,
speed, outputFormat, and instructions; Gemini supports modelId,
outputFormat, and instructions; Mistral supports modelId and
outputFormat; Cartesia supports modelId, speed, languageCode, and
outputFormat. Omit unsupported or unrequested fields.voiceId plus optional
modelId. Do not pass expressive, speed, language, or output controls to
these providers.submit_voice creates an audio asset only. Timeline placement, replacement,
trimming, and alignment happen later with timeline tools.submit_voice calls can be useful: split at
natural pauses, sentence groups, or script beat boundaries when the workflow
benefits from separately timed or placed voice clips.speedRatio, loudnessRatio, pitch, emotion,
emotionScale, performancePrompt, and explicitDialect, but not every
voice supports every expressive control. Check
references/voices.md before using them.Doubao control support for current curated voices:
vivi, xiaohe, yunzhou, xiaotian, naiqimengwa, yingtaowanzi,
wenroumama, zhixingnv, dayi, jitangnv, liuchang, ruyayichen,
morgan, qingcang, huiben, popo, yuanboxiaoshu, baqiqingshu, and
tangseng support explicit emotion / emotionScale,
performancePrompt, and ASMR-style prompt directions.shuanglangshaonian supports performancePrompt and COT/QA-style
instruction following, but does not support explicit emotion /
emotionScale or ASMR-style control.explicitDialect is only supported by vivi and can be dongbei,
shaanxi, or sichuan.ElevenLabs control support for current curated voices:
amelia, brittney, hope, jessica, arabella, jane, maria,
mark, frederick, peter, james, jon, sully, david, and alex
all support the same request-level controls; model-specific support is still
validated by ElevenLabs.eleven_v3, inline audio tags are available when the user
asks for expressive delivery such as emotion, tone, nonverbal cues, accent
hints, or local pacing. Official examples fit these useful TTS categories:
emotion/tone tags such as [happy], [sad], [angry], [excited],
[curious], [sarcastic], [crying], [annoyed], [appalled],
[thoughtful], [surprised], and [mischievously]; vocal delivery and
nonverbal cue tags such as [whispers], [laughs], [sighs], [exhales],
[inhales deeply], [clears throat], [snorts], [swallows],
[wheezing], and [coughs];
pacing/pause/local speed tags such as [slowly], [pause],
[short pause], [long pause], [rushed], and [drawn out]; and
accent/special-performance tags such as
[strong X accent], for example [strong French accent], plus [sings],
[singing], [woo], and [pirate voice]. Official examples are
non-exhaustive; similar auditory tags can be tried when the user explicitly
asks for that delivery and the tag describes how the voice should sound, not
a visual action. Write tags directly in text, close to the short phrase
they should affect. Treat tags as local guidance, not paragraph-wide controls.eleven_v3, use punctuation, text structure,
shorter generated segments, or local audio tags such as [short pause] and
[slowly] when needed.// English / multilingual via ElevenLabs
submit_voice({
provider: "elevenlabs",
text: "Hello world",
voiceId: "peter",
});
// Chinese via Doubao
submit_voice({
provider: "doubao",
text: "你好世界",
voiceId: "liuchang",
});
// With speed adjustment (Doubao only)
submit_voice({
provider: "doubao",
text: "这是一段稍快的中文旁白。",
voiceId: "liuchang",
speedRatio: 1.5,
});
// With expressive Doubao controls
submit_voice({
provider: "doubao",
text: "这次事故提醒我们,安全永远不能侥幸。",
voiceId: "liuchang",
emotion: "sad",
emotionScale: 3,
performancePrompt: "痛心但克制,语速稍慢,像新闻专题旁白",
pitch: -1,
speedRatio: 0.92,
});
// With ElevenLabs delivery controls
submit_voice({
provider: "elevenlabs",
text: "The launch changed how teams plan their daily work.",
voiceId: "peter",
speed: 0.95,
stability: 0.4,
similarityBoost: 0.8,
outputFormat: "wav_44100",
});
// MiniMax TTS (when configured) — see references/minimax-tts.md
submit_voice({
provider: "minimax",
text: "欢迎使用视频编辑助手。",
voiceId: "female-yujie",
speed: 1,
name: "VO · welcome",
});
// Cartesia shape after the user confirms the exact account voice ID.
// confirmedCartesiaVoiceId represents that supplied value, not a preset.
submit_voice({
provider: "cartesia",
text: "A concise product introduction.",
voiceId: confirmedCartesiaVoiceId,
modelId: "sonic-3",
speed: 1,
languageCode: "en",
outputFormat: "mp3",
});
When the user needs TTS and has not already chosen a concrete voice, first separate providers with curated OpenChatCut choices from providers that require an account-specific voice ID.
For Doubao, ElevenLabs, or MiniMax, read references/voices.md before recommending, rendering, or submitting an option. Use it as the only source for curated preset IDs, provider choice, display labels, tags, and bundled sample URLs. Do not create voice options from memory, translated names, or broad user descriptions.
For Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, or Cartesia, do not
offer an invented audition list or sample URL. Ask the user for the concrete
voice ID from that configured provider. A broad description such as "warm
female" is not a valid voiceId.
First determine two separate languages:
form-visual label, visual-option name,
and summary.The audition widget's submit button is fixed to the default label in this build
(submitLabel is accepted but not rendered); keep the question label and option
labels in the user conversation language, not the target narration language. For example:
English users see submit_label="Submit", Chinese users see
submit_label="提交", and Spanish users see submit_label="Enviar".
"help me generate ... voice over in Chinese" is an English conversation asking for Chinese narration, so the audition widget copy stays in English while the voice candidates come from Doubao.
For a curated provider:
references/voices.md by target narration language / provider and
explicit requirements such as gender, age range, tone, and use case.widget-forms, then call ask_followup_questions with voice options
and real bundled audio samples.submit_voice with the selected preset ID as voiceId.For a provider without a curated OpenChatCut catalog, ask for a free-text,
concrete provider voice ID instead. Do not add media or synthesize a
/voice-samples/... path. Wait for the user to supply/confirm the exact ID
before calling submit_voice.
For each curated audition option, keep value, display label, media, and
summary tied to the same preset row from references/voices.md. Use only the
sample URLs recorded there. Keep value as the preset ID and media as its
matching sample URL. Write name and summary in the user's conversation
language. The target narration language only decides the provider/voice
catalog. After submission, map the display name back to the preset ID from the
same candidate list.
English request for Chinese narration:
<widget submit_label="Submit">
<form-visual
id="voiceId"
label="For Chinese voiceover, I recommend a few voices to try:"
required="true"
>
<visual-option
value="vivi"
name="Vivi"
media="/voice-samples/doubao-vivi.mp3"
aspect-ratio="16:5"
summary="Female / young / friendly, general"
/>
<visual-option
value="xiaohe"
name="Xiaohe"
media="/voice-samples/doubao-xiaohe.mp3"
aspect-ratio="16:5"
summary="Female / young / soft, clear"
/>
<visual-option
value="yunzhou"
name="Yunzhou"
media="/voice-samples/doubao-yunzhou.mp3"
aspect-ratio="16:5"
summary="Male / young / neutral, business"
/>
</form-visual>
</widget>
Chinese request for Chinese narration:
<widget submit_label="提交">
<form-visual
id="voiceId"
label="我推荐这几个中文旁白音色,先试听一下:"
required="true"
>
<visual-option
value="morgan"
name="Morgan"
media="/voice-samples/doubao-morgan.mp3"
aspect-ratio="16:5"
summary="男 / 中年 / 低沉知识解说"
/>
<visual-option
value="zhixingnv"
name="知性女声"
media="/voice-samples/doubao-zhixingnv.mp3"
aspect-ratio="16:5"
summary="女 / 中年 / 冷静知识讲解"
/>
<visual-option
value="vivi"
name="Vivi"
media="/voice-samples/doubao-vivi.mp3"
aspect-ratio="16:5"
summary="女 / 年轻 / 亲切通用口播"
/>
</form-visual>
</widget>
For ordinary editing sound effects (SFX), do not generate first. Use the built-in Sound Effects library before generating:
browse_library with category:"sound-effects" and a query such as
"whoosh", "camera shutter", "notification", "censor beep", or
"record scratch".library:sound:<id>.edit_item, using fromFrame as the sound's
anchor/editorial moment frame:browse_library({
category: "sound-effects",
query: "short whoosh transition",
});
edit_item({
adds: [
{
type: "audio",
assetId: "library:sound:whoosh-short",
fromFrame: 120,
trackId: "A1",
},
],
});
Only generate sound effects from text descriptions with submit_sound when:
browse_library({ category:"sound-effects", query }) returns no suitable
match.// Custom/generated sound effect after the library has no suitable match
submit_sound({ prompt: "A dog barking in the distance" });
// With custom duration (0.5-22 seconds)
submit_sound({
prompt: "Thunder and heavy rain",
durationSeconds: 15,
});
// High prompt adherence
submit_sound({
prompt: "Sci-fi laser gun firing",
promptInfluence: 0.8,
});
Tips for better results:
| Field | Description | Notes |
|---|---|---|
provider | doubao, elevenlabs, minimax, inworld, fishaudio, speechify, openai, gemini, mistral, or cartesia | Required; configured choices only |
text | Text to synthesize | Required |
voiceId | Concrete provider-specific voice ID | Required except MiniMax timbre mix |
modelId | Provider model override | ElevenLabs, Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, Cartesia |
speed | Speech speed | ElevenLabs, MiniMax, OpenAI, Cartesia |
languageCode | Language hint/code | ElevenLabs, Cartesia |
outputFormat | Provider-supported output format | ElevenLabs, OpenAI, Gemini, Mistral, Cartesia |
instructions | Natural-language delivery direction | OpenAI, Gemini |
speedRatio | Speech speed | Doubao only |
name | Media-pool asset name | Optional |
| Field | Description | Notes |
|---|---|---|
prompt | Sound description | Required |
durationSeconds | Duration | 0.5-22 seconds |
promptInfluence | Prompt adherence | 0-1 |
name | Asset name | Optional |
Use the submit_voice voiceId guide and
references/voices.md for the current curated preset
list, display labels, tags, and sample URLs.
The curated catalog contains separate Doubao, ElevenLabs, and MiniMax IDs.
vivi / dayi are only Doubao; mark / amelia / james are only
ElevenLabs; female-yujie is only MiniMax. Inworld, Fish Audio, Speechify,
OpenAI, Gemini, Mistral, and Cartesia require a concrete provider-specific ID
confirmed by the user and have no bundled OpenChatCut samples.
Provider choice:
Frequently asked questions
Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete provider and voice before calling submitvoice.
The source record exposes this install command: npx skills add https://github.com/0xsline/OpenChatCut --skill "src/agent/skills/voice". Inspect the command and pinned source before running it.
Alternatives
simota/agent-skills
Collecting user feedback via NPS surveys, review analysis, sentiment analysis, feedback classification, and insight extraction reports. Use when establishing feedback loops.
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program
coreyhaines31/marketingskills
When the user wants to reduce churn, build cancellation flows, set up save offers, recover failed payments, or implement retention strategies. Also use when the user mentions 'churn,' 'cancel flow,' 'offboarding,' 'save offer,' 'dunning,' 'failed payment recovery,' 'win-back,' 'retention,' 'exit survey,' 'pause subscription,' 'involuntary churn,' 'people keep canceling,' 'churn rate is too high,' 'how do I keep users,' or 'customers are leaving.' Use this whenever someone is losing subscribers o
prowler-cloud/prowler
PostgreSQL indexing best practices for Prowler: index design, partial indexes, partitioned table indexing, EXPLAIN ANALYZE validation, concurrent operations, monitoring, and maintenance. Trigger: When creating or modifying PostgreSQL indexes, analyzing query performance with EXPLAIN, debugging slow queries, reviewing index usage statistics, reindexing, dropping indexes, or working with partitioned table indexes. Also trigger when discussing index strategies, partial indexes, or index maintenance