Source profileQuality 93/100

event4u-app/agent-config/src/skills/song-to-script/SKILL.md

song-to-script

Turn an audio track into a timed `## Scene N` script: song sections → per-scene durations, auto mode adds mood + lip-sync lines. Triggers 'music video', 'from the song', 'cut to the beat'.

Source repository stars
7
Declared platforms
0
Static risk flags
0
Last source update
2026-07-28
Source checked
2026-07-28

Decision brief

What it does—and where it fits

Turn a song into /script.md — a sequence of Scene N blocks whose duration: values sum to the track length and whose cut points land on real section boundaries. Consumed by /video:from-song, then handed to scene-expander and video-director. Never invents timing — every boundary c…

Best for

  • A music-video run needs scenes cut to the song (/video:from-song).
  • An existing script must be re-timed to a track after the edit
  • The operator already supplies a Scene N script with duration:

Not for

  • Tasks that require unconfirmed production actions or broad system permissions.
  • Environments where the pinned source and install steps cannot be inspected.

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/event4u-app/agent-config --skill "src/skills/song-to-script"
Safe inspection promptEditorial

Inspect the Agent Skill "song-to-script" from https://github.com/event4u-app/agent-config/blob/0adf49a8ae84b0ff6e2de8759eea43257e020eff/src/skills/song-to-script/SKILL.md at commit 0adf49a8ae84b0ff6e2de8759eea43257e020eff. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Procedure

    One Scene N per analysis section. duration: = end - start (rounded to 0.5 s). Then clamp the plan to the chosen model's renderable envelope — read minduration / maxduration from the model-capabilities manifest (.sh capability --model ), falling back to the provider tuning for si…

    Section shorter than minduration → merge into its neighbour.Section longer than maxduration → split into sub-scenes. WithNo valid plan exists (e.g. the whole song is shorter than
  2. 02

    Step 1: Map sections → scenes (capability-clamped, beat-aligned)

    One Scene N per analysis section. duration: = end - start (rounded to 0.5 s). Then clamp the plan to the chosen model's renderable envelope — read minduration / maxduration from the model-capabilities manifest (.sh capability --model ), falling back to the provider tuning for si…

    Section shorter than minduration → merge into its neighbour.Section longer than maxduration → split into sub-scenes. WithNo valid plan exists (e.g. the whole song is shorter than
  3. 03

    Step 2: Assign mood + action

    First decide the subject mode:

    Character mode — character.json exists: every scene's action:Style mode — no character.json: scenes describe setting, palette,Lyric segment (the vocal map places ≥1 transcribed line inside
  4. 04

    Step 3: Vocal map — transcribe, never guess (vocal tracks)

    When the track has vocals and the run intends lip-sync, build a vocal map from the real audio before assigning any dialogue::

    Transcribe the audio to timestamped lines. Adapter-first: aLabel the singer per line — map diarization labels (or unlabeledEmit /vocal-map.json:
  5. 05

    Step 4: Emit + reconcile

    Write /script.md (and /vocal-map.json when the track has vocals). Report the delta, the section→scene map, the probe method (so the operator sees whether cuts are silence-derived, energy-derived, or interval-fallback), and whether lyric timing is transcript-derived (it must be —…

    Write /script.md (and /vocal-map.json when the track has vocals). Report the delta, the section→scene map, the probe method (so the operator sees whether cuts are silence-derived, energy-derived, or interval-fallback),…

Permission review

Static risk signals and limitations

No configured static risk pattern was detected

This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score93/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars7SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
event4u-app/agent-config
Skill path
src/skills/song-to-script/SKILL.md
Commit
0adf49a8ae84b0ff6e2de8759eea43257e020eff
License
MIT
Collected
2026-07-28
Default branch
main
View the original SKILL.md

song-to-script

Turn a song into <project>/script.md — a sequence of ## Scene N blocks whose duration: values sum to the track length and whose cut points land on real section boundaries. Consumed by /video:from-song, then handed to scene-expander and video-director. Never invents timing — every boundary comes from the audio probe, and the probe's method tells this skill how musical (or not) those boundaries actually are.

When to use

  • A music-video run needs scenes cut to the song (/video:from-song).
  • An existing script must be re-timed to a track after the edit drifted from the beat (the named second consumer — re-time without a full re-author).

Do NOT use when:

  • The operator already supplies a ## Scene N script with duration: values — feed it straight to scene-expander.
  • There is no audio — use the operator brief with scene-expander directly.

Inputs

  • Audio analysis — adapter-first, probe as the floor:
    • Analysis adapter (when an audio-analysis provider is configured — see audio-adapter-contract.md): {bpm, beats, downbeats, sections:[{start,end,label,energy?}]} — real musical structure. Beats/downbeats become the candidate cut grid; section labels are musical (verse, chorus, …).
    • Audio probe — JSON from scripts/ai-video/lib/probe-audio.sh: {duration, method, warning?, sections:[{start,end,energy,label}]}.
      • method: silence — boundaries are real quiet gaps; trust them as cuts.
      • method: rms — boundaries are energy-delta inflections; usable but coarse.
      • method: intervalthe track is structurally flat (brick-walled / sustained); sections are fixed-interval, NOT musical. When method is interval (or warning is set), the emitted script header states that timing is interval-based and the operator should pass --scene-durations for musical sync. Never present interval cuts as beat-synced.
  • Model capabilities — the chosen video model's renderable envelope from the multiplexer manifest: scripts/ai-video/adapters/<provider>.sh capability --model <id>{min_duration, max_duration, audio_sync, aspect, verified}. Scene durations MUST land inside [min_duration, max_duration]; a verified: false manifest entry is surfaced in the report, never trusted silently.
  • Modebrief (operator text is the creative source) or auto (infer mood + action from energy).
  • Brief (brief mode only) — free text: story, settings, look.
  • Character lock (optional) — <project>/character.json if a human subject was locked. Absent is normal — abstract / landscape / visualiser videos have no locked subject; see Step 2.

Procedure

Step 1: Map sections → scenes (capability-clamped, beat-aligned)

One ## Scene N per analysis section. duration: = end - start (rounded to 0.5 s). Then clamp the plan to the chosen model's renderable envelope — read min_duration / max_duration from the model-capabilities manifest (<provider>.sh capability --model <id>), falling back to the provider tuning for single-model adapters:

  • Section shorter than min_duration → merge into its neighbour. With beat data, merge toward the neighbour that keeps the joined cut on a downbeat (else any beat); without beat data, merge into the shorter neighbour.
  • Section longer than max_duration → split into sub-scenes. With beat data, place every split point on the nearest downbeat (else beat) to the equal-division point — never mid-beat; without beat data, split equally.
  • No valid plan exists (e.g. the whole song is shorter than min_duration, or a section cannot be split onto any beat inside the envelope) → halt and surface the conflict with the model id and the violated bound — an unbuildable plan never reaches the renderer.

Every emitted scene satisfies min_duration ≤ duration ≤ max_duration. When the manifest entry is verified: false, say so in the report — the envelope is documented-best-effort, not a smoke-traced fact.

Step 2: Assign mood + action

First decide the subject mode:

  • Character modecharacter.json exists: every scene's action: names the locked subject, never a fresh description.
  • Style mode — no character.json: scenes describe setting, palette, and motion continuity (the recurring look), not a person. This is the valid abstract / landscape / visualiser path — do not invent a human subject to fill the slot.

Then pick the prompt source per segment — the modality switch:

  • Lyric segment (the vocal map places ≥1 transcribed line inside it) → the scene prompt derives from the lyric line itself: its imagery, subjects, and verbs seed mood: + action: (in character mode, acted by the locked subject; in style mode, rendered as setting / weather / palette — never an invented human). The line lands in dialogue: per Step 3.
  • Instrumental segment (no vocal-map line) → the scene prompt derives from the audio features: section label + energy via the intent table below. Never recycle a lyric from another segment into an instrumental one.

Then assign per scene:

  • Brief mode — distribute the brief's beats across scenes in order; the modality switch still applies (lyric segments quote the brief's matching beat through the lyric's lens), and energy modulates pacing. Do not add story the brief did not state.

  • Auto mode — derive mood per section from energy and label (probe labels and musical labels from the analysis adapter both map):

    label / energydefault scene intent
    intro / lowestablishing wide, slow camera, calm subject/scene
    verse / midnarrative motion, medium framing, follow the subject
    build / risingapproach, tightening framing
    chorus · drop / peakdynamic motion, weather/FX, fast push
    bridge · breakdown / dipclose-up / detail, quiet, single light source
    outro / fadepull-back, resolve, hold

Energy → cut frequency + motion intensity. Section energy (0..1, relative to the track mean) drives both how often the edit cuts and how hard the camera moves — chorus = faster cuts / more motion:

energy vs. track meancut length targetcamera: motion intensity
≥ mean + 0.10 (chorus / drop)short — split the section toward min_duration, one scene per 1–2 downbeat barsfast push / whip / handheld shake
within ±0.10 of mean (verse / build)medium — one scene per section or per 4-bar phrasesteady dolly, slow tighten
≤ mean − 0.10 (breakdown / outro)long — merge toward max_duration, hold shotslocked-off or slow drift

High-energy splitting and low-energy merging both stay inside the Step 1 capability envelope and land on downbeats — the energy table chooses where inside the envelope a scene length falls, never outside it.

Step 3: Vocal map — transcribe, never guess (vocal tracks)

LYRIC TIMING AND SINGER COME FROM THE TRANSCRIBED AUDIO, NEVER FROM A
BRIEF / STORY SKELETON OR A GUESSED STRETCH. NEVER PUT ONE SINGER'S
LINE ON ANOTHER SINGER'S SCENE.

When the track has vocals and the run intends lip-sync, build a vocal map from the real audio before assigning any dialogue::

  1. Transcribe the audio to timestamped lines. Adapter-first: a configured lyrics provider (e.g. audio-adapters/whisperx.sh) returns word-level timestamps plus per-line diarization labels (SPEAKER_00, …, or "?" when ambiguous):
    echo '{"audio_path":"<vocal-stem-or-song>"}' \
      | scripts/ai-video/audio-adapters/whisperx.sh analyze
    
    No lyrics provider configured → OpenAI /v1/audio/transcriptions (response_format=verbose_jsonsegments[].{start,end,text}) or local whisper as before (no speaker labels — every line starts as "?"). Either way the transcript is the only source of lyric timing.
  2. Label the singer per line — map diarization labels (or unlabeled lines) to cast names via the operator's who-sings reference (a roster, a brief that names who sings which line, or a character cast): one diarization label ↦ one cast name, consistently. If a line's singer is genuinely ambiguous (label "?", mixed-speaker line, or no roster match), keep singer: "?" and surface it — never guess a singer to fill the slot.
  3. Emit <project>/vocal-map.json: [{start, end, text, singer}], timing verbatim from the transcript.
  4. Validate — run the ground-truth enforcer before handing the map to the sign-off gate:
    scripts/ai-video/lib/validate-vocal-map.sh <project>/vocal-map.json \
      <project>/transcript.json --roster "<cast names>"
    
    It rejects re-timed lines, lyrics not in the transcript, and missing singers (exit 7, specific line named). A red validator is a halt — fix the map, never bypass.
  5. Place lines into the matching scene's dialogue: block using the transcript timing, tagged with the singer (singer: "<line>"). A scene's lip-sync subject MUST be the line's labelled singer; a "?" line gets NO lip-sync scene until the operator resolves it.

No vocals / no transcript / no lip-sync intent → leave dialogue: empty; the scene is performance / B-roll. Never fabricate lyrics, never re-time a line off the brief, and in style mode dialogue: stays empty (lip-sync needs a character subject). The /video:from-song sign-off gate (its Step 6) shows this map for approval before any render.

Step 4: Emit + reconcile

Write <project>/script.md (and <project>/vocal-map.json when the track has vocals). Report the delta, the section→scene map, the probe method (so the operator sees whether cuts are silence-derived, energy-derived, or interval-fallback), and whether lyric timing is transcript-derived (it must be — never brief-derived). If the sum cannot be reconciled (e.g. provider max-duration forces more time than the song has), halt and surface the conflict — do not pad silently.

Step 5: Validate before handoff

Concrete checks (all must pass before the script is handed to scene-expander):

  • Assert Σ(duration) == probe.duration within ±1.0 s; report the exact delta. A larger delta → halt, do not pad.
  • Verify every scene boundary equals a probe section boundary (or a --scene-durations value) — no invented cut points.
  • Confirm no scene duration: exceeds the model's max_duration or falls below min_duration (model-capabilities manifest, or the provider tuning for single-model adapters).
  • Verify every lyric-segment scene derives its prompt from its own vocal-map line and every instrumental scene from audio features — no cross-segment lyric recycling (modality switch).
  • Ensure every ## Scene N carries all five keys (duration · mood · action · camera · dialogue), and that dialogue: is empty in style mode.

Output format

  1. script.md opens with the derivation header — # <project> — derived from <song-file> (<mode> mode · cuts: <method>) — so the probe method stays visible downstream.
  2. One ## Scene N block per cut carrying exactly the keys duration · mood · action · camera · dialoguescene-expander consumes this verbatim; keep the keys exact.
  3. dialogue: stays empty unless operator-supplied lyrics cover the section — detected vocal energy alone never fills it.
# <project> — derived from <song-file> (<mode> mode · cuts: <method>)

## Scene 1
duration: 6.0
mood: establishing, cold, pre-storm
action: <subject from character.json, OR style description in style mode>
camera: slow push-in
dialogue:

## Scene 2
duration: 4.5
mood: build, rising tension
action: close on <subject / detail>, wind picking up
camera: handheld tighten
dialogue:
  - "<subject>: \"<lyric line for this section, if any>\""

Gotcha

  • method: interval is the brick-walled-master signal, not a bug. A compressed modern master has near-constant RMS and no silence, so the probe degrades to fixed intervals. That is the honest floor — surface it and point the operator at --scene-durations; never dress interval cuts up as beat-synced.
  • A vocal section without supplied lyrics is B-roll, not lip-sync. Detected vocal energy alone does not authorise dialogue: — only operator-supplied lyrics do.
  • Style mode is the default for a no-character run, not an error path. Landscape / abstract / visualiser videos never get a fabricated human subject.

Do NOT

  • Do NOT invent timing. Every cut maps to a probe boundary or a --scene-durations value — never to taste.
  • Do NOT present interval-fallback cuts as beat-synced. Always surface the probe method.
  • Do NOT emit a clip outside the provider's min/max duration — split/merge in Step 1 instead.
  • Do NOT fabricate lyrics or story beyond the brief / detected vocals.
  • Do NOT invent a human subject in style mode; defer identity to character.json only when a lock exists.
  • Do NOT pad a unreconcilable timing sum — halt and surface it.

See also