Best for
- Use when: (1) iterating on a clip with natural-language edits instead of regenerating ("make the phone invisible, keep everything else the same"), (2) generating 3-10s 720p clips with synthesized audio, rendered on-scre…
calesthio/OpenMontage/.agents/skills/gemini-omni/SKILL.md
Generate and conversationally edit short videos with Google Gemini Omni Flash (`gemini-omni-flash-preview`). Use when: (1) iterating on a clip with natural-language edits instead of regenerating ("make the phone invisible, keep everything else the same"), (2) generating 3-10s 720p clips with synthesized audio, rendered on-screen text, or timecoded beats, (3) binding reference images to roles with <FIRST_FRAME>/<IMAGE_REF_N> prompt tags, (4) editing an existing uploaded video. Accessed via the `g
Decision brief
Gemini Omni is Google DeepMind's video generation and editing model family, announced at I/O 2026. The first model, Gemini Omni Flash (gemini-omni-flash-preview, developer access since June 30, 2026), generates 3-10 second clips at 720p/24fps with synthesized audio via the Gemin…
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/calesthio/OpenMontage --skill ".agents/skills/gemini-omni"Inspect the Agent Skill "gemini-omni" from https://github.com/calesthio/OpenMontage/blob/cd9f3c1f03368be87b140af494914b8ee4e3c7a4/.agents/skills/gemini-omni/SKILL.md at commit cd9f3c1f03368be87b140af494914b8ee4e3c7a4. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Route through videoselector for generation operations. Editing (editvideo) is a direct-tool operation — call geminiomnivideo from the registry, because the multi-turn interaction state lives outside the selector's model.
Describe scene + camera + lighting + motion + audio. Official example:
Schedule beats with bracketed ranges or natural language — this maps directly onto OpenMontage scene-plan timings:
Audio is synthesized automatically; direct it in the prompt: "Include calm background music," "The audio is a low tinny radio broadcast in the background." Rendered text works and can be timed:
Pass local images via referenceimagepaths (they are sent in order), then bind them to roles inside the prompt with tags. indexes from 0 in the order supplied:
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 90/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 50,127 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Gemini Omni is Google DeepMind's video generation and editing model family, announced at I/O 2026. The first model, Gemini Omni Flash (gemini-omni-flash-preview, developer access since June 30, 2026), generates 3-10 second clips at 720p/24fps with synthesized audio via the Gemini Interactions API. Its differentiator in the OpenMontage fleet is stateful conversational editing: each generation returns an interaction_id, and a follow-up call with previous_interaction_id edits that video in place — no other wrapped provider can refine a clip without regenerating it.
OpenMontage wraps it as gemini_omni_video (native Gemini API, no gateway). It shares GOOGLE_API_KEY/GEMINI_API_KEY with google_imagen and google_tts — one key, three capabilities. Paid tier only: ~$0.10 per second of output video (billed as 5,792 output tokens/sec at $17.50/1M).
Other documented routes are available when the direct Google key is not the chosen provider:
| Route | OpenMontage call | Important limitation |
|---|---|---|
| fal.ai | gemini_omni_fal | T2V, I2V, reference video, and edit endpoints; no Google interaction ID is returned |
| Runway | runway_video, model: "gemini_omni_flash" | T2V/I2V/V2V; video edits accept up to five image references |
| ComfyUI Partner Node | comfyui_video, model_family: "gemini_omni_flash" | Hosted paid node; requires network, Comfy login, and credits |
Use the direct gemini_omni_video route for stateful conversational editing.
Gateway routes return ordinary provider tasks and cannot preserve Google's
previous_interaction_id workflow. The fal edit endpoint can still be iterated
by feeding each output video URL into the next edit call.
| Use it for | Prefer another provider for |
|---|---|
| Iterative refinement — generate, review, then edit the same clip in layers | One-shot cinematic hero clips (→ Seedance 2.0, see seedance-2-0) |
| Editing an existing/uploaded clip (restyle, add/remove objects, change text) | Clips longer than 10s or above 720p |
| On-screen rendered text and word-by-word text beats | Seed-reproducible generations (no seed support) |
| Reference-image-bound subjects/styles via prompt tags | First/last-frame interpolation (→ veo_video) |
| Timecode-scheduled multi-beat clips from one prompt | Non-English narration (English only fully supported) |
Route through video_selector for generation operations. Editing (edit_video) is a direct-tool operation — call gemini_omni_video from the registry, because the multi-turn interaction state lives outside the selector's model.
Describe scene + camera + lighting + motion + audio. Official example:
Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air.
negative_prompt parameter: "No dialogue," "No extra sound effects."Schedule beats with bracketed ranges or natural language — this maps directly onto OpenMontage scene-plan timings:
[0-3s] A person is walking [3-6s] They stop and turn around
"After 3 seconds, a woman enters the scene." / "At 5s the chorus starts in the background audio."
Audio is synthesized automatically; direct it in the prompt: "Include calm background music," "The audio is a low tinny radio broadcast in the background." Rendered text works and can be timed:
One word on the screen at a time: 'did, you, know, that, Omni, can, do, awesome, text?' Each word appears for 1s.
<FIRST_FRAME> / <IMAGE_REF_N> tags)Pass local images via reference_image_paths (they are sent in order), then bind them to roles inside the prompt with tags. <IMAGE_REF_N> indexes from 0 in the order supplied:
in the style of <IMAGE_REF_0> a woman <IMAGE_REF_1> is walking
[0-3s] A studio fashion sequence. Starting with woman <IMAGE_REF_0>, she is
holding <IMAGE_REF_1> [3-6s] Then we see the man <IMAGE_REF_2> holding <IMAGE_REF_3>
<FIRST_FRAME> makes an image the opening frame: <FIRST_FRAME> a woman is walking.Editing prompts are the opposite of generation prompts: short and surgical. Overly descriptive edit prompts cause unintended changes.
interaction_id in its result data.previous_interaction_id with operation="edit_video" and describe only the delta.Official good/bad pairs:
| Avoid | Instead |
|---|---|
| "In the video of the man sitting on the sofa, please add a small black cat..." | "Add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same." |
| "Please remove the cell phone... and fill in the background so it looks like..." | "Make the phone invisible. Keep everything else the same." |
Other working edit prompts: "Make this video anime" / "Put a fashionable hat on this person" / "Change the lighting to be more dramatic" / "Change the text on the sign to say 'Omni Flash'".
Gotcha — store: editing via previous_interaction_id only works if the prior call kept the interaction server-side (store defaults to true in gemini_omni_video). Set store=false only for one-shot generations you will never edit.
Editing uploaded videos: pass input_video_path instead of previous_interaction_id; the tool uploads it via the Files API. Unavailable in the EEA, Switzerland, and the UK (editing generated videos works everywhere).
16:9 or 9:16. All output carries an invisible SynthID watermark.Frequently asked questions
Gemini Omni is Google DeepMind's video generation and editing model family, announced at I/O 2026. The first model, Gemini Omni Flash (gemini-omni-flash-preview, developer access since June 30, 2026), generates 3-10 second clips at 720p/24fps with synthesized audio via the Gemin…
The source record exposes this install command: npx skills add https://github.com/calesthio/OpenMontage --skill ".agents/skills/gemini-omni". Inspect the command and pinned source before running it.