sonichi/sutando

image-generation

Generate and edit images using Gemini Flash Image, and generate videos using Veo. Supports text-to-image, image editing, text-to-video, and image-to-video.

72CollectingRuns scripts
See how to use itView GitHub source
npx skills add https://github.com/sonichi/sutando --skill "skills/image-generation"
Automated source guide

Source checked Jul 28, 2026·Refresh due Oct 26, 2026

Reorganized from the pinned upstream SKILL.md

Turn image-generation's source instructions into a guide you can follow

According to the pinned SKILL.md from sonichi/sutando: Generate images and videos using Gemini APIs.

npx skills add https://github.com/sonichi/sutando --skill "skills/image-generation"
Check the pinned source

Best fit

  • "Generate a hero image for my project"
  • "Create a short video of a sunset timelapse"
  • "Edit this photo to remove the background"

Bring this context

  • google-genai Python package (pip3 install google-genai)
  • GEMINIAPIKEY in .env or environment
  • Pillow (pip3 install Pillow) — for image generation/editing

Expected outputs

  • python3 "$SKILLDIR/scripts/generate.py" --prompt "A cute robot mascot" --output mascot.png
  • google-genai Python package (pip3 install google-genai)
  • GEMINIAPIKEY in .env or environment

Key source sections

Read image-generation through these 5 source sections

Sections are extracted automatically from the pinned SKILL.md and link back to the source.

01

Usage

Review the “Usage” section in the pinned source before continuing.

SKILL.md · Usage
Review and apply the “Usage” source section.
02

Image Generation (Gemini Flash Image)

Text-to-image: Generate images from text descriptions

SKILL.md · Image Generation (Gemini Flash Image)
Text-to-image: Generate images from text descriptionsImage editing: Modify existing images with natural languageBackground replacement: Change or enhance backgrounds
03

Video Generation (Veo)

Text-to-video: Generate video clips from text prompts

SKILL.md · Video Generation (Veo)
Text-to-video: Generate video clips from text promptsImage-to-video: Animate a reference image with a prompt- Text-to-video: Generate video clips from text prompts - Image-to-video: Animate a reference image with a prompt
04

When to Use

"Generate a hero image for my project"

SKILL.md · When to Use
"Generate a hero image for my project""Create a short video of a sunset timelapse""Edit this photo to remove the background"
05

Text-to-image

python3 "$SKILLDIR/scripts/generate.py" --prompt "A futuristic city skyline at night"

SKILL.md · Text-to-image
python3 "$SKILLDIR/scripts/generate.py" --prompt "A futuristic city skyline at night"

SkillSignal prompt templates

Provide the task, context, and acceptance criteria

These prompts were written by SkillSignal from the source structure; they are not upstream text.

Task-start prompt

Confirm source fit, inputs, and outputs before acting.

Use image-generation to help me with: [specific task]. Context: [files, data, or background]. Constraints: [environment, scope, and prohibited actions]. Before acting, check the pinned SKILL.md and explain which sections apply, what inputs are still missing, and what you will deliver.

Source-guided execution

Make the Agent explicitly follow the key extracted sections.

Apply the pinned image-generation source to [task]. Pay particular attention to these source sections: “Usage”, “Image Generation (Gemini Flash Image)”, “Video Generation (Veo)”, “When to Use”, “Text-to-image”. Preserve the important decision at each step. Mark facts not covered by the source as “needs confirmation” instead of inventing them. Then verify the result against my acceptance criteria: [criteria].

Result-review prompt

Check omissions, permissions, and source drift before delivery.

Review the current image-generation result: (1) does it satisfy the original task; (2) were any applicable steps or limits in the pinned SKILL.md missed; (3) did it perform any unauthorized file, command, network, or data action; and (4) which conclusions remain unverified? List issues first, then fix only what the source or user authorization supports.

Output checklist

Verify each item before delivery

The task matches the purpose documented in the SKILL.md.

The source section “Usage” has been checked.

The source section “Image Generation (Gemini Flash Image)” has been checked.

The source section “Video Generation (Veo)” has been checked.

The source section “When to Use” has been checked.

Inputs, constraints, and acceptance criteria are explicit.

Unverified facts, compatibility, and outcome claims are clearly marked.

Any file, command, network, or data action has been reviewed.

Choose a different workflow

When another Skill is the better fit

image-generation

Generate an image from a brief — provider-agnostic blueprint then provider-specific translation, with ref-image/seed reuse for consistency. Use when generating/creating an image.

A separate implementation from event4u-app/agent-config; compare its source, maintenance signals, and permission requirements.

Open source detail

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

A separate implementation from coreyhaines31/marketingskills; compare its source, maintenance signals, and permission requirements.

Open source detail

churn-prevention

When the user wants to reduce churn, build cancellation flows, set up save offers, recover failed payments, or implement retention strategies. Also use when the user mentions 'churn,' 'cancel flow,' 'offboarding,' 'save offer,' 'dunning,' 'failed payment recovery,' 'win-back,' 'retention,' 'exit survey,' 'pause subscription,' 'involuntary churn,' 'people keep canceling,' 'churn rate is too high,' 'how do I keep users,' or 'customers are leaving.' Use this whenever someone is losing subscribers o

A separate implementation from coreyhaines31/marketingskills; compare its source, maintenance signals, and permission requirements.

Open source detail

FAQ

What does image-generation do?

Generate images and videos using Gemini APIs.

How do I start using image-generation?

The catalog detected this source-specific install command: npx skills add https://github.com/sonichi/sutando --skill "skills/image-generation". Inspect the command and pinned source before running it.

Which Agent platforms does it declare?

No dedicated Agent platform is declared in the pinned source record.

Repository stars
359
Repository forks
79
Quality
72/100
Source repository last pushed

Quality breakdown

Based on traceable docs and repository signals; stars are not treated as quality.

72/100
Documentation23/30
Specificity21/25
Maintenance18/20
Trust signals10/25

Compare before choosing

Related Agent Skills and source variants

These links are selected from shared tasks, functions, stacks, platforms, and same-name variants. Compare the source owner, documentation, permissions, and maintenance signals.

image-generation by event4u-app

Generate an image from a brief — provider-agnostic blueprint then provider-specific translation, with ref-image/seed reuse for consistency. Use when generating/creating an image.

ab-testing by coreyhaines31

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

churn-prevention by coreyhaines31

When the user wants to reduce churn, build cancellation flows, set up save offers, recover failed payments, or implement retention strategies. Also use when the user mentions 'churn,' 'cancel flow,' 'offboarding,' 'save offer,' 'dunning,' 'failed payment recovery,' 'win-back,' 'retention,' 'exit survey,' 'pause subscription,' 'involuntary churn,' 'people keep canceling,' 'churn rate is too high,' 'how do I keep users,' or 'customers are leaving.' Use this whenever someone is losing subscribers o

design-intelligence by event4u-app

Grounded design brief from the adopted corpus — style, WCAG-checked color tokens, typography, layout pattern, anti-patterns. Use on ui-design-brief or any which-style/palette/font/chart decision.

design-system-capture by event4u-app

Write and maintain DESIGN.md + PRODUCT.md — captures visual decisions and interaction patterns so design tasks stay consistent across sessions without re-scanning past work.

View original Skill.mdThis page is parsed directly from the repository SKILL.md without editorial rewriting. Collected: Jul 28, 2026 · about 2 min

Media Generation

Generate images and videos using Gemini APIs.

Image Generation (Gemini Flash Image)

  • Text-to-image: Generate images from text descriptions
  • Image editing: Modify existing images with natural language
  • Background replacement: Change or enhance backgrounds
  • Hero/banner creation: Create branded images with text overlays
  • Style transfer: Apply artistic styles to photos

Video Generation (Veo)

  • Text-to-video: Generate video clips from text prompts
  • Image-to-video: Animate a reference image with a prompt

When to Use

  • "Generate a hero image for my project"
  • "Create a short video of a sunset timelapse"
  • "Edit this photo to remove the background"
  • "Make a video from this image"
  • "Generate a logo with a dark theme"

Usage

# Text-to-image
python3 "$SKILL_DIR/scripts/generate.py" --prompt "A futuristic city skyline at night"

# Edit an existing image
python3 "$SKILL_DIR/scripts/generate.py" --input photo.jpg --prompt "Add dramatic clouds"

# Specify output path
python3 "$SKILL_DIR/scripts/generate.py" --prompt "A cute robot mascot" --output mascot.png

# Text-to-video
python3 "$SKILL_DIR/scripts/generate.py" --video --prompt "A timelapse of a city at sunset"

# Video with portrait aspect ratio
python3 "$SKILL_DIR/scripts/generate.py" --video --prompt "Ocean waves" --aspect 9:16

# Image-to-video (animate a reference image)
python3 "$SKILL_DIR/scripts/generate.py" --video --input scene.jpg --prompt "Animate this scene with gentle wind"

# Specify output
python3 "$SKILL_DIR/scripts/generate.py" --video --prompt "Dancing robot" --output robot.mp4

Options

FlagDescriptionDefault
--promptText prompt describing what to generate(required)
--inputInput image path(s) for editing/referenceNone
--outputOutput file pathgenerated-{timestamp}.png or .mp4
--modelGemini model to usegemini-2.5-flash-image / veo-3.1-generate-preview
--videoGenerate video instead of imagefalse
--aspectVideo aspect ratio16:9
--qualityJPEG quality (1-100, images only)90

Requirements

  • google-genai Python package (pip3 install google-genai)
  • GEMINI_API_KEY in .env or environment
  • Pillow (pip3 install Pillow) — for image generation/editing

Notes

  • Video generation takes 1-3 minutes (polling every 10s)
  • Generated videos are stored on Google servers for 2 days
  • Gemini may refuse some prompts (people's faces, copyrighted characters, etc.)
  • For image editing, be explicit: "keep the subject unchanged, only modify the background"
  • Image output format inferred from extension (.jpg, .png, .webp)
  • Maximum input image size: ~20MB
Source repo
sonichi/sutando
Skill path
skills/image-generation/SKILL.md
Commit SHA
6a8f0fccd32e
Repository license
MIT
Data collected