Best for
- Use when the user wants to train a WAN or Z-Image LoRA; covers local + RunPod setup, dataset prep, key params, and using the result in a ComfyUI workflow.
artokun/comfyui-mcp/plugin/skills/ai-toolkit-trainer/SKILL.md
Train custom LoRAs with ostris AI-Toolkit — covers WAN 2.2/2.1 (people, styles, video motion) and Z-Image (Turbo & Base, low-VRAM image LoRAs). Use when the user wants to train a WAN or Z-Image LoRA; covers local + RunPod setup, dataset prep, key params, and using the result in a ComfyUI workflow.
Decision brief
Train custom LoRAs with ostris AI-Toolkit — covers WAN 2. 2/2.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/artokun/comfyui-mcp --skill "plugin/skills/ai-toolkit-trainer"Inspect the Agent Skill "ai-toolkit-trainer" from https://github.com/artokun/comfyui-mcp/blob/0852abe2c68d9fe9e2af89c54cd039357f08ae6c/plugin/skills/ai-toolkit-trainer/SKILL.md at commit 0852abe2c68d9fe9e2af89c54cd039357f08ae6c. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
The installer comes in two generations — both clone ostris/ai-toolkit, set up Torch for your GPU, and launch the web UI. Put it in a folder whose full path has NO spaces (e.g. C:\AI-Toolkit).
Installs into the persistent volume /workspace/ai-toolkit; idempotent (re-run just relaunches the UI). Use RunPod's PyTorch 2.8.0 template, 100GB disk. It installs apt deps, clones the repo, makes a venv, installs Torch (torchaudio included), installs nvm + Node 22, then builds/…
In the UI: create a Job, point it at a dataset folder, pick the model (WAN variant or Z-Image), set params, start. Jobs run in the Python backend, so you can close the browser. (Or bypass the UI: copy a config/examples/.yml, edit, python run.py config/.yml.)
AI-Toolkit pairs each sample with a same-basename .txt caption and auto-resizes/buckets aspect ratios (no pre-cropping).
Captions: natural-language; include a unique trigger word for a person/character.
Permission review
The documentation asks the agent to create, modify, or delete local files.
In the UI: create a **Job**, point it at a dataset folder, pick the model (WAN variant or Z-Image), set params, start. Jobs run in the Python backend, so you can close the browser. (Or bypass the UI: copy a `config/examples/*.yml`, edit, `pThe documentation asks the agent to read local files, directories, or repositories.
Param tables are aggregated starting points (community/training-guide sources), **not** read from the repo's `config/examples/*.yml` — open the actual WAN / Z-Image example config in your clone and tune. See "Unverified".The documentation asks the agent to read local files, directories, or repositories.
The **param tables** (both WAN and Z-Image) are synthesized starting points, **not** read from the repo's `config/examples/*.yml`. Open the actual example config in your clone and adjust.Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 88/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 485 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
AI-Toolkit by ostris is "the ultimate training toolkit for finetuning diffusion models" (MIT license) — a standalone trainer with its own web UI, NOT a ComfyUI custom node. It runs a Node.js UI front end over a Python (run.py) training backend, and trains LoRAs for many model families — here we cover WAN 2.2 / 2.1 video models and Z-Image (Turbo & Base).
https://github.com/ostris/ai-toolkit (cloned by the installers).python run.py config/<job>.yml. UI: a Node.js app under ui/ that schedules/monitors jobs (you don't have to keep the UI open while a job runs)..safetensors LoRA you drop into ComfyUI models/loras/ and load with LoraLoaderModelOnly.Best for:
For low-VRAM anime image LoRAs on a different stack (kohya sd-scripts), see the sibling anima-lora-trainer.
Two LoRA kinds (WAN): a WAN image LoRA trains on still images (cheaper, ~24GB-class, good for identity/style); a WAN video LoRA trains on short clips (heavier — best on cloud — good for motion). Z-Image is image-only.
The installer comes in two generations — both clone ostris/ai-toolkit, set up Torch for your GPU, and launch the web UI. Put it in a folder whose full path has NO spaces (e.g. C:\AI-Toolkit).
AI-TOOLKIT_AUTO_INSTALL.bat: expects Git, Python 3.10.x, and Node 18+ already in PATH.AI-TOOLKIT_AUTO_INSTALL-V2.bat (recommended): uses an embedded Python 3.10.11, auto-installs Git + Node, builds a clean PATH without your system Python, and adds aggressive pip/curl retries — far fewer prerequisites and the more robust choice. (Used for the Z-Image Turbo LoRA training release.)Both are CUDA-aware and select the Torch wheel by GPU generation:
| Choice | GPU | CUDA | Torch index | Torch packages |
|---|---|---|---|---|
| 1 | RTX 50-series (Blackwell) | 12.8 | https://download.pytorch.org/whl/cu128 | torch==2.7.0 torchvision==0.22.0 |
| 2 | RTX 40 / 30 / 20 and older | 12.6 | https://download.pytorch.org/whl/cu126 | torch==2.7.0 torchvision==0.22.0 |
Each then: clones ostris/ai-toolkit; downloads two launcher scripts (LAUNCHER-TOOLKIT.bat, SECURE_LAUNCHER-TOOLKIT.bat, from https://huggingface.co/Aitrepreneur/FLX/resolve/main/); makes the venv; installs Torch from the chosen index; pip install -r requirements.txt; then cd ui && npm run build_and_start.
AI-TOOLKIT_AUTO_INSTALL-RUNPOD.sh (and -V2.sh)Installs into the persistent volume /workspace/ai-toolkit; idempotent (re-run just relaunches the UI). Use RunPod's PyTorch 2.8.0 template, 100GB disk. It installs apt deps, clones the repo, makes a venv, installs Torch (torchaudio included), installs nvm + Node 22, then builds/starts the UI.
| Choice | GPU | Stream | Torch spec |
|---|---|---|---|
| 1 | RTX 5000-series (Blackwell) | cu128 | torch==2.7.0+cu128 torchvision==0.22.0+cu128 torchaudio==2.7.0+cu128 |
| 2 | Ada / Hopper / Ampere, older | cu126 | torch==2.7.0 torchvision==0.22.0 torchaudio==2.7.0 |
Ports & auth: UI on 8675, Jupyter on 8888. Set AI_TOOLKIT_AUTH (UI password) before launch. Reach it at https://${RUNPOD_POD_ID}-8675.proxy.runpod.net. GPU recs: RTX 4090/5090 for image (WAN t2i/t2v, Z-Image) LoRAs; RTX 6000 Pro (Blackwell) for heavy WAN video / high-res / high-rank jobs.
LAUNCHER-TOOLKIT.bat (local) or SECURE_LAUNCHER-TOOLKIT.bat (password-protected) from the ai-toolkit folder..sh — it detects the install and starts the UI instantly on :8675.In the UI: create a Job, point it at a dataset folder, pick the model (WAN variant or Z-Image), set params, start. Jobs run in the Python backend, so you can close the browser. (Or bypass the UI: copy a config/examples/*.yml, edit, python run.py config/<job>.yml.)
AI-Toolkit pairs each sample with a same-basename .txt caption and auto-resizes/buckets aspect ratios (no pre-cropping).
my_dataset/
001.png 001.txt
002.jpg 002.txt
Short clips + a .txt per clip; caption the motion/camera move. Per-clip frames via the job's num_frames (e.g. 81). Markedly heavier — prefer cloud GPUs.
WAN 2.2 14B is a Mixture-of-Experts with a high-noise expert (structure/motion) and a low-noise expert (detail). AI-Toolkit trains both via Multi-stage.
| Param | Default | Notes |
|---|---|---|
| Linear rank / dim | 16 | 16 simple; 16–32 complex/cinematic |
| Learning rate | 5e-5 (identity) | 7e-5–1e-4 style; high LR → plasticky skin |
| Steps | 1500–2500 | stop before overbaking |
| Resolution | 512 (or 768) | bucketed; 768 costs more VRAM |
num_frames (video) | 81 | per-clip frame count |
| Multi-stage | High + Low = ON | trains both experts |
| Switch Every | 10 | raise to 20–50 if offload swapping is slow |
| Optimizer / Quant | AdamW8bit / 4-bit ARA or float8 | fits 14B on consumer cards |
Z-Image is a ~6B single-stream model — no hi/lo multi-stage (leave Multi-stage OFF; you train one model). It's the lightest target here: the headline of the Z-Image releases is training on very low VRAM.
| Param | Starting point | Notes |
|---|---|---|
| Linear rank / dim | 16–32 | 32 for detailed characters/styles |
| Learning rate | 1e-4 | lower (5e-5) for tighter identity |
| Steps | 1500–3000 | dataset-dependent |
| Resolution | 768 (or 1024) | Z-Image's native range |
| Multi-stage | OFF | single-stream model, not WAN's MoE |
| Optimizer / Quant | AdamW8bit / float8 | enables sub-12GB training |
Train on Base, deploy anywhere. Z-Image Base is the finetuning-friendly model; a LoRA trained on Base generally applies to the Turbo workflow too. Use the z-image-xy-plot pack to grid-compare your trained LoRAs.
Param tables are aggregated starting points (community/training-guide sources), not read from the repo's
config/examples/*.yml— open the actual WAN / Z-Image example config in your clone and tune. See "Unverified".
<your_lora>.safetensors into ComfyUI models/loras/.LoraLoaderModelOnly:
LoraLoaderModelOnly on the Z-Image model path (see the z-image-base / z-image-turbo packs). Strength 0.7–1.0.{ "class_type": "LoraLoaderModelOnly",
"inputs": { "model": ["<base_model>", 0],
"lora_name": "<your_lora>.safetensors",
"strength_model": 1.0 } }
No module named 'torchaudio' when starting a job (AI-Toolkit). The venv's Torch stack is mismatched. Fix: activate the AI-Toolkit venv (venv\Scripts\activate), then pip uninstall torch torchaudio torchvision -y and pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 (or your CUDA's index). Only affects the AI-Toolkit install, not ComfyUI.self and mat2 must have the same dtype (ComfyUI-WanVideoWrapper, WAN usage). Re-clone ComfyUI-WanVideoWrapper in custom_nodes/ and reinstall its requirements.txt, then restart ComfyUI.pip install onnxruntime==1.20.1 in the affected venv.AI_TOOLKIT_AUTH is set and you're on the 8675 proxy URL.config/examples/*.yml. Open the actual example config in your clone and adjust..bat files are downloaded from a third-party HuggingFace repo (Aitrepreneur/FLX); review before running on a security-sensitive machine.Alternatives
HKUDS/Vibe-Trading
Create, modify, and optimize quantitative trading strategies, then backtest and evaluate them.
alirezarezvani/claude-skills
ISO 13485 Quality Management System implementation and maintenance for medical device organizations. Provides QMS design, documentation control, internal auditing, CAPA management, and certification support. Use when working with medical device quality systems, preparing for ISO 13485 audits, managing regulatory compliance documentation, setting up corrective actions, or building audit preparation programs. Useful for quality management, audit preparation, regulatory compliance, medical device d
MoizIbnYousaf/marketing-cli
Use when the user wants to generate an image or video via Higgsfield AI. Covers 30+ models: Soul V2, Seedance 2.0, Kling 3.0, Veo 3.1, GPT Image 2, Nano Banana 2. Also covers Marketing Studio — branded ad video/image with avatars and products. Use whenever: "generate an image", "make a video", "animate this photo", "image-to-video", "img2vid", "edit this image with AI", "produce a clip", "create an ad", "make a UGC video", "marketing video", "brand video", "TV spot", "import product from URL", "
davepoon/buildwithclaude
Automate Figma tasks via Rube MCP (Composio): files, components, design tokens, comments, exports. Always search tools first for current schemas.