huggingface/skills/skills/huggingface-spaces/SKILL.md
huggingface-spaces
Build, deploy, and maintain applications on Hugging Face Spaces — Gradio / Docker / Static SDKs, ZeroGPU and dedicated hardware, model loading, debugging, buckets, inference providers, community grants. Use whenever the user asks to create or host an app on Hugging Face, port code onto ZeroGPU, fix a Space that won't build or run, or otherwise work with `hf spaces …`, `@spaces.GPU`, Space README frontmatter, or the `spaces` Python package.
- Source repository stars
- 10,895
- Declared platforms
- 0
- Static risk flags
- 2
- Last source update
- 2026-08-04
- Source checked
- 2026-08-04
Decision brief
What it does—and where it fits
Hugging Face Spaces host machine-learning applications. There are 1M+ today; each Space is a git repo. This skill covers creating, building, debugging, and maintaining them.
Not for
- Tasks that require unconfirmed production actions or broad system permissions.
- Environments where the pinned source and install steps cannot be inspected.
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/huggingface/skills --skill "skills/huggingface-spaces"Inspect the Agent Skill "huggingface-spaces" from https://github.com/huggingface/skills/blob/32f8bb0928e95fc9d47ca9fbf69cbfbaf2bc2bda/skills/huggingface-spaces/SKILL.md at commit 32f8bb0928e95fc9d47ca9fbf69cbfbaf2bc2bda. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
0. Getting ready
1. Check the hf CLI is installed: which hf. If not, pip install -U huggingfacehub. 2. Check the user is logged in: hf auth whoami. If not, run hf auth login — it prints a URL and a one-time code; ask the user to open the URL and enter the code, then login completes automatically…
Check the hf CLI is installed: which hf. If not, pip install -U huggingfacehub.Check the user is logged in: hf auth whoami. If not, run hf auth login — it prints a URL and a one-time code; ask the user to open the URL and enter the code, then login completes automatically (OAuth, no token needed).…Note whoami's canPay and isPro flags — they gate hardware choices below. A free (isPro=False) account can only host Static Spaces and up to 2 ZeroGPU Spaces. - 02
1. What a Space is
A Space is a git repo with three possible SDKs:
Gradio — most Spaces. Python, fast iteration, supports ZeroGPU.Docker — arbitrary container. Use when you need a non-Python stack or a pre-built template (Streamlit, Argilla, Shiny, etc. — full list at https://huggingface.co/docs/hub/spaces-sdks-docker). Does not support ZeroGPU.Static — plain HTML, or a React/Svelte/Vue project built at deploy time. Use for in-browser ML (transformers.js / WebGPU / WebAssembly / onnxruntime-web), project pages, interactive reports, or Spaces that orchestrate o… - 03
Hardware tiers
Static Spaces are free for everyone and need no hardware. Gradio and Docker Spaces run on compute and require a paid plan to create — PRO for personal accounts, Team or Enterprise for organizations — with one exception: free personal accounts in good standing (verified email, ac…
Static Spaces are free for everyone and need no hardware. Gradio and Docker Spaces run on compute and require a paid plan to create — PRO for personal accounts, Team or Enterprise for organizations — with one exception:…So on a free account ZeroGPU is the only way to host a Gradio Space. cpu-basic is not the safe fallback it used to be — it is gated too.ZeroGPU (zero-a10g) — dynamic, per-request GPU allocation on NVIDIA RTX PRO 6000 Blackwell (sm120). Two sizes: large (half MIG, 48 GB, 1× quota) and xlarge (full, 96 GB, 2× quota). Free for the Space creator; Space visi… - 04
2. Look for an existing demo first
Before deciding how to build anything, search for prior art:
Before deciding how to build anything, search for prior art:If someone has built a similar Space, read its app.py and requirements.txt — that gives you the working pattern. Saves a lot of blind iteration. Mention to the user what you found before committing to an approach. - 05
3. Decide SDK and hardware
Follow the user's explicit request first. If they were vague:
Default for a public ML demo: Gradio + ZeroGPU. Use this unless something below applies.The model's only inference path is non-PyTorch (ONNX / TF / JAX / vLLM as the MAIN model, with heavy init): dedicated GPU.But: marginal non-torch tools (a small ONNX preprocessor, a TF utility) inside a torch-main pipeline are fine on ZeroGPU. The hijack only patches torch; init the non-torch lib inside @spaces.GPU and pay the short per-ca…
Permission review
Static risk signals and limitations
Writes files
The documentation asks the agent to create, modify, or delete local files.
Push files with `hf upload <namespace>/<name> . --repo-type space`. **`--repo-type space` is required** — `hf upload` defaults to a *model* repo and will otherwise upload to (and silently create) a model repo of the same name. Add `--excludReads files
The documentation asks the agent to read local files, directories, or repositories.
| When to read | File |Evidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 89/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 10,895 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- huggingface/skills
- Skill path
- skills/huggingface-spaces/SKILL.md
- Commit
- 32f8bb0928e95fc9d47ca9fbf69cbfbaf2bc2bda
- License
- Apache-2.0
- Collected
- 2026-08-04
- Default branch
- main
View the original SKILL.md
Hugging Face Spaces
Hugging Face Spaces host machine-learning applications. There are 1M+ today; each Space is a git repo. This skill covers creating, building, debugging, and maintaining them.
0. Getting ready
Before anything else:
- Check the
hfCLI is installed:which hf. If not,pip install -U huggingface_hub. - Check the user is logged in:
hf auth whoami. If not, runhf auth login— it prints a URL and a one-time code; ask the user to open the URL and enter the code, then login completes automatically (OAuth, no token needed). Alternatively, pass a write-scoped token from https://huggingface.co/settings/tokens with--token. - Note
whoami'scanPayandisProflags — they gate hardware choices below. A free (isPro=False) account can only host Static Spaces and up to 2 ZeroGPU Spaces.
The hf-cli skill teaches an agent every hf command and is the recommended companion to this one. Install it with hf skills add hf-cli (add --claude --global to install for Claude Code as well, user-level).
1. What a Space is
A Space is a git repo with three possible SDKs:
- Gradio — most Spaces. Python, fast iteration, supports ZeroGPU.
- Docker — arbitrary container. Use when you need a non-Python stack or a pre-built template (Streamlit, Argilla, Shiny, etc. — full list at https://huggingface.co/docs/hub/spaces-sdks-docker). Does not support ZeroGPU.
- Static — plain HTML, or a React/Svelte/Vue project built at deploy time. Use for in-browser ML (transformers.js / WebGPU / WebAssembly / onnxruntime-web), project pages, interactive reports, or Spaces that orchestrate other Spaces. No hardware needed.
Hardware tiers
Static Spaces are free for everyone and need no hardware. Gradio and Docker Spaces run on compute and require a paid plan to create — PRO for personal accounts, Team or Enterprise for organizations — with one exception: free personal accounts in good standing (verified email, account older than 30 days) can host up to 2 ZeroGPU Spaces.
So on a free account ZeroGPU is the only way to host a Gradio Space. cpu-basic is not the safe fallback it used to be — it is gated too.
ZeroGPU (zero-a10g) — dynamic, per-request GPU allocation on NVIDIA RTX PRO 6000 Blackwell (sm_120). Two sizes: large (half MIG, 48 GB, 1× quota) and xlarge (full, 96 GB, 2× quota). Free for the Space creator; Space visitors consume their own daily quota (~5 min free / 40 min Pro / 60 min Enterprise). Gradio-only, PyTorch-first. Hosting caps per account: 2 free personal, 10 PRO, 50 Team / Enterprise org.
cpu-basic — 2 vCPU / 16 GB, no hourly cost but needs a paid plan. For data viz, API-proxy Spaces, small CPU-bound models.
Dedicated GPU (T4, L4, A10G, L40S, A100, H200) — billed to the Space creator by the hour. List + pricing: hf spaces hardware. Only the creator can attach these, and only if canPay=True. Use when ZeroGPU genuinely doesn't fit — non-PyTorch main model with heavy init, very-large-model long-context inference, etc.
If the user needs hardware they can't pay for — a dedicated GPU, or a Gradio Space beyond the free 2-ZeroGPU cap — they can still create a Static Space (free for everyone), push the app there, and request a community grant. See references/grants.md.
For the authoritative reference: https://huggingface.co/docs/hub/spaces-overview
2. Look for an existing demo first
Before deciding how to build anything, search for prior art:
hf spaces search "<model name or task>" --sdk gradio --limit 10
If someone has built a similar Space, read its app.py and requirements.txt — that gives you the working pattern. Saves a lot of blind iteration. Mention to the user what you found before committing to an approach.
3. Decide SDK and hardware
Follow the user's explicit request first. If they were vague:
- Default for a public ML demo: Gradio + ZeroGPU. Use this unless something below applies.
- The model's only inference path is non-PyTorch (ONNX / TF / JAX / vLLM as the MAIN model, with heavy init): dedicated GPU.
- But: marginal non-torch tools (a small ONNX preprocessor, a TF utility) inside a torch-main pipeline are fine on ZeroGPU. The hijack only patches torch; init the non-torch lib inside
@spaces.GPUand pay the short per-call init cost.
- But: marginal non-torch tools (a small ONNX preprocessor, a TF utility) inside a torch-main pipeline are fine on ZeroGPU. The hijack only patches torch; init the non-torch lib inside
- Tiny / CPU-bound model, or API-proxy Space:
cpu-basic— but it needs a paid plan. On a free account, put it onzero-a10gwith a no-op decorated function (ZeroGPU requires at least one) and keep the real work outside it — nothing ever requests a GPU, so no quota is burned. Seereferences/inference-providers.md. - Browser-side ML or project page: Static.
- Container with non-Python stack: Docker.
Sourcing the model
- GitHub repo — clone locally to read structure. If it already has a Gradio demo, the minimal viable path is to adapt it onto ZeroGPU (see
references/zerogpu.md). Otherwise: read the README + inference code, prefer the PyTorch path, estimate VRAM (bf16 ≈params_B × 2GB; 48 GB fits ≤24B params at bf16, or much larger with quantization — seereferences/zerogpu.mdfor quantization on ZeroGPU). - HF model repo — read its README, follow any linked GitHub.
- Paper / blog post — look for an official or unofficial implementation. Don't reimplement unless trivial or the user explicitly asks.
- Vague request — search Spaces first; surface results.
If the model genuinely won't fit, check Inference Providers as an alternative: see references/inference-providers.md. This avoids hosting the model at all.
4. Create the Space
hf repos create <namespace>/<name> --type space --space-sdk <gradio|docker|static> \
[--flavor zero-a10g|cpu-basic|<paid-flavor>] \
[--secrets KEY=val] [--env KEY=val] \
--public|--private|--protected \
--exist-ok
--space-sdkis required.--flavorselects hardware.zero-a10gis the (legacy) identifier for ZeroGPU. Omitting it givescpu-basic— which is itself gated behind a paid plan, so on a free account pass--flavor zero-a10gexplicitly. Runhf spaces hardwarefor the full paid list and pricing.- Visibility:
--public(anyone can view),--private(only you),--protected(app is reachable but git repo / Files tab is private). --secrets KEY=valbecomes an environment variable inside the Space and is not visible to visitors. Use for API keys, gated-repo tokens (HF_TOKEN=hf_…), etc. Can also be set later viahf spaces secrets set <id> KEY=val.--env KEY=valis visible to visitors — use only for non-sensitive config (GRADIO_SSR_MODE=false,PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, etc.).
Note:
hardware:in the README YAML is silently ignored — hardware is only set via--flavorat creation, or later viahf spaces settings <id> --hardware <name>.
5. Build the app
The Space now exists at https://huggingface.co/spaces/<namespace>/<name> but is empty.
README.md frontmatter
Always required:
---
title: ...
emoji: 🚀 # pick something representative
colorFrom: blue # red|yellow|green|blue|indigo|purple|pink|gray (only these)
colorTo: indigo
sdk: gradio # gradio | docker | static
sdk_version: 6.15.1 # latest stable unless you have a reason*
app_file: app.py # gradio only (docker / static use Dockerfile / index.html)
short_description: ... # ≤ 60 chars (server rejects longer)
python_version: "3.12" # ZeroGPU officially supports 3.10.13 and 3.12.12
startup_duration_timeout: 30m # default; bump to 1h for big LLMs / heavy downloads
---
* Default to the current latest stable, and look up what that is (pip index versions gradio, or the version a freshly-created Space defaults to) — the number above is a placeholder that goes stale, don't reuse it. Only pin older when the latest genuinely doesn't work for this Space: a custom component pins it, or you're adapting an existing demo and don't want to rewrite for 5.x→6.x breaking changes. If you need a 5.x, pick 5.50.0 (latest of the series; still supports custom components).
All frontmatter options: https://huggingface.co/docs/hub/spaces-config-reference
Minimal ZeroGPU Gradio app
import spaces # MUST come before torch / diffusers / transformers
import torch
import gradio as gr
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained("<repo>", torch_dtype=torch.bfloat16).to("cuda")
@spaces.GPU(duration=60)
def generate(prompt: str):
"""Generate an image from a text prompt.""" # docstring → API / MCP tool description
return pipe(prompt).images[0]
gr.Interface(fn=generate, inputs=gr.Text(), outputs=gr.Image()).launch(mcp_server=True)
Three rules — full treatment in references/zerogpu.md:
import spacesbefore torch / any CUDA-touching import. It monkey-patchestorch.cuda.*; once CUDA is initialized in the main process, it's too late.- Load the model at module scope,
.to("cuda")eagerly. ZeroGPU intercepts the call, packs weights to disk, and streams them into VRAM on the first@spaces.GPUentry. Lazy loading inside the decorator costs every user. - Decorate the function Gradio binds. Estimate
durationto the realistic worst case (smaller = higher queue priority and tighter quota check). For input-dependent runtime, pass a callable.
Examples, docstrings, and MCP
- Add
gr.Exampleswhenever it makes sense (the app takes input and representative inputs exist) — prefer the model/repo's own official examples. Keep example rows to the few inputs a user actually varies (prompt, image) and give the handler defaults for the rest (steps, seed, guidance) so a row is["a prompt"], not a wall of knobs. Usecache_examples=True, cache_mode="lazy". Seereferences/gradio.md. - Give every API-triggered function a docstring and type hints. Each Gradio event handler is exposed over the API; the docstring + signature are what a caller — and the MCP tool schema — sees.
- Launch with
demo.launch(mcp_server=True)(Gradio 5+) so the Space doubles as an MCP server: each API function becomes an MCP tool described by its docstring and hints.
requirements.txt
Short version:
- Do NOT list:
gradio,spaces,huggingface_hub(preinstalled and platform-managed; pinning them causes resolution failures or silently breaks the ZeroGPU runtime). - Do list if you use them:
torchvision,torchaudio(not preinstalled), plus everything else (diffusers,transformers,accelerate,sentencepiece, …). - ZeroGPU only accepts torch
2.8.0,2.9.1,2.10.0,2.11.0. Default to leaving torch unpinned (the runtime preinstalls the latest). Only pin when a dep forces it. - For prebuilt CUDA-extension wheels (
flash_attn,xformers,pytorch3d,nvdiffrast,diff_gaussian_rasterization,torchmcubes): use the prebuilt Blackwell wheels athttps://huggingface.co/datasets/multimodalart/zerogpu-blackwell-wheels/tree/main/wheels. Full mapping + caveats inreferences/requirements.md.
Per-SDK depth
- Gradio patterns (themes,
gr.Examples, streaming, custom HTML components,gr.Server):references/gradio.md. - Docker: https://huggingface.co/docs/hub/spaces-sdks-docker. Examples:
hf spaces list --filter docker. - Static: https://huggingface.co/docs/hub/spaces-sdks-static. For built SPAs, set
app_build_command: npm run buildandapp_file: dist/index.htmlin frontmatter. - ZeroGPU specifics (decorator semantics, sizing, AoTI, generators, concurrency, pickle /
gr.Stateacross the worker boundary):references/zerogpu.md— read this whenever the Space targets ZeroGPU.
6. Iterate on the Space, not locally
Try to build a release candidate from the user quest locally and push it — then use the live URL as your test loop. The Space environment is the only one that matters; do not try to test locally. python3 -m py_compile app.py is the maximum local check worth doing before pushing.
Push files with hf upload <namespace>/<name> . --repo-type space. --repo-type space is required — hf upload defaults to a model repo and will otherwise upload to (and silently create) a model repo of the same name. Add --exclude "**/__pycache__/**" so local bytecode caches aren't committed into the Space.
Once pushed, pick the cheapest update mechanism for each change — hot-reload for pure Python edits, hf upload for code-only files hot-reload can't touch, full rebuild only when requirements.txt / Dockerfile / README frontmatter actually changed. Full ladder + footguns (hot-reload poisoning factory reboot, runtime.sha lag, etc.) in references/debugging.md.
7. Verify
Don't trust RUNNING alone — the app can be running but broken. Four steps, in order:
A. Alive? Stage + hardware:
hf spaces info <ns>/<name> --expand runtime
B. Logs clean post-boot? Read the run log to confirm startup finished without warnings or silent fallbacks:
hf spaces logs <ns>/<name> --tail 200
Look for model-load completion, no import warnings, no "falling back to CPU" / dtype downgrade messages, no RUNNING masking a half-broken app.
C. API actually responds. With logs still tailing in another terminal (hf spaces logs <ns>/<name> --follow), call the endpoint:
from gradio_client import Client, handle_file
import os
c = Client("<ns>/<name>", token=os.environ["HF_TOKEN"], httpx_kwargs={"timeout": 600})
print(c.view_api()) # discover endpoints — don't guess
result = c.predict(..., api_name="/generate")
D. Sniff output AND logs. HTTP 200 ≠ correct output. Check both:
head = open(result, "rb").read(16)
# glTF / \x89PNG / RIFF…WEBP / RIFF…WAVE / [4:8]==b"ftyp" → png/jpg/webp/wav/mp4
And look at the run log emitted during the call — silent fallbacks (model snapping to a different size, missing optional dep, dtype downgrade) only show up there.
Full smoke-test patterns (streaming endpoints, OAuth-gated Spaces, gr.Server custom routes): references/debugging.md.
8. Permanent storage (buckets)
Spaces are stateless — /data is wiped on restart. If the Space needs to persist user uploads, generations, logs, or interact with a long-lived store, mount a bucket:
hf buckets create <ns>/<bucket-name> # --private optional
hf spaces volumes set <ns>/<space> -v hf://buckets/<ns>/<bucket-name>:/data # read-write at /data
Buckets are paid storage; check canPay and confirm with the user. Full patterns (read-fast / write-durable, public bucket URLs, model-cache anti-pattern): references/buckets.md.
9. When things break
Order of operations:
- Read the logs:
hf spaces logs <id> --build --follow(build error) orhf spaces logs <id> --follow(runtime error). Find the first error, not the last. - Grep
references/known-errors.mdfor the error string. Check if this is a known issue before trying your own fix — most common ZeroGPU / Gradio / dependency errors have a 1–2 line fix there. - Iterate using the cheapest rung from
references/debugging.md. The vast majority of issues resolve with log-reading + smoke-test loops; interactive dev mode + SSH is a heavy-hammer last resort.
If you solve an error that wasn't in the known-errors list, suggest the user PR it back to this skill so future runs benefit.
Reference index
| When to read | File |
|---|---|
| How ZeroGPU works + correct patterns (decorator, sizing, pickle, generators, real-time, AoTI) | references/zerogpu.md |
| Iterate + debug: logs, rung ladder, smoke testing (and dev mode + SSH as a last resort) | references/debugging.md |
| Error-string lookup — the single place for all error symptoms (Spaces, ZeroGPU, Gradio, deps) | references/known-errors.md |
| Pinning deps, picking wheels, torch-family alignment | references/requirements.md |
gr.Examples (add when it makes sense), themes, custom HTML components, gr.Server, MCP server (mcp_server=True) | references/gradio.md |
| Persistent storage, public bucket URLs | references/buckets.md |
| Community grant requests (hardware the user can't pay for) | references/grants.md |
| Provider proxy (zero-VRAM big LLM via Cerebras / Fireworks / Together / etc.) | references/inference-providers.md |
| 3D Spaces: generation, CUDA extensions, output formats, and model recipes (incl. gaussian splatting) | references/3d-generation.md |
Alternatives
Compare before choosing
wshobson/agents
brand-landingpage
Brand-first landing page designer — runs a brand-identity interview (colors, typography, shape language), then generates and iterates on a polished landing page via Stitch with deployment-ready HTML. Use when the user asks to create, design, or build a landing page, homepage, or marketing page and has no established visual direction. Skip when they have a design mockup, need a dashboard or app UI, are working at component level, building a multi-page app, or restyling with known design tokens —
wanshuiyin/Auto-claude-code-research-in-sleep
experiment-bridge
Use it for deployment and documentation tasks; the detail page covers purpose, installation, and practical steps.
ruvnet/ruflo
github-release-management
Comprehensive GitHub release orchestration with AI swarm coordination for automated versioning, testing, deployment, and rollback management
Galaxy-Dawn/claude-scholar
git-workflow
This skill should be used when the user asks to "create git commit", "manage branches", "follow git workflow", "use Conventional Commits", "handle merge conflicts", or asks about git branching strategies, version control best practices, pull request workflows. Provides comprehensive Git workflow guidance for team collaboration.