📊 Jobs: retries, rescheduling, and port exposure
Hugging Face Jobs get a batch of quality-of-life improvements this release. Jobs can now retry automatically on failure with the new attempts parameter, and you can start a fresh Job from an existing Job's saved spec (including secrets and hardware settings) with rerun_job / hf jobs rerun. Scheduled Jobs can be rescheduled — change the cron expression or preset without recreating the Job — via update_scheduled_job_schedule / hf jobs scheduled reschedule. Finally, you can now expose ports on Jobs: use expose for token-protected access and expose_public for unauthenticated access, both at creation time and on a running Job with update_job_expose / hf jobs expose.
# Retry a job up to 2 times on failure, and rerun it later from its saved spec
>>> hf jobs run --attempts 3 --detach python:3.12 python train.py
>>> hf jobs rerun <job_id>
# Reschedule a scheduled job to run every day at 9:00
>>> hf jobs scheduled reschedule <scheduled_job_id> "0 9 * * 1"
# Expose ports on a running Job (8000 token-protected, 9000 public)
>>> hf jobs expose <job_id> 8000 --public 9000- [Jobs] Add rerun and retry attempts by @Wauplin in #5025
- [Jobs] Allow rescheduling scheduled jobs by @Wauplin in #5026
- [Jobs] Manage public and private exposed ports by @Wauplin in #5027
📚 Documentation: Jobs guide
🤗 Track your ZeroGPU quota
The Hub now exposes a ZeroGPU quota endpoint, and huggingface_hub wraps it with a new HfApi.get_zero_gpu_quota() method returning a ZeroGpuQuota dataclass (values in GPU-seconds). This is particularly useful for apps, MCP servers and agents built on top of ZeroGPU Spaces: check how much quota is left and when it resets, and warn your users before they hit the daily limit. The same information is available from the CLI with hf spaces zero-gpu-quota, which also prints a hint to purchase credits when the quota is running low.
>>> from huggingface_hub import get_zero_gpu_quota
>>> quota = get_zero_gpu_quota()
ZeroGpuQuota(base=2400, remaining=1810, resets_at=datetime.datetime(2026, 9, 30, 9, 12, 3, tzinfo=datetime.timezone.utc), overquota_used=0)>>> hf spaces zero-gpu quota
✓ ZeroGPU quota (in GPU-seconds)
remaining: 1810
base: 2400
resets_at: 2026-09-30T18:12:03+00:00
overquota_used: 0- [Spaces] Add get_zero_gpu_quota and hf spaces zero-gpu-quota by @Wauplin in #5040
- [CLI] Rename 'hf spaces zero-gpu-quota' to 'hf spaces zero-gpu quota'- #5055 [Shipped as part of v2.1.1]
📚 Documentation: Manage your Space guide
⚡ HfFileSystem.get_file downloads Xet files with hf_xet
HfFileSystem.get_file now downloads Xet-backed files through hf_xet (the same code path as hf_hub_download) when writing to a local path, letting hf_xet write directly to disk instead of going through Python. It reuses the Xet hash already returned by info() (part of the tree listing), so no extra HTTP call is needed. This mostly benefits non-streaming datasets loads from hf://buckets/... paths: benchmarks on HF Jobs show get_file dropping from ~120s to ~7s on a 2 GB parquet file, and load_dataset("hf://buckets/...") from ~130s to ~17s. The classic HTTP path is still used for non-Xet files, when hf_xet is not installed, or when writing to a file-like object. As a side effect, a missing remote file no longer leaves an empty local file behind.
- [HfFileSystem] Download Xet files with hf_xet in get_file by @hanouticelina in #4990
🧠 Inference Endpoints: Model Catalog migrates to the new v1 API
The catalog methods now use the new /api/v1/catalog API. list_inference_catalog returns InferenceCatalogModel dataclasses with their tested deployment recipes (hardware + engine combinations), and takes server-side filters (accelerator, engine, license, task, search, limit). create_inference_endpoint_from_catalog can deploy an exact recipe with the new recipe_id and gguf_file arguments. The CLI follows suit: hf endpoints catalog ls prints one row per recipe and supports the filters, while hf endpoints catalog deploy gains --recipe and --gguf-file.
>>> hf endpoints catalog ls --engine vllm --task text-generation --search llama --limit 10
REPO_ID TASK LICENSE ACCELERATOR ENGINE GGUF_FILE RECIPE_ID
----------------------------------- ------------------ ---------- ----------- -------- ----------------------------- -------------------------
meta-llama/Llama-3.1-8B-Instruct text-generation Llama 3.1 gpu vllm sizzling-biryani-g4xsi1ac
bartowski/QwQ-32B-Preview-GGUF text-generation Apache 2.0 gpu llamacpp QwQ-32B-Preview-Q8_0.gguf baked-orange-m863gx7d📚 Documentation: Inference Endpoints guide
🔒 Sandboxes require scoped credentials
Sandbox security is tightened in this release. Pooled sandboxes with missing, empty or malformed capability tokens are now refused instead of silently substituting the host-management credential, and the client is pinned to sbx-server 0.7.0, which removes host-token acceptance on per-sandbox routes entirely. Older pool hosts that do not provide scoped tokens must be upgraded/recycled: new clients cannot reach legacy hosts and legacy clients cannot reach new sandboxes (documented, accepted breaking change). The sandbox concepts and guides have been updated to describe the strict behavior, including how to safely use proxy_headers with external HTTP clients (disable redirects!).
- [Sandbox] Require scoped credentials for pooled sandboxes by @Wauplin in #4782
- [Sandbox] Bump sbx-server to 0.7.0 by @Wauplin in #5028
💔 Breaking Change
Note
The Inference Endpoints catalog migration (#4928) and the Sandbox credential changes (#4782, #5028) are breaking changes detailed in their highlight sections above.
🖥️ CLI
- [CLI] Escape control characters in listings and table output by @Wauplin in #5029
- [CLI] Suggest
hf auth loginwhenhf auth whoamifails by @moon-bot-app[bot] in #5033 - [CLI] Align
hf auth whoaminot-logged-in message withhf auth tokenby @Wauplin in #5034 - [CLI] Raise if a bare --secrets NAME is not set locally in hf jobs by @Wauplin in #5044
- [CLI] Hint to use --global after local 'hf skills add/update' by @Wauplin in #5024
- [CLI] nit: update
hf uploadhelp text by @davanstrien in #5046
🔧 Other QoL Improvements
- [HfApi] Support resource_group_id in duplicate_repo by @Wauplin in #5038
- [Repos] Warn when duplicated repo files are still being copied by @Wauplin in #5041
- Support secrets on Job-triggered webhooks by @moon-bot-app[bot] in #5009
- [HfApi] Add
visibilitytocreate_bucketby @hanouticelina in #5053
🐛 Bug and typo fixes
- [Buckets] Fix copy_files treating lexical siblings as an existing destination directory by @AnishPatel526 in #5001
- [HTTP] Fix repo_id in errors for query strings and download URLs by @Wauplin in #5008
- Fix scalar tags splitting into single-character list in card front matter by @zeenat28-ui in #5017
- [datasets sql] Run the DuckDB CLI subprocess with UTF-8 encoding by @825pranav in #5020
- [Auth] Store token names as UTF-8 instead of locale encoding by @Evolian-o in #5021
- [Dataclasses] Accept bare typing.List/Dict/Set in @strict validation by @anishmehta24 in #5032
- [Utils] Retry the first pagination page as well by @Wauplin in #5035
📖 Documentation
- [docs] xet caching and concurrency updates by @cbensimon in #4883
- [Spaces] Update docs for cpu-basic gating and dev mode by @Wauplin in #5037
- [docs] Fix hf.cos typo in repository guide link by @pratikgx in #5016
- [Docs] Fix links in i18n issue template by @pratikgx in #5051
- [Docs] Fix duplicated words in docstring and comments by @Ad1th in #5052