pypi huggingface-hub 2.1.0
[v2.1.0] Better Jobs tooling, ZeroGPU quota tracking, and revamped Model Catalog

latest release: 2.1.1
5 hours ago

📊 Jobs: retries, rescheduling, and port exposure

Hugging Face Jobs get a batch of quality-of-life improvements this release. Jobs can now retry automatically on failure with the new attempts parameter, and you can start a fresh Job from an existing Job's saved spec (including secrets and hardware settings) with rerun_job / hf jobs rerun. Scheduled Jobs can be rescheduled — change the cron expression or preset without recreating the Job — via update_scheduled_job_schedule / hf jobs scheduled reschedule. Finally, you can now expose ports on Jobs: use expose for token-protected access and expose_public for unauthenticated access, both at creation time and on a running Job with update_job_expose / hf jobs expose.

# Retry a job up to 2 times on failure, and rerun it later from its saved spec
>>> hf jobs run --attempts 3 --detach python:3.12 python train.py
>>> hf jobs rerun <job_id>

# Reschedule a scheduled job to run every day at 9:00
>>> hf jobs scheduled reschedule <scheduled_job_id> "0 9 * * 1"

# Expose ports on a running Job (8000 token-protected, 9000 public)
>>> hf jobs expose <job_id> 8000 --public 9000

📚 Documentation: Jobs guide

🤗 Track your ZeroGPU quota

The Hub now exposes a ZeroGPU quota endpoint, and huggingface_hub wraps it with a new HfApi.get_zero_gpu_quota() method returning a ZeroGpuQuota dataclass (values in GPU-seconds). This is particularly useful for apps, MCP servers and agents built on top of ZeroGPU Spaces: check how much quota is left and when it resets, and warn your users before they hit the daily limit. The same information is available from the CLI with hf spaces zero-gpu-quota, which also prints a hint to purchase credits when the quota is running low.

>>> from huggingface_hub import get_zero_gpu_quota
>>> quota = get_zero_gpu_quota()
ZeroGpuQuota(base=2400, remaining=1810, resets_at=datetime.datetime(2026, 9, 30, 9, 12, 3, tzinfo=datetime.timezone.utc), overquota_used=0)
>>> hf spaces zero-gpu quota
✓ ZeroGPU quota (in GPU-seconds)
  remaining: 1810
  base: 2400
  resets_at: 2026-09-30T18:12:03+00:00
  overquota_used: 0
  • [Spaces] Add get_zero_gpu_quota and hf spaces zero-gpu-quota by @Wauplin in #5040
  • [CLI] Rename 'hf spaces zero-gpu-quota' to 'hf spaces zero-gpu quota'- #5055 [Shipped as part of v2.1.1]

📚 Documentation: Manage your Space guide

⚡ HfFileSystem.get_file downloads Xet files with hf_xet

HfFileSystem.get_file now downloads Xet-backed files through hf_xet (the same code path as hf_hub_download) when writing to a local path, letting hf_xet write directly to disk instead of going through Python. It reuses the Xet hash already returned by info() (part of the tree listing), so no extra HTTP call is needed. This mostly benefits non-streaming datasets loads from hf://buckets/... paths: benchmarks on HF Jobs show get_file dropping from ~120s to ~7s on a 2 GB parquet file, and load_dataset("hf://buckets/...") from ~130s to ~17s. The classic HTTP path is still used for non-Xet files, when hf_xet is not installed, or when writing to a file-like object. As a side effect, a missing remote file no longer leaves an empty local file behind.

🧠 Inference Endpoints: Model Catalog migrates to the new v1 API

The catalog methods now use the new /api/v1/catalog API. list_inference_catalog returns InferenceCatalogModel dataclasses with their tested deployment recipes (hardware + engine combinations), and takes server-side filters (accelerator, engine, license, task, search, limit). create_inference_endpoint_from_catalog can deploy an exact recipe with the new recipe_id and gguf_file arguments. The CLI follows suit: hf endpoints catalog ls prints one row per recipe and supports the filters, while hf endpoints catalog deploy gains --recipe and --gguf-file.

>>> hf endpoints catalog ls --engine vllm --task text-generation --search llama --limit 10
REPO_ID                             TASK               LICENSE    ACCELERATOR ENGINE   GGUF_FILE                     RECIPE_ID
----------------------------------- ------------------ ---------- ----------- -------- ----------------------------- -------------------------
meta-llama/Llama-3.1-8B-Instruct    text-generation    Llama 3.1  gpu         vllm                                   sizzling-biryani-g4xsi1ac
bartowski/QwQ-32B-Preview-GGUF      text-generation    Apache 2.0 gpu         llamacpp QwQ-32B-Preview-Q8_0.gguf     baked-orange-m863gx7d
  • [Endpoints] Migrate catalog to the new /api/v1 API by @Wauplin in #4928

📚 Documentation: Inference Endpoints guide

🔒 Sandboxes require scoped credentials

Sandbox security is tightened in this release. Pooled sandboxes with missing, empty or malformed capability tokens are now refused instead of silently substituting the host-management credential, and the client is pinned to sbx-server 0.7.0, which removes host-token acceptance on per-sandbox routes entirely. Older pool hosts that do not provide scoped tokens must be upgraded/recycled: new clients cannot reach legacy hosts and legacy clients cannot reach new sandboxes (documented, accepted breaking change). The sandbox concepts and guides have been updated to describe the strict behavior, including how to safely use proxy_headers with external HTTP clients (disable redirects!).

  • [Sandbox] Require scoped credentials for pooled sandboxes by @Wauplin in #4782
  • [Sandbox] Bump sbx-server to 0.7.0 by @Wauplin in #5028

💔 Breaking Change

  • [Buckets] Make include/exclude patterns case-sensitive on all platforms by @Wauplin in #5007

Note

The Inference Endpoints catalog migration (#4928) and the Sandbox credential changes (#4782, #5028) are breaking changes detailed in their highlight sections above.

🖥️ CLI

  • [CLI] Escape control characters in listings and table output by @Wauplin in #5029
  • [CLI] Suggest hf auth login when hf auth whoami fails by @moon-bot-app[bot] in #5033
  • [CLI] Align hf auth whoami not-logged-in message with hf auth token by @Wauplin in #5034
  • [CLI] Raise if a bare --secrets NAME is not set locally in hf jobs by @Wauplin in #5044
  • [CLI] Hint to use --global after local 'hf skills add/update' by @Wauplin in #5024
  • [CLI] nit: update hf upload help text by @davanstrien in #5046

🔧 Other QoL Improvements

  • [HfApi] Support resource_group_id in duplicate_repo by @Wauplin in #5038
  • [Repos] Warn when duplicated repo files are still being copied by @Wauplin in #5041
  • Support secrets on Job-triggered webhooks by @moon-bot-app[bot] in #5009
  • [HfApi] Add visibility to create_bucket by @hanouticelina in #5053

🐛 Bug and typo fixes

  • [Buckets] Fix copy_files treating lexical siblings as an existing destination directory by @AnishPatel526 in #5001
  • [HTTP] Fix repo_id in errors for query strings and download URLs by @Wauplin in #5008
  • Fix scalar tags splitting into single-character list in card front matter by @zeenat28-ui in #5017
  • [datasets sql] Run the DuckDB CLI subprocess with UTF-8 encoding by @825pranav in #5020
  • [Auth] Store token names as UTF-8 instead of locale encoding by @Evolian-o in #5021
  • [Dataclasses] Accept bare typing.List/Dict/Set in @strict validation by @anishmehta24 in #5032
  • [Utils] Retry the first pagination page as well by @Wauplin in #5035

📖 Documentation

  • [docs] xet caching and concurrency updates by @cbensimon in #4883
  • [Spaces] Update docs for cpu-basic gating and dev mode by @Wauplin in #5037
  • [docs] Fix hf.cos typo in repository guide link by @pratikgx in #5016
  • [Docs] Fix links in i18n issue template by @pratikgx in #5051
  • [Docs] Fix duplicated words in docstring and comments by @Ad1th in #5052

🏗️ Internal

  • Post-release: bump version to 2.1.0.dev0 by @huggingface-hub-bot[bot] in #4999
  • [CI] Bump doc-builder workflows pin to 63ec6f1 by @Wauplin in #5010
  • [Tests] Reduce redundant Hub calls in repo settings and move tests by @Wauplin in #5036

Don't miss a new huggingface-hub release

NewReleases is sending notifications on new releases.