github xorbitsai/inference v4.0.0

3 hours ago

What's new in 4.0.0 (2026-10-09)

These are the changes in inference v4.0.0.

New features

  • feat(image): add jina-ocr-v1 OCR support by @Minamiyama in #5588
  • feat(api): add stateless OpenAI Responses API (/v1/responses) by @bluefish-08 in #5590
  • feat(tts): support Confucius4-TTS by @Minamiyama in #5591
  • feat(logs): add runtime log center by @Minamiyama in #5592
  • feat(auth): require globally unique API key names by @m199369309 in #5597
  • feat: add native NIXL backend for vLLM PD deployments by @qinxuye in #5604
  • feat: correlate audit and runtime logs by request ID by @m199369309 in #5610
  • feat: add GPU-first tiered Xavier snapshot storage by @qinxuye in #5615
  • feat: enable GPU-first Xavier caching with NIXL transfers by @qinxuye in #5616
  • feat(webui): unify runtime and historical log views by @m199369309 in #5635
  • FEAT: support SGLang Xavier KV sharing and native GPU PD by @qinxuye in #5631
  • feat: retain GPU weights across vLLM and SGLang engine reloads by @qinxuye in #5640
  • feat: add MLX Xavier KV cache sharing and PD separation by @qinxuye in #5638
  • FEAT: support Xavier P/D across vLLM and SGLang by @qinxuye in #5639
  • FEAT: support bidirectional Xavier P/D between NVIDIA and MLX by @qinxuye in #5641
  • feat(llm): support Ling-3.0-tiny with vLLM and SGLang by @Minamiyama in #5645
  • feat(embedding): add EmbeddingGemma 2 support by @Minamiyama in #5644
  • FEAT: report cache hit tokens across vLLM, SGLang and MLX by @qinxuye in #5647
  • feat: add cross-platform installation and local service management by @OliverBryant in #5178
  • FEAT: Add distribution extension points for API composition and worker admission by @qinxuye in #5650
  • feat: add configurable network interface MAC address lookup by @llyycchhee in #5654

Enhancements

  • ENH: Support Qwen3.5 recurrent-state PD handoff by @qinxuye in #5630
  • ENH: Complete vLLM V1 pipeline parallel execution with xoscar by @qinxuye in #5637
  • perf: batch Xavier KV reads and add transfer diagnostics by @qinxuye in #5606
  • perf: batch Xavier KV reads across layers by @qinxuye in #5614
  • perf: preserve BF16 bits in Xavier KV transfers by @qinxuye in #5613
  • perf: avoid gathering already published Xavier GPU snapshots by @qinxuye in #5620
  • perf: reduce Xavier GPU transfer padding with persistent small slabs by @qinxuye in #5621
  • perf: batch Xavier KV loads across scheduled requests by @qinxuye in #5622
  • PERF: overlap Xavier KV loads with vLLM decoding by @qinxuye in #5623
  • PERF: reduce Xavier GPU snapshot eviction overhead by @qinxuye in #5624
  • PERF: avoid repeated Xavier snapshot size scans by @qinxuye in #5625
  • perf: reduce Xavier cold-prefill snapshot overhead by @qinxuye in #5627
  • perf: default Xavier PD to GPU handoff with tiered KV history by @qinxuye in #5628
  • PERF: reduce Xavier V1 cache overhead and allow CUDA Graph by @qinxuye in #5643
  • REF: extract shared Xavier core and Torch backend by @qinxuye in #5629

Bug fixes

Documentation

  • DOC: update new models for v3.5.0 by @qinxuye in #5587
  • doc: add v3.5.0 to release notes index and locale catalogs by @martinma51 in #5594
  • doc: clarify Xinference Enterprise, Cloud, and Model API offerings by @qinxuye in #5605
  • doc: discuss Docker Swarm deployment considerations by @qiulang in #5617
  • doc: simplify README quick start by @qinxuye in #5652
  • doc: refresh community installation and 4.0 release guides by @qinxuye in #5653

Others

  • TEST: Cover multi-prefill and multi-decode GPU routing by @qinxuye in #5632

New Contributors

Full Changelog: v3.5.0...v4.0.0

Don't miss a new inference release

NewReleases is sending notifications on new releases.