What's new in 4.0.0 (2026-10-09)
These are the changes in inference v4.0.0.
New features
- feat(image): add jina-ocr-v1 OCR support by @Minamiyama in #5588
- feat(api): add stateless OpenAI Responses API (/v1/responses) by @bluefish-08 in #5590
- feat(tts): support Confucius4-TTS by @Minamiyama in #5591
- feat(logs): add runtime log center by @Minamiyama in #5592
- feat(auth): require globally unique API key names by @m199369309 in #5597
- feat: add native NIXL backend for vLLM PD deployments by @qinxuye in #5604
- feat: correlate audit and runtime logs by request ID by @m199369309 in #5610
- feat: add GPU-first tiered Xavier snapshot storage by @qinxuye in #5615
- feat: enable GPU-first Xavier caching with NIXL transfers by @qinxuye in #5616
- feat(webui): unify runtime and historical log views by @m199369309 in #5635
- FEAT: support SGLang Xavier KV sharing and native GPU PD by @qinxuye in #5631
- feat: retain GPU weights across vLLM and SGLang engine reloads by @qinxuye in #5640
- feat: add MLX Xavier KV cache sharing and PD separation by @qinxuye in #5638
- FEAT: support Xavier P/D across vLLM and SGLang by @qinxuye in #5639
- FEAT: support bidirectional Xavier P/D between NVIDIA and MLX by @qinxuye in #5641
- feat(llm): support Ling-3.0-tiny with vLLM and SGLang by @Minamiyama in #5645
- feat(embedding): add EmbeddingGemma 2 support by @Minamiyama in #5644
- FEAT: report cache hit tokens across vLLM, SGLang and MLX by @qinxuye in #5647
- feat: add cross-platform installation and local service management by @OliverBryant in #5178
- FEAT: Add distribution extension points for API composition and worker admission by @qinxuye in #5650
- feat: add configurable network interface MAC address lookup by @llyycchhee in #5654
Enhancements
- ENH: Support Qwen3.5 recurrent-state PD handoff by @qinxuye in #5630
- ENH: Complete vLLM V1 pipeline parallel execution with xoscar by @qinxuye in #5637
- perf: batch Xavier KV reads and add transfer diagnostics by @qinxuye in #5606
- perf: batch Xavier KV reads across layers by @qinxuye in #5614
- perf: preserve BF16 bits in Xavier KV transfers by @qinxuye in #5613
- perf: avoid gathering already published Xavier GPU snapshots by @qinxuye in #5620
- perf: reduce Xavier GPU transfer padding with persistent small slabs by @qinxuye in #5621
- perf: batch Xavier KV loads across scheduled requests by @qinxuye in #5622
- PERF: overlap Xavier KV loads with vLLM decoding by @qinxuye in #5623
- PERF: reduce Xavier GPU snapshot eviction overhead by @qinxuye in #5624
- PERF: avoid repeated Xavier snapshot size scans by @qinxuye in #5625
- perf: reduce Xavier cold-prefill snapshot overhead by @qinxuye in #5627
- perf: default Xavier PD to GPU handoff with tiered KV history by @qinxuye in #5628
- PERF: reduce Xavier V1 cache overhead and allow CUDA Graph by @qinxuye in #5643
- REF: extract shared Xavier core and Torch backend by @qinxuye in #5629
Bug fixes
- fix: address recurring Windows and Metal CI failures by @qinxuye in #5593
- fix(llm): keep text outside Qwen tool call tags as content by @MohammadHijjawi97 in #5595
- fix(core): name the model in the load-retry warning by @David-Wu1119 in #5596
- fix(ui): restore GPU selection for recommended model types by @m199369309 in #5598
- fix(monitoring): map audit request metadata in Elasticsearch by @m199369309 in #5601
- fix: constrain missing host TorchCodec to compatible Torch versions by @qinxuye in #5599
- fix: load parent editable installs in model virtual environments by @qinxuye in #5600
- fix(llm): keep parallel DeepSeek V3.1 tool calls with different arguments by @MohammadHijjawi97 in #5602
- fix(llm): return every DeepSeek-R1 tool call in a tool calls block by @MohammadHijjawi97 in #5603
- fix: use public SentenceTransformer device property by @m199369309 in #5608
- fix: align synchronous RPC structured logging by @m199369309 in #5612
- fix: reuse Xavier actor RPC loop across EngineCore connectors by @qinxuye in #5607
- fix: retry Router Agent bootstrap operations by @m199369309 in #5609
- fix: preserve launch history device semantics by @m199369309 in #5611
- fix(llm): preserve DeepSeek V3.2 tool argument types by @SulimanAbdulrazzaq in #5618
- fix: release request limit when metrics RPC is cancelled by @tuanzirwar in #5626
- fix: validate Xavier KV compatibility before accepting remote hits by @qinxuye in #5619
- fix(llm): keep finish_reason and empty content for tool responses without calls by @MohammadHijjawi97 in #5633
- fix(audit): skip successful metrics scrape events by @m199369309 in #5636
- fix: resolve SGLang JIT cache paths before runtime startup by @qinxuye in #5642
- BUG: make log rotation safe across Windows processes by @Ricardo-M-L in #5303
- fix(api): honour encoding_format=base64 in /v1/embeddings by @MohammadHijjawi97 in #5646
- fix(frontend): remove flags from language selector by @maoyuehui in #5648
- fix(llm): stream text held back as a possible tool call tag by @hpdkhoa in #5651
Documentation
- DOC: update new models for v3.5.0 by @qinxuye in #5587
- doc: add v3.5.0 to release notes index and locale catalogs by @martinma51 in #5594
- doc: clarify Xinference Enterprise, Cloud, and Model API offerings by @qinxuye in #5605
- doc: discuss Docker Swarm deployment considerations by @qiulang in #5617
- doc: simplify README quick start by @qinxuye in #5652
- doc: refresh community installation and 4.0 release guides by @qinxuye in #5653
Others
New Contributors
- @MohammadHijjawi97 made their first contribution in #5595
- @SulimanAbdulrazzaq made their first contribution in #5618
- @tuanzirwar made their first contribution in #5626
- @hpdkhoa made their first contribution in #5651
Full Changelog: v3.5.0...v4.0.0