github triton-inference-server/server v2.72.0
Release 2.72.0 corresponding to NGC container 26.08

3 hours ago

Triton Inference Server

The Triton Inference Server provides a cloud inferencing solution optimized for both CPUs and GPUs. The server provides an inference service via an HTTP or GRPC endpoint, allowing remote clients to request inferencing for any model being managed by the server. For edge deployments, Triton Server is also available as a shared library with an API that allows the full functionality of the server to be included directly in an application.

New Features and Improvements

  • Added a log callback option to the tritonserver C and Python APIs and to the common Logger, allowing an embedding application to route Triton log records into its own logging stack instead of a file or stream.
  • Fixed dynamic-batcher starvation caused by waiting_consumer_count drift, where the scheduler could stop dispatching work despite queued requests.
  • Restored preserve_ordering semantics by reserving the completion-queue slot at enqueue time rather than at completion.
  • Core: Improved model-readiness reporting in TRITONSERVER_ServerModelIsReady — non-availability errors are now returned and unresolved models correctly report ready=false.
  • Model loading now catches exceptions rather than propagating them out of the load path, and uses logic_error so both invalid_argument and out_of_range are handled uniformly.
  • TensorRT backend: Fixed CUDA graph execution for non-batching models.
  • vLLM backend: Corrected JSON input parsing.
  • OpenVINO backend: Enabled OpenVINO model generation on SBSA (aarch64).
  • Python client: Removed the upper bound on the grpcio package requirement, so client installations are no longer pinned to an outdated gRPC.
  • Build: Added experimental build presets, allowing a build to be driven from a file via --build-presets-file instead of a long command line.
  • Build: Default build parallelism is now capped by available memory, with matching caps for the ONNX Runtime and OpenVINO backend builds, preventing out-of-memory failures during container builds.
  • Aligned protobuf and related dependencies with the gRPC 1.81.1 bump across the server, core, client, third-party and Model Analyzer components.
  • Documentation: Added a RHEL / manylinux build tutorial.

Known Issues

  • The Triton TensorRT-LLM Backend container image is not included in this release.
  • vLLM's v0 API and Ray are affected by vulnerabilities. Users should consider their own architecture and mitigation steps which may include but should not be limited to:
    • Do not expose Ray executors and vLLM hosts to a network where any untrusted connections might reach the host.
    • Ensure that only the other vLLM hosts are able to connect to the TCP port used for the XPUB socket. Note that the port used is random.
  • When using Valgrind or other leak detection tools on AGX-Thor or DGX-Spark systems, you might see memory leaks attributed to NvRmGpuLibOpen.
  • Valgrind or other memory leak detection tools may occasionally report leaks related to DCGM. These reports are intermittent and often disappear on retry.
  • CuPy has issues with the CUDA 13 Device API in multithreaded contexts. Avoid using tritonclient cuda_shared_memory APIs in multithreaded environments until fixed by CuPy.
  • TensorRT calibration cache may require size adjustment in some cases, which was observed for the IGX platform.
  • The core Python binding may incur an additional D2H and H2D copy if the backend and frontend both specify device memory to be used for response tensors.
  • A segmentation fault related to DCGM and NSCQ may be encountered during server shutdown on NVSwitch systems. A possible workaround for this issue is to disable the collection of GPU metrics tritonserver --allow-gpu-metrics false ....
  • When using TensorRT models, if auto-complete configuration is disabled and is_non_linear_format_io:true for reformat-free tensors is not provided in the model configuration, the model may not load successfully.
  • When using Python models in decoupled mode, users need to ensure that the ResponseSender goes out of scope or is properly cleaned up before unloading the model to guarantee that the unloading process executes correctly.
  • Triton Inference Server with vLLM backend currently does not support running vLLM models with tensor parallelism sizes greater than 1 and the default "distributed_executor_backend" setting when using explicit model control mode. In attempt to load a vllm model (tp > 1) in explicit mode, users could potentially see failure at initialize step: could not acquire lock for <_io.BufferedWriter name='<stdout>'> at interpreter shutdown, possibly due to daemon threads. For the default model control mode, after server shutdown, vllm related sub-processes are not killed. Related vllm issue: vllm-project/vllm#6766. Please specify "distributed_executor_backend":"ray" in the model.json when deploying vllm models with tensor parallelism > 1.
  • When loading models with file override, multiple model configuration files are not supported. Users must provide the model configuration by setting parameter "config" : "<JSON>" instead of custom configuration file in the following format: "file:configs/<model-config-name>.pbtxt" : "<base64-encoded-file-content>".
  • TensorRT-LLM backend provides limited support of Triton extensions and features.
  • The TensorRT-LLM backend may core dump on server shutdown. This impacts server teardown only and will not impact inferencing.
  • The Java CAPI is known to have intermittent segfaults.
  • Some systems which implement malloc() may not release memory back to the operating system right away causing a false memory leak. This can be mitigated by using a different malloc implementation. Tcmalloc and jemalloc are installed in the Triton container and can be used by specifying the library in LD_PRELOAD. NVIDIA recommends experimenting with both tcmalloc and jemalloc to determine which one works better for your use case.
  • Auto-complete may cause an increase in server start time. To avoid a start time increase, users can provide the full model configuration and launch the server with --disable-auto-complete-config.
  • Auto-complete does not support PyTorch models due to lack of metadata in the model. It can only verify that the number of inputs and the input names matches what is specified in the model configuration. There is no model metadata about the number of outputs and datatypes. Related PyTorch bug: pytorch/pytorch#38273
  • Triton Client PIP wheels for ARM SBSA are not available from PyPI and pip will install an incorrect Jetson version of Triton Client library for Arm SBSA. The correct client wheel file can be pulled directly from the Arm SBSA SDK image and manually installed.
  • Traced models in PyTorch seem to create overflows when int8 tensor values are transformed to int32 on the GPU. Refer to pytorch/pytorch#66930 for more information.
  • Triton cannot retrieve GPU metrics with MIG-enabled GPU devices.
  • Triton metrics might not work if the host machine is running a separate DCGM agent on bare-metal or in a container.

Client Libraries and Examples

The client libraries and examples are available in this release exclusively via the Ubuntu 24.04–based NGC Container. The SDK container includes the client libraries and examples, Performance Analyzer, and Model Analyzer. See Getting the Client Libraries for more information.

ManyLinux Assets (early access)

This release was compiled with AlmaLinux 8.9 based out of manylinux_2_34 and can be used on RHEL 9 and later versions.
See the included README.md for complete details about installation, verification, and support.
This release supports ensembles. Confirm CUDA, TensorRT, ONNX Runtime, PyTorch, and Python versions in the shipped README.md.
Some optional backend features such as the PyTorch backend's TorchTRT extension are not currently supported.

Don't miss a new server release

NewReleases is sending notifications on new releases.