Triton Inference Server
The Triton Inference Server provides a cloud inferencing solution optimized for both CPUs and GPUs. The server provides an inference service via an HTTP or GRPC endpoint, allowing remote clients to request inferencing for any model being managed by the server. For edge deployments, Triton Server is also available as a shared library with an API that allows the full functionality of the server to be included directly in an application.
The client libraries and examples are available in this release exclusively via the Ubuntu 24.04–based NGC Container. The SDK container includes the client libraries and examples, Performance Analyzer, and Model Analyzer. See Getting the Client Libraries for more information.
This release was compiled with AlmaLinux 8.9 based out of New Features and Improvements
developer_tools C API wrapper as deprecated in the documentation.
cudaIpcMemHandle_t bytes via getPtr() for compatibility with cuda-bindings 13.x.
preserve_ordering behavior by reserving the completion-queue slot at enqueue time.
default_model_filename autofill incorrectly applying to Python-runtime PyTorch models.
libOpenCL.so.1.
ReportStatistics timestamp arguments and added a guard to the AOTI package loader in the PT2/AOTI backend.
libtritonserver from auditwheel processing in the tritonfrontend wheel build.
UNAVAILABLE error instead of a NULL-dereference crash.
tritonfrontend wheel to hatchling with PEP 639 metadata.
Known Issues
tritonserver --allow-gpu-metrics false ....
is_non_linear_format_io:true for reformat-free tensors is not provided in the model configuration, the model may not load successfully.
ResponseSender goes out of scope or is properly cleaned up before unloading the model to guarantee that the unloading process executes correctly.
initialize step: could not acquire lock for <_io.BufferedWriter name='<stdout>'> at interpreter shutdown, possibly due to daemon threads. For the default model control mode, after server shutdown, vllm related sub-processes are not killed. Related vllm issue: vllm-project/vllm#6766. Please specify "distributed_executor_backend":"ray" in the model.json when deploying vllm models with tensor parallelism > 1.
"config" : "<JSON>" instead of custom configuration file in the following format: "file:configs/<model-config-name>.pbtxt" : "<base64-encoded-file-content>".
--disable-auto-complete-config.
Client Libraries and Examples
ManyLinux Assets (early access)
manylinux_2_34 and can be used on RHEL 9 and later versions.
See the included README.md for complete details about installation, verification, and support.
This release supports ensembles. Confirm CUDA, TensorRT, ONNX Runtime, PyTorch, and Python versions in the shipped README.md.
Some optional backend features such as the PyTorch backend's TorchTRT extension are not currently supported.