Performance Optimizations
Intel 64/AMD64 Processors
- Introduced initial support for AI Compute Extensions (ACE) instructions. This functionality is not dispatched by default and requires opt-in with environment variable
ONEDNN_MAX_CPU_ISA=AVX10_2_ACE. - Improved performance of future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).
- Improved performance of
fp8matmul with block-wise weights on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids). - Improved performance of
bf16andf16matmul with relaxed accumulation mode on Intel Xeon processors with Intel AMX instruction set support. - Improved performance of strided deconvolution on Intel Xeon Scalable processors (formerly Sapphire Rapids).
- Improved performance of floating-point Scaled Dot Product Attention (SDPA) training forward propagation subgraph with Graph API.
Intel Graphics
- Improved performance of future discrete GPUs based on Xe3p-XPC architecture (codename Crescent Island).
- Improved performance of future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
- Improved performance of Scaled Dot Product Attention (SDPA) training forward and backpropagation subgraphs.
- Improved performance of Scaled Dot Product Attention (SDPA) subgraphs with asymmetric head sizes.
- Improved performance of grouped matmul with small group sizes.
- Improved performance of
f8matmul with source and weights scales mask3. - Improved performance of
f16ands8matmul withu4weights andN=1. - Reduced convolution and deconvolution primitives creation time.
AArch64 Processors
- Improved the performance of
u8ands8matmul on platforms with SVE support. - Improved the performance of
f32matmul on platforms with 128-bit SVE vector lengths. - Improved the performance of
f32depthwise convolution on platforms with ASIMD support. - Improved the performance of convolutions.
- Improved the performance of
eltwise_logon platforms with ASIMD support. - Improved the performance of
bf16eltwise for thegelu_erf,swish,gelu_tanh,exp,log, andsqrtalgorithms. - Improved the performance of
f32, andf16PReLU. - Improved the performance of eltwise post-ops.
- Improved the performance of the
logsoftmaxalgorithm for the softmax primitive on platforms with ASIMD support. - Reduced penalties on small utility functions on clang builds by changing the default stack-protection level from
alltostrong.
RISC-V Processors
- Improved performance of
f32binary, eltwise, pooling, softmax, and logsoftmax on processors withVextension support. - Improved performance of
f16matmul, eltwise, and softmax on processors withZvfhextension support, including softmax with non-contiguous axes. - Extended RVV-optimized implementations to
bf16binary, eltwise, pooling, softmax, batch normalization, and resampling, and improvedbf16matmul performance on processors withZvfbfwmaextension support. - Introduced RVV-optimized forward resampling with nearest-neighbor and linear interpolation for
f32andf16data types. - Introduced RVV-optimized shuffle for
f32,s32,f16, andbf16data types. - Extended the RVV matmul implementation to support unsigned 8-bit source and weights.
Functionality
Functional API
- Introduced the
binary_mul_inplacealgorithm for binary post-ops. Unlikebinary_mul, the new algorithm allows the matmul destination tensor to be used as one of its inputs. An optimized implementation is available for matmul on Intel GPUs. - [experimental] Extended eltwise post-ops support in grouped matmul with all supported algorithms. Optimized implementation is available on Intel GPUs.
- [experimental] Extended grouped matmul with support for backpropagation cases (2D grouped by 3D dense and 2D grouped by 2D grouped) covering
f32,f16andbf16data types. Optimized implementation is available for Intel GPUs. This is an experimental feature that requires opt-in withONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ONbuild option.
Graph API
- Extended DynamicQuantize and DynamicDequantize operation to support the new
maskattribute.
Usability
Common
- Updated
mxfp8downconversion implementations to saturate instead of overflowing. New behavior is consistent with OCP MX specification and aligned with preferred behavior in PyTorch. - Version number can now be used as a passable approximation of pi.
Intel 64/AMD64 processors
- Cleaned up implicit narrowing conversions and removed suppression of MSVC compiler warning C4244.
Intel Graphics
- Refactored verbose profiling implementation for Level Zero runtime to avoid spurious synchronizations.
- Introduced support for concurrent primitive execution with the Level Zero runtime on Intel GPUs.
- [experimental] Introduced support for verbose profiling based on sycl_ext_oneapi_profiling_tag SYCL extension. This is an experimental feature that requires opt-in with
ONEDNN_EXPERIMENTAL_ENABLE_SYCL_PROFILING_TAG=ONbuild option.
AArch64 Processors
- Introduced initial asynchronous runtime support to AArch64 platforms for the matmul, convolution, eltwise, binary, lnorm, and reorder primitives.
- Fixed a memory leak in convolutions on platforms with SVE support.
Validation
- Updated benchdnn
smokeandCItest sets for matmul using parameter space sampling approach. - [experimental] Extended benchdnn
--groupedknob withbalanced,hot, anddecodestrategies for offset generation to generate MoE-style group distributions in grouped matmul validation. - Extended benchdnn graph driver: operation attribute removal via
--op-attrsknob, scalar tensor support via--in-shapesknob, tensor property rewriting via--tensor-propertyknob, operation removal via--op-kind.
Deprecated Functionality
- BLAS-like API including
dnnl::sgemm,dnnl::gemm_u8s8s32, anddnnl::gemm_s8s8s32functions is deprecated and will be removed in future releases. If you are using this API consider switching to matmul primitive.
Breaking changes
- Removed optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th generation Intel Core processors. oneDNN remains functional on these platforms and dispatches a generic OpenCL implementation.
- Removed optimizations for processors with Intel SSE4.1 and Intel AVX instruction sets. oneDNN remains functional on these platforms and dispatches a generic C++ implementation.
- Removed optimizations for
tf32fpmath_modein matmul on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).
Thanks to our Contributors
This release contains contributions from the project core team as well as Abhishek Kumar @abhishek-iitmadras, Aditya Singh @adityasingh2400, Akihiro Tabuchi @Akihiro-Tabuchi, AragornOfKebroyd @AragornOfKebroyd, Aron Xu @happyaron, @AyushSinghBaiswar, Codrut Irimie @CodrutIrimieARM, Crefeda Rodrigues @cfRod, elimor01 @MorelElian, Emilio Cota @cota, Ishita Shreya @ishita-shreya, Kamil Jackiewicz @kjackiew, Kamil Wieloch @kwieloch-intel, Keerthana KT @Keerthana-64, Léandre LE DUC @leduclean, Leon Kennedy @leoken01, Megha Sangtani @megha-sangtani, Mohammed Bilgrami @mohbil01, Nikhil Gupta @nikhil-arm, PiotrReiterIntel @PiotrReiterIntel, Puneet Matharu @puneetmatharu, @rinatrap, Thiago Macieira @thiagomacieira, @Tiwari-Avanish, Udit Kumar Agarwal @uditagarwal97, @velonica0, Wang hongyan @ww8191201-coder, and @xinghai-zh.
Each of them contributed an invaluable slice of the pi.