ONNX Runtime 1.31.0 stabilizes model-package and EPContext data APIs, improves CPU model loading and quantized MoE inference, expands device-based execution-provider selection, and strengthens model-loading and runtime reliability. These notes cover changes since ONNX Runtime 1.30.0. CUDA and WebGPU kernel updates are summarized here; detailed provider notes are covered by their separate plugin EP releases.
Highlights
- Promoted the model-package API and EPContext data callbacks to stable C and C++ APIs, with EPContext callback support added across language bindings (#33166, #32265).
- Reduced CPU session-initialization overhead by parallelizing eligible weight prepacking, and accelerated block-wise INT4/INT8 QMoE experts with MLAS QNBit kernels and grouped expert dispatch (#31691, #32644, #32668).
- Added x86 FP16 LayerNorm/RMSNorm acceleration, Arm KleidiAI SVE2.1 FP16 GEMM support, and additional RISC-V vector kernels (#32715, #32670, #32710).
- Enabled CoreML participation in device-based EP selection, including Apple Neural Engine selection through
PREFER_NPU(#31975). - Added workspace-memory accounting and verification for constrained-memory graph partitioning, alongside broader model, graph, and tensor validation (#31962, #32189).
Announcements & Compatibility
- CUDA Plugin EP v0.3.0 and WebGPU Plugin EP v0.5.0 are matched with ONNX Runtime 1.31.0. For contrib-op compatibility, core and the plugin must use the same contributed-operator schemas; building both from the same ONNX Runtime revision is recommended. Schema mismatches can cause incorrect execution or crashes and may not be detected during plugin registration.
GatherBlockQuantizedbehavior change: out-of-range indices now produce zero output slices across CPU, CUDA, and WebGPU for both integer and floating-point quantized formats. This deliberately differs from ONNXGathererror semantics (#32480).- Experimental API migration: model-package consumers should use
OrtApi::GetModelPackageApiand the stableOrtModelPackageApitable. EPContext callback registration now uses stable APIs; the corresponding experimental lookup entries were removed. The stable native callback contract carries callback/state pointers, with payload limits enforced by application callbacks, bindings, or EPs rather than a native read-options policy (#33166, #32265). - ACL EP is deprecated and will be removed in a future release. Source builds using
--use_aclnow receive a deprecation warning (#32564). - The ONNX dependency remains at 1.22.0. The initial ONNX 1.23 integration was reverted pending upstream fixes; this release does not ship that integration's new opset or FLOAT6 support (#33073).
- Supported native builds now default to 1DS telemetry. Windows TraceLogging remains selectable with
--use_windows_telemetryoronnxruntime_USE_WINDOWS_TELEMETRY=ON; telemetry can be disabled with--no_telemetry(#32384).
New Features
Core APIs & Runtime
- Stabilized the model-package C and C++ APIs for loading packages, selecting compatible model variants, and creating sessions (#33166).
- Stabilized named EPContext data read/write callbacks for application-managed storage. Added EP support negotiation so registered callbacks do not silently fall back to filesystem I/O (#32265).
- Extended the Compile API to write external initializers and EP context data to buffers, enabling in-memory compilation workflows for models larger than 2 GB (#31347).
- Added workspace-memory accounting, source reporting, and post-partition reservation verification. Set
session.strict_workspace_verification=1to fail initialization on workspace declaration overruns or nonzero orphaned reservations; strict verification is not supported for ORT-format model loads (#31962, #32189). - Added the opt-in source-build option
onnxruntime_DISABLE_DEVICE_DISCOVERY=ONto skip platform accelerator probing and report only the CPU device (#32592).
Shared Contrib Operators
These updates affect core operator schemas and cross-provider behavior, not only the separately released GPU plugins.
- Extended
EngramGateandNGramHashMappingfor Qwen4-Exp, and addedVarlenNGramHashMappingfor packed DeepSeek Engram batching with request-local history. Packed hashing also supports Qwen4-Exp head offsets, EOS resets, and segment boundaries (#32285, #32358, #32467). - Added the
activationattribute toGatedRMSNormfor SiLU, its Swish alias, and Sigmoid gating across CPU, CUDA, and WebGPU. SiLU remains the default (#32512). - Added weightless hyper-connection operators
BranchwiseRMSNorm,ScaledSiLU,HyperConnectionPreMix, andHyperConnectionPostMix, with CPU, CUDA, and WebGPU implementations (#32687). - Extended
GatherBlockQuantizedto FP8/FP4 data, full-axis blocks, and scale broadcasting, and improved CPU UINT8 dequantization. See the out-of-range-index behavior change above (#32480, #32621).
Plugin Execution Provider Infrastructure
- Added support for statically linked plugin EPs, with a static WebGPU plugin build configuration (#32395).
- Preserved caller streams in single-tensor plugin data transfers and handled null plugin allocator returns (#32666, #32683).
- Deduplicated AutoEP custom-op domains and corrected EP assignment reporting after NHWC layout transformation (#32711, #33184).
- Added reference guidance for compiled-model compatibility validation, including multi-device scenarios (#29168).
Execution Provider Updates
CUDA EP
CUDA updates include sparse attention and packed continuous batching, expanded paged speculative decoding, INT2 and mixed-width QMoE execution, opt-in MatMul auto-tuning, improved MXFP4 support on Blackwell consumer GPUs, and reliability fixes.
See the CUDA Plugin EP v0.3.0 release notes for detailed features, configuration options, limitations, and fixes. Most CUDA-specific PRs are covered there rather than repeated here. Use the matched core/plugin combination and observe the contrib-op schema compatibility requirements above.
WebGPU EP
- Expanded generative AI support with sparse attention, packed continuous-batching indexers, GatedDeltaNet, and broader quantized MoE support, including block-scaled FP8 and INT2/mixed-width weights (#32528, #32529, #32533, #32618, #32510, #32794, #32878).
- Improved native provider integration and performance, including explicit session-stream concurrency, subgroup-matrix kernels on AMD GPUs, and standalone Dawn API package support (#32587, #32667, #33006).
Detailed WebGPU kernel additions, optimizations, and correctness fixes will be covered in the separate WebGPU Plugin EP v0.5.0 release notes.
CoreML EP
- Exposed CoreML through
OrtEpFactory, device enumeration,SessionOptionsAppendExecutionProvider_V2, and EP selection policies. NPU selection requires iOS 16/macOS 13 or later; GPU selection requires iOS 15/macOS 12 or later. In builds advertising both CoreML and WebGPU for the same Apple GPU,PREFER_GPU/MAX_PERFORMANCEselect CoreML through the existing EP-name tie-breaker; WebGPU remains available through explicit device selection (#31975).
WebNN EP
- Added
GroupNormalization,GroupNorm, andSkipGroupNormsupport (#32606). - Supported NHWC
Resizeby inferringresample2daxes and fixedConvTransposeshape-arithmetic overflow (#32628, #32661).
Other Execution Providers
- Fixed a CANN deadlock during parallel execution and concurrent XNNPACK allocator initialization (#29740, #32489).
- Validated DirectML upload-heap allocations and corrected
ResizeROI axis expansion (#32745, #32747).
CPU & Core Optimizations
Model Loading, MLAS, and CPU Kernels
- Parallelized eligible CPU weight prepacking during session initialization using the intra-op thread pool. Cross-session shared prepacked-weight caching retains its serialized path (#31691).
- Accelerated block-wise INT4/INT8 QMoE CPU experts with prepacked MLAS QNBit kernels and grouped dispatch. Added a
QMoEaccuracy_levelattribute and theORT_QMOE_CPU_QNBIT_GEMMoverride; INT4 uses FP32 activations by default, whileaccuracy_level=4opts into INT8 activation computation. INT8-weight experts use INT8 activations on the new path (#32644, #32668). - Added opt-in CPU
MatMulNBitsFP32 throughput mode withmlas.qnbit.force_fp32=1for eligible x86/x64 FP32-input, 4-bit, block-size-32, accuracy-level-4 nodes. It selects one numerical path for the whole session, is disabled by default, and trades single-request latency for batched throughput (#31723). - Added AVX2/FMA/F16C FP16 LayerNorm/RMSNorm kernels and an AVX-512 sliding-window NCHWc depthwise-convolution kernel (#32715, #32749).
- Integrated KleidiAI SVE2.1 FP16 GEMM with per-thread vector-length validation, retaining SME/SME2 priority where available. Updated KleidiAI to 1.30.0 (#32670).
- Optimized RISC-V RVV kernels and added RVV LinearAttention and FP32 minimum/maximum reduction kernels (#32540, #32710).
- Reduced coordinate and memory-access work in CPU
Col2Im/Im2col, accelerated adjacent-axis-groupTranspose, reduced repeated scans in 1DMaxPool, and optimizedBatchNormalizationfor NC inputs (#32698, #32800, #32817, #32891, #32936, #33113). - Enabled CPU INT8/UINT32
Wherekernels and independent scale/output types inDequantizeLinear; avoided unused dynamic QGEMM packed-buffer allocations (#32846, #32728, #32879).
Graph Optimization & Quantization Tools
- Added BatchNormalization fusion into
ConvTranspose(#29542). - Prevented unsupported CPU GELU/QuickGELU fusions, preserved public and shared graph values during fusion, and limited transpose movement through shared outputs to cases where it can cancel (#32427, #32431, #32991, #32868).
- Fixed LayerNorm fusion with zero normalized dimensions, Reshape fusion with node-produced shape inputs, duplicate node names during QDQ fusion, and
MatMulNBitsbias fusion with prepacked weights (#33079, #33091, #33115, #32973). - Retained optimized models when quantization preprocessing skips shape inference, raised model IR to version 10 when the quantizer emits INT4/UINT4 tensors, and avoided intermediate overflow when computing quantization ranges (#32803, #33068, #33104).
- Corrected symbolic shape inference for
Shapestart/end, missingConstantOfShapevalues during FP16 conversion, and packedGatedDeltaNetinputs. Preserved nested-graph serialization and FP16Castfallback consistency (#33157, #33089, #32708, #32727).
Security & Reliability
Model Loading and Graph Validation
- Limited sparse-initializer dense allocations, checked sparse value cardinality, validated inline raw initializer sizes and ORT-format string initializers, and included string payloads in initializer size limits (#32684, #32753, #32866, #32751, #32867).
- Rejected duplicate model-local functions, recursive function graph attributes, non-control-flow nodes carrying subgraphs, and inputs supplied to zero-input schemas; bounded AOT function expansion (#32750, #32864, #32641, #32633, #32662).
- Validated optional graph inputs,
Constantnodes, ORT-format input argument counts, and contrib-op shape inference; fixed out-of-bounds reads involving Loop/Scan initializers and shape-inference input data (#32940, #32865, #32607, #32609, #32681). - Bound external-data validation to opened files and rejected model-package-controlled output-file options, including optimized-model output paths (#33063, #33059, #33096).
- Strengthened optimizer validation for initializer types, shape propagation, QDQ weight shapes and axes, and LabelEncoder fusion attributes (#32647, #32653, #32648, #32649, #32726).
Runtime and Operator Correctness
- Serialized captured-graph replay with session runs and graph release for EPs that disallow concurrent runs. Checked CPU fallback assignments recursively in nested subgraphs (#32902, #32862).
- Corrected conversions between ABI-stable C API tensor element types and ONNX TensorProto numbering, particularly
UINT2,INT2, andFLOAT8E8M0, without renumbering public enum values. Fixed reuse of packed sub-byte buffers for full-byte tensors (#33066, #32611). - Strengthened recurrent-input rank, LayerNormalization axis, NCHWc layout, dynamic-pad, and
Resizescale validation; used checked arithmetic for CPU LSTM and DynamicQuantizeLSTM state extents (#32348, #32555, #32552, #32557, #33060, #32654, #32722, #33128). - Fixed CPU MHA shared-cache length handling, large wrap-mode padding, empty
TopK/DFTinputs, zero-channelConvTranspose, and real-window handling for complexSTFTsignals (#32723, #32725, #32724, #33064, #33136). - Corrected Arm64 FP16 GEMM Inf/NaN comparisons and NEON/SVE GELU/Erf overflow handling, fixed CPU INT2 per-axis
QuantizeLinearoutput on the last axis, and avoided degenerate CPU GQA FlashAttention tiles when L2 cache size is unknown (#31761, #32631, #33040, #32786, #33202).
Language Bindings & Web
- Added EPContext read/write callbacks in C#, Python, and Java (#32270, #32271, #32305).
- Added EPContext read callbacks in Node.js/React Native, Objective-C/Swift, and Rust. The JavaScript option is read-only and explicitly unsupported on Web/WASM (#32308, #32309, #32310).
- Added LoRA adapter support to the standard, non-proxy ONNX Runtime Web run path (#32886).
- Added Boolean tensor support in the Swift package and validated DLPack capsules before conversion (#32215, #33100).
Build, Packaging & Dependencies
- Validated CPU kernel registrations at compile time and improved CPU provider build time with an Eigen precompiled header (#29649, #32616).
- Fixed Windows ARM64
--enable_arm_neon_nchwchandling and added Windows ARM64 GitHub Actions coverage (#32094, #32791). - Improved no-exceptions error handling for provider bridge initialization, optional OpenVINO probing, model-package metadata parsing, and partition configuration; fixed Clang portability issues (#33051, #33112, #32883, #32894, #32853, #33140).
- Validated Protobuf consistency when ONNX comes from an installed package and repaired free-threaded Python wheel builds after the ONNX rollback (#31346, #33105).
- Updated Transformers to 5.3.0 and coremltools to 9.0, standardized telemetry dependency handling, and refreshed JavaScript dependencies and test lockfiles (#32752, #32586, #32705, #32996, #33003, #33194).
Contributors
Thanks to our 80 human contributors for this release!
@1duo, @abeclowicz, @aciddelgado, @adrastogi, @AkZGhost, @andife, @anubhavagr, @apsonawane, @artyomrabosh, @baijumeswani, @bbernhar, @bheu, @bingcheng1998, @bmehta001, @chilo-ms, @crvineeth97, @cursoragent, @damdoo001-arm, @danielsongmicrosoft, @edgchen1, @eserscor, @feich-ms, @galleonli, @GopalakrishnanN, @gramalingam, @gyagp, @hanbitmyths, @hariharans29, @Honry, @Hosi121, @hulkbig, @javier-intel, @jchen10, @jiafatom, @justinchuby, @kjg0724, @kuanyul-qti, @kunal-vaishnavi, @kylo5aby, @linzuojian, @maxim-davgalev, @mei1127, @melkap01-Arm, @miaobin, @mirounga, @mustjab, @NVMadhava, @prathikr, @pratikgx, @qjia7, @qti-yuduo, @raashish1601, @rohan-patnaik, @rvandermeulen, @saiyambharara, @saravanan-a-r, @sayanshaw24, @SharmaRithik, @shiyi9801, @SIDDARTHAREDDY8, @SiluPanda, @skottmckay, @sylvesterkaczmarek, @TANGBUDU, @the0cp, @tianleiwu, @titaiwangms, @velonica0, @vijay-kapse, @wangw-1991, @xadupre, @xhcao, @xiaofeihan1, @xiaohanAMD, @xiaoyu-work, @yairbenavraham, @ykhrustalev, @Yuhx141, Zapakio6, @Zestion
Full Changelog: rel-1.30.0...rel-1.31.0
Release highlights were prepared with AI assistance.