Summary
Burn 0.22.0 brings five months of improvements to training and inference, along with our biggest API update yet. Over the last four years, we've kept APIs stable where possible while collecting your feedback on what could be improved. This release brings Burn's API much closer to what we want for 1.0.
The biggest change is that models and tensors no longer carry a backend generic. Devices select where and how they run, and editing a model requires much less recompilation. The release also adds LoRA and QLoRA support, parameter groups and multiple optimizers, graph capture, and improvements to remote execution. New tensor operations, optimizers, and kernel improvements round out the release.
For more details, check out the release post on our website.
No More B: Backend
Burn is still built around the Backend trait, with composable support for autodiff, kernel fusion, and remote execution. User code now selects these capabilities through Device, while models are plain types:
use burn::{
module::Module,
nn::{Linear, LinearConfig},
tensor::Device,
};
#[derive(Module, Debug)]
pub struct Model {
linear: Linear,
}
// Requires the cuda feature. Other options include Device::wgpu(..) and Device::flex().
let device = Device::cuda(0);
let model = Model {
linear: LinearConfig::new(4, 2).init(&device),
};More of the infrastructure is now implemented in Rust: Pliron replaces the MLIR-based intermediate representation, and Turso replaces the bundled SQLite dependency. CubeCL also adds LLVM-based CUDA and AMD GPU compilation paths. These changes reduce reliance on bundled C and C++ components and make more of the stack accessible to Rust contributors, though LLVM remains a native dependency for the backends that use it. Clean builds can take longer even as model-edit rebuilds become much faster.
Changelog
Breaking changes and migration
See Migrating to Burn 0.22 in the Burn Book for the complete migration guide, including before-and-after examples for devices, autodiff, modules, training, datasets, tensor operations, and custom integrations.
Start by removing backend parameters such as Model<B> and Tensor<B, D>, selecting a Device, and enabling autodiff with device.autodiff() before initializing a training model and its inputs. Candle has been removed in this release, and NdArray and LibTorch are being deprecated.
Records now use burnpack. Legacy MessagePack, binary, and JSON recorder files must be loaded with an older compatible Burn version and exported through burn-store before importing the weights into 0.22. Follow the checkpoint migration instructions; this transfers model parameters, not legacy optimizer or scheduler state.
What's Changed
Tensor & Modules
- [Breaking] Remove
Tensorbackend generic and add high-levelDevicestruct (#4717) @laggui - Refactor tensor kind traits (#4957) @laggui
- refactor(tensor): add
BridgeTensorto bridge high-level tensor API with dispatch (#4958) @laggui - Bool tensor API improvements (#4955) @Cielbird
- fix(nn): conv1d initializer fan out (#4973) @laggui
- feat(ops): autodiff for rfft/irfft (#4956) @cofinite
- feat: add
Calibration::AbsMeanfor BitNet b1.58 ternary weight quantization (#4989) @simeon-kepp - refactor(tensor)!: replace
TensorKind::id()withconst KINDand renameTensorKindId->Kind(#5018) @laggui - refactor(tensor): seal
TensorKindandParametertraits and consolidate param impls (#5022) @laggui - fix(module): clearer error when generic used in both module and skip fields (#5028) @jaweed3
- fix(tensor): detect NaN element-wise in contains_nan to avoid f16 false positives (#5027) @zhan-wei-919
- fix(quantization): fallback (dequant -> op -> quant) implementation for slice, gather, select and expand (#5020) @Andy2887
- feat(nn): add rotary encoding reset (#5042) @nicolaslara
- feat(module): add Identity Module, extend Activation and Normalization. (#5058) @crutcher
- feat(nn): add MultiHeadAttention configuration flags for Linear::bias. (#5059) @crutcher
- fix(param): cloning shares lazy initialization state (#5045) @laggui
- fix(tensor): prevent unit dim strides from corrupting
split_stridesoutput (#5071) @ArthurBrussee - feat(tensor): hypot function (#5065) @skewballfox
- feat(module): param group (#5086) @Charles23R
- fix(tensor): resolve unsupported bool storage to device default (#5094) @yasumorishima
- feat(autodiff): implement prod and prod_dim (#5106) @leonardwetzel
- Implement int_scatter_nd for autodiff backend (#5114) @weiji14
- Feat/lora qlora (#5139) @Charles23R
- Feat/spread param groups (#5154) @Charles23R
- LstmState: Clone, Debug, initial, shape, stack, slice, (un)squeeze (#5167) @crutcher
- Feat/relu6 activation (#5143) @Bellman281
- feat: add hardtanh and tanhshrink activation functions (#5144) @Bellman281
- feat: add threshold activation function (#5145) @Bellman281
- feat: add margin ranking loss (#5151) @Bellman281
- feat: add HingeEmbeddingLoss (#5174) @Bellman281
- feat: add soft margin loss (#5149) @Bellman281
- feat: add RReLU activation module (#5150) @Bellman281
- feat: add Mish, SiLU and LogSigmoid activation modules (#5183) @Bellman281
- fix: compute margin^p without std-only f64::powi (no_std build) (#5176) @Bellman281
- Add LstmState::chunk method (#5178) @jaweed3
- feat: add CosineSimilarity module (#5186) @Bellman281
- feat: add PairwiseDistance module (#5185) @Bellman281
- feat: add TripletMarginLoss (#5175) @Bellman281
- feat: add GaussianNLLLoss (#5182) @Bellman281
- feat: add PixelShuffle and PixelUnshuffle modules (#5184) @Bellman281
- feat: support negative GLU dimensions (#5203) @mattisonchao
- fix(autodiff): zero-safe cumprod gradient (#5141) @AnayGarodia
- feat: support negative softmax dimensions (#5205) @mattisonchao
- Fix norm accumulation (#5211) @nathanielsimard
- feat: support negative cumulative dimensions (#5213) @mattisonchao
- feat: support negative FFT dimensions (#5214) @mattisonchao
- feat: support negative sort and top-k dimensions (#5215) @mattisonchao
- feat: support negative dimensions in statistics and norms (#5225) @mattisonchao
- feat: Tensor<D, Int>::square() support (#5224) @crutcher
- feat: support negative dimensions across tensor APIs (#5226) @mattisonchao
- fix(autodiff): fix deadlock when a Backward step runs a nested backward (#5194) @laurigates
- refactor(tensor): make
_elemcomparison methods aliases to unify_scalarmethods (#5230) @laggui - fix(tensor): preserve dimension check errors (#5232) @mattisonchao
- feat(tensor): add mask_select (boolean mask indexing) default impl (#5238) @Yashiru
- feat(burn-nn): support negative dimensions (#5241) @mattisonchao
- feat: add
Tensor::empty_like(#5242) @crutcher - feat(tensor): add
extractmethod; simplifyinplaceimplementation. (#5207) @crutcher - refactor(burn-tensor): remove duplicate dimension checks (#5243) @mattisonchao
- feat(tensor): add
tensor.select_dim(dim, index)(#5246) @crutcher - feat(nn): add fold4d / Fold4d module (inverse of unfold4d) (#5173) @Bellman281
- feat(tensor): add QR decomposition (#5102) @anagaev
- feat(quant): support UE4M3 scale params end to end on CPU (#5253) @ThierryCantin-Demers
- fix(quant): round quantization scales up instead of to nearest (#5261) @laggui
- fix(tensor):
TensorDataconvert to bool stores by truthiness (#5269) @SamuelBelanger - fix(tests): skip some qr tests in f16 (#5272) @laggui
- feat(nn): add AdaptiveAvgPool3d module (#5115) @jaweed3
- fix(tensor): topk check for
k <= shape[dim](#5307) @laggui - test(burn-tensor): run the doc examples instead of only compiling them (#5302) @4ktLuffy
- Feat/autodiff quantized ops (#5317) @nathanielsimard
- feat(module): generalize weight reparameterization abstraction (#5311) @laggui
- feat(module): add record param group in the collector and allow loading unused params (#5336) @nathanielsimard
- fix(quant): dequantize transposed quantized tensor (#5281) @subotac
- fix(tensor)!: standardize NaN propagation in extrema reductions (#5290) @jcwal1516
- refactor(nn): add functional
batch_normforward operation (#5347) @laggui - fix(tests): batch_norm f16 tol (#5348) @laggui
- fix(module): build LoRA factors under a persistent-allocation window (#5362) @nathanielsimard
- fix(module): build LoRA factors at an explicit dtype over a packed base (#5364) @nathanielsimard
- fix(module): preserve LoRA factor dtype when composing weights (#5365) @laggui
- feat(tensor)!: expand
TensorDataconversion and access APIs (#5316) @crutcher @laggui - fix(tensor): preserve dtype in padding operations (#5386) @laggui
- feat(quant)!: support two-level quantization scales (#5262) @ThierryCantin-Demers
- fix(nn): use mean reduction in group norm (#5410) @laggui
- fix(tests): correct FFT fixture and linalg tolerances (#5415) @laggui
- fix(tensor): prevent softplus from overflowing to inf (#5374) @Mikyx-1
- fix(tensor): validate TensorData byte length on deserialization (#5360) @robertomeroni
- refactor(tensor): move padding operations to backend (#5505) @laggui
- feat(tensor): support mul indexing updates on NdArray and Flex (#5325) @mattisonchao
- feat(module): generalize
Paramwith module-owned flags (#5498) @laggui - fix(module): keep the param mapper across valid() and from_inner (#5509) @nathanielsimard
- test(quantization): strengthen quantized layout coverage (#5503) @Mikyx-1
- feat(autodiff): implement topk backward (#5531) @Mikyx-1
- fix(autodiff): zero-safe prod gradients (#5534) @Mikyx-1
- feat(tensor): support assign indexing updates across backends (#5532) @mattisonchao
- fix(autodiff): prod backward broadcasting (#5547) @laggui
- feat(tensor): add batched SVD decomposition to linalg (#5259) @sehaxe
- feat(module): separate gradient control from module freezing (#5537) @laggui
- fix(tensor): validate matmul batch broadcast in TensorCheck (#5555) @ax1s-x1zz
- feat(tensor): replace
no_gradwith explicit autodiff conversions (#5557) @laggui - feat(tensor)!: add multi-axis vector norm variants and update empty
max_abs_dimssemantics (#5539) @Sadik00789 - feat(tensor): add assert_shape! and debug_assert_shape! macros (#5554) @antimora
- feat(tensor): apply TensorCheck to remainder, powi, powf, hypot, atan2 (#5564) @ax1s-x1zz
- fix(nn): exclude pad tokens from cross-entropy normalization (#5569) @Mikyx-1
- feat(tensor): add
is_autodiffandis_trackedstate inspection (#5571) @laggui - feat(tensor): validate matmul rank in TensorCheck (#5580) @ax1s-x1zz
- feat(linalg)!: extract tensor linalg into
burn-linalgextension crate (#5572) @laggui - test(backend): cover empty-axis autodiff reductions (#5598) @lorenzozanee
- Feat/storage tiled carrier (#5631) @louisfd
- fix(nn): honor count_include_pad for asymmetric average pooling (#5592) @DivyamTalwar
- perf(autodiff): reduce a broadcast gradient over all dims at once (#5623) @Marc-AnthonyG
- fix(tensor): avoid cosine similarity denominator underflow (#5585) @Mikyx-1
- feat(autodiff): support graph-preserving cross-backend transfers (#5624) @laggui
- test(burn-std): import vec! macro in layout.rs tests for no-std (#5639) @ax1s-x1zz
- test(backend): cover empty-axis autodiff product reductions (#5641) @lorenzozanee
- fix(linalg): add autotune feature propagation (#5648) @laggui
- fix(autodiff)!: retain input nodes until child registration (#5647) @laggui
- fix(autodiff): explicitly reject consumed graph reuse and preserve reusable leaves (#5645) @laggui
- feat(tensor): implement Min and Max scatter/select_assign across backends (#5582) @ax1s-x1zz
- perf(autodiff): avoid redundant traversal in checkpoint topological sort (#5679) @haydenflinner
- fix(linalg): support negative even lp norm orders (#5693) @Mikyx-1
- perf(nn): use a dedicated BatchNorm training op with closed-form backward (#5676) @Marc-AnthonyG @laggui
- feat(tensor): add einsum with runtime and macro APIs (#5583) @Mikyx-1
- fix(autodiff): preserve gradients for zero scalar exponents (#5692) @SRaswan
- feat(signal)!: extract tensor signal into
burn-signalextension crate (#5720) @laggui - fix(core): let an init_mapper parameter train on an autodiff device (#5705) @ThierryCantin-Demers
- refactor(module)!: merge
AutodiffModuleintoModule(#5721) @laggui - fix(tensor): fix quiet softmax for negative infinity slices (#5743) @Mikyx-1
- fix(core): count a module's parameters without initializing them (#5746) @ThierryCantin-Demers
- fix(tensor): preserve f64 precision in degree/radian conversions (#5745) @Mikyx-1
- fix(tensor): return NaN for infinite dividends in fmod_scalar (#5754) @Mikyx-1
- feat(core): a lazy parameter initializes on the device it moves to (#5739) @ThierryCantin-Demers
- fix(module): preserve lazy initialization when collecting devices (#5778) @laggui
- fix(tensor): keep input dtype on empty slice/repeat/cat results and one_hot_fill (#5790) @antimora
- fix(std): ignore unit axes when checking contiguity (#5738) @Liberxue
- fix(linalg): import Vec in fusion svd for no_std builds (#5804) @antimora
- fix(tensor): reject quantized inputs in *_like ops and one_hot (#5827) @antimora
- fix(autodiff): correct repeat_dim gradient ordering (#5837) @lorenzozanee
- Qa/tiled storage (#5839) @louisfd
- fix(module): preserve reparameterizations during validation (#5821) @laggui
- refactor(tensor): remove into_tiled from the public API (#5846) @laggui
- fix(tensor): make TensorData fields private and validate construction (#5838) @antimora
- feat(tensor)!: accept output_size or scale_factor in InterpolateOptions (#5849) @antimora
- fix(nn): validate scalar module configurations (#5860) @Mikyx-1
- Refactor/storage tile naming (#5894) @louisfd
- fix(tensor): display quantized tensors as metadata without panicking (#5868) @onenewcode
- fix(nn): clamp probabilities in smoothed cross-entropy (#5892) @Arthur031221
- fix(nn): apply attention dropout after softmax (#5875) @Mikyx-1
- test(autodiff): tighten weak tests and cover untested backward paths (#5889) @antimora
- refactor(tensor)!: move pooling args into options structs with asymmetric padding (#5853) @antimora
- test(backend): run the empty select and roll tests (#5916) @antimora
- fix(tensor): roll shift direction to match torch.roll (#5912) @Arthur031221
- fix(nn): omit unused forget gate in coupled LSTM (#5862) @Mikyx-1
- fix(nn): validate BatchNorm configuration (#5930) @Mikyx-1
- fix(quantization): round-trip a tensor packed along an outer axis (#5960) @louisfd
- fix(derive): use field init shorthand in generated Config::new (#5981) @jwric
- feat(quantization): a storage-tiled quantized tensor keeps its tiles (#5982) @louisfd
Training, Optimizers & Datasets
- Add A-FINE image quality metric (#4894) @Capataina
- Make RL event types mod public (#4951) @laggui
- feat(metric): add BLEU score training metric (#4937) @kimjune01
- feat(metric): add ROUGE-L score metric (#4967) @jaweed3
- feat(metric)!: add multiclass and multi-label support to AUROC metric (#4960) @manucouto1
- fix(train): guard TUI metric navigation against empty state (#4987) @SAY-5
- feat(train): add training and evaluation progress logger and add them to the event processors (#4980) @Zatiji
- fix(train): scroll TUI metric tabs to keep selected visible (#4995) @LucaCappelletti94
- feat(train): add mouse support to TUI metric navigation (#4998) @LucaCappelletti94
- refactor(rl): add inference device and update dqn example (#5009) @Charles23R
- refactor(train)!: transform all Progress items into a global progress struct (#5012) @Zatiji
- fix(train): use recv_timeout instead of try_recv. (#5021) @jeandudey
- fix(dataset): fully qualify HuggingFace dataset repo ids (#5050) @ThierryCantin-Demers
- fix: forward features flag in burn Cargo.toml for burn-vision (#5078) @Marc-AnthonyG
- feat(metric): add AUC-PR (Average Precision) training metric (#4963) @manucouto1
- fix(train): load checkpoint on the correct device (#5084) @laggui
- fix(data): reuse persistent dataloader workers across epochs (#5073) @manucouto1
- feat(train): add custom checkpointers (#5097) @Zatiji
- Fix checkpointing + minor stuff (#5109) @Charles23R
- fix metric logging when checkpoint + burn-train integration tests (#5116) @Charles23R
- Feat/multi optimizers (#5121) @Charles23R
- feat(dataset): make image extension matching case-insensitive (#5152) @veezhang
- chore: change loop logic for data iterators and batchers (#5198) @Zatiji
- fix(tui): silently unwind on training kill signal (#5223) @laggui
- Feat/add get many in dataset trait (#5220) @Zatiji
- feat: add sequential learning rate scheduler (#5204) @AnayGarodia
- fix(optim)!: remove erroneous restart in
CosineAnnealingLrScheduler(#5191) @DeathSurfing - fix(burn-train): correctly aggregate per-epoch metrics that require global statistics (#5218) @laggui
- refactor(dataset)!: replace
Vec<PixelDepth>withPixelData(#5271) @Charles23R - Change visibility of optimizer (#5314) @nathanielsimard
- Keep metric items on their device instead of syncing them to the host (#5319) @nathanielsimard
- fix(train): flush async metrics before checkpoint/early-stop (#5320) @TsaoLun
- feat(vision): add color conversion and blurring transforms (#5339) @Charles23R
- fix(vision): keep even filter2d output size (#5414) @Soundcreates
- fix(optim): keep tensor data in optimizer records (#5407) @original4422
- feat(optim): add LAMB optimizer (#5495) @Mikyx-1
- fix(train): aggregate Dice statistics across batches (#5515) @Mikyx-1
- fix(train): return final ROUGE-L value (#5517) @Mikyx-1
- fix(rl): preserve deterministic mode in async batches (#5574) @Ultronen
- fix(train): exclude padded samples from accuracy aggregation (#5578) @Mikyx-1
- fix(optim): keep a parameter's checkpointing strategy across an update (#5618) @Marc-AnthonyG
- refactor(dataset): back SqliteDataset with Turso instead of rusqlite (#5546) @antimora
- fix(burn-dataset): cap MNIST item counts at the split size (#5667) @li-jin-quan
- chore(deps): reduce dataframe dataset dependencies with polars-core (#5680) @laggui
- fix(train): handle tied scores in average precision (#5684) @Mikyx-1
- chore(deps): trim zip default features and fix burn-dataset nlp feature (#5715) @antimora
- fix(train): flush partial gradient accumulation and step LR per optimizer update (#5729) @Mikyx-1
- feat(optim): add Adafactor optimizer (#5775) @Mikyx-1
- feat(optim): add Lion optimizer (#5741) @Mikyx-1
- test(optim): strengthen optimizer save-load round-trip coverage (#5799) @laggui
- fix: remove rl from default features (#5806) @antimora
- feat(train): add labels to training progress loggers (#5805) @Charles23R
- fix(optim): compute gradient norm clipping in F32 for half precision (#5843) @antimora
- fix(tui): handle already-joined thread in manual close (#5866) @orbitwebsites-cloud
- fix(train): compute AUROC with sorted score groups (#5900) @Mikyx-1
- refactor(train)!: propagate execution errors through training and evaluation (#5896) @Charles23R
- fix(train): validate top-k accuracy bounds (#5948) @Mikyx-1
- fix(train): validate MS-SSIM configuration (#5950) @Mikyx-1
Backends & Performance
- Refactor interpolate from cubek (#4928) @SamuelBelanger
- Cubek-pool integration (#4948) @SamuelBelanger
- Feature gate the backend ops extension when no backend feature is enabled (#4953) @laggui
- Refactor/isolate burn backend deps (#4954) @nathanielsimard @laggui
- Re-export BurnConfig with fusion/autodiff getters (#4959) @truffle-dev
- refactor(extension): feature gate
Tensor::from/into_primitiveand addfrom/into_bridgew/TensorKindIdvalidation (#4961) @laggui - fix(fusion): shared tensor (#4962) @nathanielsimard
- refactor(cubecl): update to reference changes (#4974) @wingertge
- fix(dispatch):
WgpuandWebGpure-export feature gating (#4981) @laggui - fix(fusion): resolve tensor into_ir (re-entrancy bug) (#4984) @laggui
- feat(cube): interpolate nearest exact mode (#4982) @SamuelBelanger @louisfd
- perf(core)!: improve compile times via opaque inner types to break dependency chain (#4977) @nathanielsimard
- refactor: use obfuscate from burn-std (#4997) @nathanielsimard
- feat(dispatch): add remote backend (#4994) @nathanielsimard
- refactor(backends)!: remove associated element types & replace with device defaults (#5000) @laggui
- fix(ndarray): broadcast remainder operands (#5002) @puneetdixit200
- feat(wgpu): use
WgpuRuntimecompiler generic to enable specialized aliases (#5001) @nathanielsimard - fix(dispatch): use the correct webgpu/vulkan/metal/wgpu device (#5010) @laggui
- refactor(cubecl): add lifetime to
View(#4999) @wingertge - added FloatKind::F64 arm for FuseType implementation (#5011) @Andy2887
- fix(flex): max_pool3d_backward indices type (#5017) @laggui
- perf(flex): bulk-copy inner-contiguous run in to_contiguous (#5019) @jkaczman
- feat(extension): add support for custom output types wrapping tensor primitives (#5030) @laggui
- perf: improve compilation speed with into-scalar (#5037) @nathanielsimard
- feat: add missing
Devicestaging and memory cleanup methods (#5048) @laggui - feat(cube): interpolate autotune (#5044) @SamuelBelanger
- feat(fusion): add CubeCL fusion lowering for int bitwise ops (#5057) @iamorlando
- Improve compilation speed with into-scalar (#5061) @nathanielsimard
- feat(extension): support async fn + futures and add
Tensorlow-level backend primitive interop (#5049) @laggui - perf(ndarray): read select indices from a contiguous slice (#5066) @IvanUkhov
- feat(cubecl): wire logical any/all reductions (#5051) @zhan-wei-919
- perf(ndarray): reshape standard-layout arrays without copying (#5067) @IvanUkhov
- fix(fusion): use dtype_to_storage_type in autotune checks (#5085) @laggui
- fix(autotune): feature gate interpolate autotune strategy (#5090) @laggui
- feat(autotune): support cpu gemm (#5081) @louisfd
- refactor(backend): move device to TensorMetadata (#5089) @skewballfox
- fix(cubecl): materialize lazy device bytes in from_data (#5095) @laggui
- fix: guard against divide-by-zero in matmul autotune TMA priority (#5117) @ltouati
- Fuse Cat operations (#5133) @nathanielsimard
- Perf/fusion dag (#5135) @nathanielsimard
- fix(fusion): don't fuse single-use view ops (#5134) @nathanielsimard
- Fix fusion plan re-exploration for sync segments with a drained tail (#5137) @nathanielsimard
- Fix fused in-place aliasing (never fired since #2870/#3263) (#5138) @nathanielsimard
- Fix interpolate autotune key (#5136) @SamuelBelanger
- Feat/graph capture (#5146) @nathanielsimard
- Refactor how graph are captured (#5148) @nathanielsimard
- Feat/autotune throughput (#5157) @SamuelBelanger @ThierryCantin-Demers
- Feat/cubecl memory (#5158) @nathanielsimard
- Fix matmul bounds check when launch (#5160) @nathanielsimard
- drop from foreign stream drains home stream (#5166) @Charles23R
- fix(ndarray): propagate NaN in argmax/argmin/cummin/cummax (#5155) @jaweed3
- Fix/attention fallback activation gate (#5187) @nathanielsimard
- perf(fusion): use cubek's dedicated any/all reduce instruction (#5113) @zhan-wei-919
- chore: activate burn-wgpu burn-backend default cubecl feature flag (#5199) @ThierryCantin-Demers
- fix(flex): support GQA/MQA in attention (q_heads % kv_heads == 0) (#5142) @AnayGarodia
- feat(burn-cubecl): topk_with_indices (#5200) @ThierryCantin-Demers
- feat(extension): support struct and enum inputs for backend extensions (#5221) @ThierryCantin-Demers
- perf(cube): pooling and interpolate optimization (#5112) @SamuelBelanger @ThierryCantin-Demers
- fix(tests): nhwc_relayout expected shapes and dtypes equality (#5229) @laggui
- feat(ndarray): native multiply-free ternary matmul for Q2S (BitNet b1.58) (#5075) @simeon-kepp
- fix(burn-cubecl): scatter_nd values offset for multi-dim batches (#5196) @artyompal
- feat(backend): add default int_argtopk impl and native tch support (#5236) @IssaAlBawwab
- fix(cube): skip unsupported accelerated kernels during autotune (#5239) @SamuelBelanger
- Update to new cubecl environment (#5251) @nathanielsimard
- feat(fusion): custom fusion optimization (#5240) @nathanielsimard
- fix(burn-cubecl): fix cubek double buffering matmul (#5252) @laggui
- feat(burn-cubecl): max/min_dim_with_indices in a single launch (#5273) @ThierryCantin-Demers
- Fix: wgpu fused reduce broadcast produced wrong values (#5268) @Charles23R
- Fix/fusion read composition (#5282) @nathanielsimard
- feat(ndarray): support assign indexing updates for floats and ints (#5245) @mattisonchao
- feat(backend): promote mask_select to a backend op with dispatch (#5283) @Yashiru
- fix(fusion): decline to fuse unsupported quant scale params (#5279) @ThierryCantin-Demers
- fix(cubecl): use real out_channels in conv_transpose2d autotune key (#5289) @SamuelBelanger
- refactor(tests): add flex features to enable simd or rayon, disable allocation tracking (nicer output) (#5285) @laundmo
- perf(flex): precompute interpolation axis taps (#5274) @huahuadeliaoliao
- feat(autotune): centralize roofline bounds registration for matmul, reduce, and attention (#5188) @ThierryCantin-Demers @SamuelBelanger
- fix(dispatch): remove incorrect checkpointing debug assertion in grad_replace (#5294) @4ktLuffy
- fix(cubecl): update rev for constant bitcasts fix (#5306) @laggui
- perf(flex): improve add_bias and conv_plane_accumulate autovectorization (#5267) @laundmo
- fix(burn-tch): panic broadcasting a zero-sized dimension on the tch backend (#5305) @hdimer
- fix(cubecl): add missing device throughput on rocm and cpu (#5312) @ThierryCantin-Demers
- fix(cubecl): allow metadata-only reshape of sub-byte quantized tensors (#5321) @ThierryCantin-Demers
- fix(dispatch): keep a WgpuDevice conversion when wgpu features unify (#5323) @ThierryCantin-Demers
- fix(extension): dispatch autodiff ops to the correct backend (#5326) @laggui
- chore!: remove deprecated burn-candle backend (#5343) @antimora
- chore: deprecate burn-ndarray backend (#5344) @antimora
- fix(reduce): return the identity for a zero-length axis, and reject the extrema (#5333) @robertomeroni
- Migrate to pliron (#5324) @wingertge
- fix(burn-cubecl): broadcast a size-1 dim against a zero-sized dim (#5335) @Yashiru
- fix(ndarray): avoid invalidating in-flight handles in UnsafeSharedRef (#5346) @robertomeroni
- fix(tests): flex simd and rayon feature graph (#5350) @laggui
- fix(flex): replace all remaining cases of function pointer passing (#5358) @laundmo
- fix(flex): make float_sign branchless so it compiles on Xtensa (#5331) @Bellman281
- perf(flex): reduce int cast codegen size (#5357) @laundmo
- fix(dispatch): preserve autodiff device for non-float tensors (#5340) @tianrking
- fix(burn-tch): preserve parent storage through permute/swap_dims (#5376) @andoriyu
- perf(flex): collapse both operand layouts into a joint loop nest for broadcast binary ops (#5123) @Andy2887
- perf(flex): pooling impl optimization by rewriting inner kernel loop (#5361) @laundmo
- feat(capture): add graph capture backend producing
GraphIr(#5377) @laggui - fix(fusion): reject a padded reference layout (#5395) @ThierryCantin-Demers
- fix(flex): preserve non-finite padding semantics in conv (#5394) @tandede
- fix!: address fusion and autodiff edge cases (#5400) @Charles23R @laggui @nathanielsimard
- perf(cubecl): reduce, don't pool, for a 1x1 adaptive average pool (#5405) @nathanielsimard
- fix(capture): preserve initializers across capture scopes (#5408) @laggui
- test(capture): update cross-device tensor movement contract (#5409) @laggui
- fix(metal): update cubecl rev for MSL capability detection (#5396) @laggui
- Perf/fusion layout propagation (#5406) @nathanielsimard
- refactor(burn): route backend features through the dispatch chain (#5369) @laurigates
- fix(cubecl): adapt matmul integration to cubek changes (#5421) @laggui
- perf(cubecl): compute a depthwise weight gradient in one kernel (#5422) @nathanielsimard
- perf(fusion): let a concatenation choose its own layout (#5417) @nathanielsimard
- perf(cubecl): cut a pointwise weight gradient's contraction (#5424) @nathanielsimard
- Let a caller install and read a device's memory pools (#5425) @nathanielsimard
- perf(conv): route strided pointwise convolutions through the matmul path (#5429) @ThierryCantin-Demers
- perf(cubecl): compute a dense weight gradient over the input's columns (#5428) @nathanielsimard
- fix(cubecl): restrict GEMV CPU autotune priority to 1D vector workloads (#5423) @Sadik00789
- refactor(cubecl): interpolate autotuning with cubek strategies and roofline bounds (#5502) @SamuelBelanger
- fix(router): support rfft and irfft (#5513) @Mikyx-1
- feat(tch): support mul indexing updates (#5511) @mattisonchao
- feat(cubecl): support mul indexing updates (#5512) @mattisonchao
- test(cubecl): rename indexing tests to reflect reference matching (#5519) @laggui
- Refactor/cubecl runtimes (#5528) @nathanielsimard
- refactor(backend): remove specialized indexing add operations (#5523) @laggui
- chore: deprecate burn-tch backend (#5525) @antimora
- perf(cubecl): offer
Gemmon CPU for general matmuls (#5538) @ThierryCantin-Demers - fix(fusion): a panicking operation must not poison its stream (#5535) @nathanielsimard
- refactor(dispatch): unify routing and strengthen autodiff contract (#5520) @laggui
- fix(dispatch): restore autodiff context promotion (#5551) @laggui
- Refactor/fusion write scope (#5540) @nathanielsimard
- fix(ndarray): preserve SIMD unary layout and recip precision (#5553) @laggui
- feat(backend)!: support asymmetric padding in conv1d and conv2d (#5521) @laggui
- perf(conv): accumulate direct convolution channels in vector registers on CPU (#5545) @ThierryCantin-Demers
- Chore/update cubecl runtime erasure (#5556) @nathanielsimard
- fix(dispatch): decouple CubeCL runtime facade crates (#5570) @laggui
- feat(cubecl): support AdaptiveAvgPool3d (#5372) @jcwal1516
- perf(fusion): let a padded tensor vote for the block's layout (#5620) @Marc-AnthonyG
- perf(fusion): settle only the unfusable head of a block (#5622) @Marc-AnthonyG
- feat(cubecl): implement integer powi operations (#5616) @laggui
- perf(flex): write slice_assign in place when the destination is uniquely owned (#5603) @hdimer
- test(fusion): change f32 expected value to oracle (#5633) @laggui
- perf(cubecl): scatter-add runs one unit per value with atomic adds (#5621) @Marc-AnthonyG
- fix(cubecl): gate storage_tiled tests behind runtime features (#5635) @laggui
- perf(flex): optimize views, copies, in-place ops, and Rayon parallelism (#5613) (#5617) @Sadik00789
- perf(cubecl): walk elementwise kernels in operands' memory order (#5625) @Marc-AnthonyG
- fix(flex): avoid overflow in boolean select updates (#5629) @Ultronen
- fix(flex): support unsigned dtypes in int_argmax/int_argmin (#5630) @Punisheroot
- fix(flex): guard sum fast path with layout_covers_storage_once (#5634) @antimora
- refactor(cubecl): call cubek's direct convolution routine (#5646) @ThierryCantin-Demers
- fix(cubecl): update cubek to fix sums over overlapping views (#5654) @laggui
- fix(flex): validate gather_nd and scatter_nd coordinates (#5640) @stack3mpty
- fix(flex): round nearest grid_sample ties to even (#5663) @yuefdev
- fix(flex): make u64 right shift logical (#5661) @yuefdev
- Update/cubecl crate layout (#5670) @nathanielsimard
- chore(deps)!: make backend tracing opt-in (#5675) @laggui
- fix(flex): honor asymmetric padding in conv3d (#5650) @onenewcode
- perf(flex): parallelize attention across (batch, head) pairs (#5612) (#5636) @Sadik00789
- fix(burn-ndarray): sign(NaN) must be 0, not its hidden sign bit (#5665) @SRaswan
- fix(flex): handle remainder edge cases (#5652) @onenewcode
- fix(flex): propagate NaN through max pooling (#5662) @yuefdev
- fix(cubecl): same-runtime device moves without a peer transport, and quantized moves (#5678) @ThierryCantin-Demers
- fix(flex): propagate NaN through relu and clamp_min/clamp_max (#5658) @Liberxue
- feat(extension): generate Fusion implementations for backend extensions (#5673) @laggui
- fix(cubecl): update dependencies to fix empty tensor readback (#5704) @laggui
- fix(cubecl): propagate NaN through two-sided clamp (#5718) @Liberxue
- fix(dispatch): require explicit backend selection and support backend-free builds (#5722) @laggui
- fix(flex): handle NaN clamp bounds (#5717) @Liberxue
- feat(tensor): device identity and one entry per physical GPU (#5688) @ThierryCantin-Demers
- Feat/profiling (#5726) @nathanielsimard
- chore: remove recursion limit workarounds and update CubeCL configs (#5736) @laggui
- fix(fusion): avoid autotuning empty reductions (#5725) @lorenzozanee
- fix(ndarray): preserve NaN through float clamps (#5733) @lorenzozanee
- Chore/cubecl device capacity (#5763) @nathanielsimard
- fix(cubecl): handle broadcasting in mask_where and mask_fill (#5767) @laggui
- Chore/cubecl environment records (#5771) @nathanielsimard
- fix(backends): correct average-pooling gradients with ceil mode (#5653) @slobodaapl
- test(backend): skip ue4m3 scale accuracy cases where no 8-bit type exists (#5794) @vboussot
- fix(flex): compute f32 layer_norm variance in two passes (#5780) @antimora
- perf(cubecl): NHWC direct kernel for conv_transpose2d (#5760) @onenewcode
- Chore/cubecl adaptive pool (#5797) @nathanielsimard
- perf(flex): reduce ConvTranspose time and memory use (#5761) @Hosi121
- perf(cubecl): walk cast in its input's memory order (#5796) @vboussot
- Update cubecl and cubek to the tune plan record (#5810) @nathanielsimard
- fix(test): align dtype-support expectations with the runtimes (#5814) @onenewcode
- Metal/wgpu msl tests (#5815) @louisfd
- fix(backend): match PyTorch pooling output size in ceil_mode (#5798) @antimora
- fix(flex): mask shift amounts to each int dtype's own width (#5808) @antimora
- perf: use grouped convolution for fold4d (#5811) @AnkitNakhawa
- perf(router): avoid rescanning the graph when scoring fusion (#5850) @ThierryCantin-Demers
- Share the autotune roofline policy and bound fused matmul and reduce tuning (#5822) @SamuelBelanger
- fix(fusion): keep the default vectorization axis for matmul operands the epilogue reads (#5871) @nathanielsimard
- fix(backend)!: separate device settings queries from initialization (#5869) @laggui
- perf(flex): upgrade macerator to 0.5.0 and hoist SIMD dispatch out of hot loops (#5865) @antimora
- fix(fusion): resolve reduce reference strides against the reference's argument list (#5847) @antimora
- fix(router): close the fuser on an upload so the upload can be freed (#5851) @ThierryCantin-Demers
- fix(cubecl): reject reshape that splits a quantization block across rows (#5844) @antimora
- fix(router): serve unsigned int tensors in the interpreter (#5877) @ThierryCantin-Demers
- fix(flex): enforce quantization block alignment (#5888) @laggui
- test(fusion): relax half-precision absolute tolerance for matmul epilogue view (#5893) @laggui
- fix(fusion): update cubek to reject invalid VecMat kernel shapes (#5890) @laggui
- fix(fusion): avoid remapping fused matmul input offsets twice (#5902) @laggui
- fix(cubecl): skip the reduce launch when another axis leaves the output empty (#5881) @antimora
- perf(fusion): defer a cross-thread drop until its producer runs instead of draining the stream (#5901) @ThierryCantin-Demers
- fix(cube): propagate NaNs in clamp and stabilize GPU tests (#5909) @laggui
- fix(metal): require native MSL for explicit Metal devices (#5921) @laggui
- feat(wgpu): add device initialization options (#5933) @laggui
- fix(tch): flatten the broadcast
alphabefore calling tch's prelu (#5925) @hiro-nakanishi - fix(cubecl): skip large deppthwise filters with fewer than 32 channels (#5935) @laggui
- perf(conv): add im2col data-gradient path for dense convolutions (#5813) @onenewcode
- fix(tensor): pin wgpu backend during device enumeration (#5942) @Pewpenguin
- perf(cube): update cubecl for llvm TF32 support on CUDA (#5979) @laggui
Distributed & Remote Execution
- fix(remote): add remote device settings fetched during client init (#5008) @laggui
- fix(remote): correct multi-streams and drastically improve performance (#5029) @nathanielsimard
- refactor(distributed)!: add
DistributedContextandall_reduceto high-level API (#5013) @laggui @Charles23R - chore: cleanup remote feature flags (#5034) @laggui
- Feat/multi device remote (#5036) @nathanielsimard
- Remote backend fusion: client-side op-graph caching (#5088) @nathanielsimard
- Add remote backend extension (#5101) @nathanielsimard
- perf(remote): materialize tensor readback off the shared runtime (#5104) @nathanielsimard
- Feat/iroh remote backend (#5111) @jwric
- Fix/remote async read (#5126) @nathanielsimard
- fix(remote): leave the process-wide log subscriber to the program (#5724) @ThierryCantin-Demers
- fix(remote): turn off iroh GSO until iroh#4555 is fixed (#5819) @ThierryCantin-Demers
- fix(remote): try again while a server is not reachable yet (#5824) @ThierryCantin-Demers
- fix(remote): notice a WebSocket peer that vanished without closing (#5825) @ThierryCantin-Demers
- fix(remote): blocking reads inside a tokio task stall after 128 (#5829) @ThierryCantin-Demers
- fix(remote): end a server session when its client disconnects (#5828) @ThierryCantin-Demers
- test(remote): serve the tests on ports the OS picks, with burn-remote linked once (#5873) @ThierryCantin-Demers
- fix(remote): let a WebSocket server read what its clients send (#5903) @ThierryCantin-Demers
- fix(remote): refuse a session for a device the server does not host (#5904) @ThierryCantin-Demers
- fix(remote): keep remote devices off local GPUs' runner threads (#5898) @ThierryCantin-Demers
- fix(remote): clean up a session whose task panicked (#5878) @ThierryCantin-Demers
- feat(remote): configure an Iroh server's relays, port and authorizer, and dial one by its id (#5910) @ThierryCantin-Demers
- test(remote): feed a queued reader from tensors dropped on another thread (#5919) @ThierryCantin-Demers
- fix(remote): gate what partial feature builds leave unused (#5920) @ThierryCantin-Demers
- feat(remote)!: connecting to a remote device returns a Result (#5922) @ThierryCantin-Demers
- feat(remote)!: one client type and one server type for the remote API (#5938) @ThierryCantin-Demers
- fix(remote): open streams only on connections a server dialed (#5980) @ThierryCantin-Demers
Model Storage & Import
- fix(features): expose safetensors/pytorch support directly (#4985) @crutcher
- fix: include RMS_NORM in normalization layer detection for safetensors adapter (#5023) @ogghead
- fix(store): add .allow_partial(true) hint in TensorNotFound error w/ docs (#5032) @jaweed3
- feat(store): extract burnpack format to burn-pack + add minimal record (#5064) @nathanielsimard
- Refactor Record - Serialization & Deserialization of Module, Optimizer & LrScheduler (#5083) @nathanielsimard
- feat(burn-store): add FloatCastAdapter for target-driven float dtype casting (#5164) @laurigates
- fix(store): restore persisted ParamId on load_record and Applier apply (#5177) @jaweed3
- fix(store): preserve module param id on load (#5180) @nathanielsimard
- fix(burn-store): prevent CPU exhaustion via pickle memo bomb (#5120) @Infinty-ux
- fix(burn-core): record params through their on_save mapper (#5208) @ThierryCantin-Demers
- feat(pack): stream tensors on demand when writing burnpack files (#5349) @antimora
- refactor(store)!: replace TensorSnapshot with burn_pack::Tensor (#5411) @antimora
- fix(store): abort instead of double-dropping when a mapper unwinds (#5488) @antimora
- fix(store): preserve extensionless burnpack paths (#5494) @rioyu123
- fix(store): respect PyTorch tensor strides (#5392) @original4422
- fix(store): exclude adapter-matched tensors from unused (#5536) @apoorvdarshan
- fix(store): return deserialization error instead of panicking in nested Deserializer (#5496) @Sadik00789
- fix(pack): validate size from aligned data offset (#5530) @Mikyx-1
- fix(store): write safetensors files atomically (#5489) @antimora
- fix(store): reject a loaded tensor whose dtype is the wrong kind (#5490) @antimora
- fix(store): reject invalid path filter regex (#5558) @Mikyx-1
- fix(pack): validate tensor byte lengths when reading (#5576) @Mikyx-1
- fix(store): harden and restructure the PyTorch reader (#5593) @antimora
- test(pack): check streaming memory with a drop hook, not a global allocator (#5666) @yuefdev
- fix(store): collect PyTorch tensors nested in lists/tuples with indexed names (#5664) @SIDDARTHAREDDY8
- refactor(store): extract the PyTorch reader into a burn-free pytorch-reader crate (#5656) @antimora
- fix(pytorch-reader): report values load_config cannot represent instead of defaulting (#5728) @antimora
- perf(pytorch-reader): read stored ZIP entries outside the archive lock (#5714) @antimora
- fix(pytorch-reader): remove unsafe visitor cloning and deserialization panics (#5700) @Sadik00789
- fix(pytorch-reader): open a checkpoint once and never reopen it by path (#5737) @antimora
- fix(pytorch-reader): handle unsupported values in deserialize_any (#5764) @laggui
- fix(burn-store): scope contiguous index mapping per prefix (#5750) @antimora
- fix(pytorch-reader): thiserror source chains + reject non-checkpoint files (#5749) @superroket169
- fix(pytorch-reader): accept checkpoints saved with compute_crc32=False (#5756) @antimora
- feat(pytorch-reader): load torch.save(model) full-model pickles (#5766) @antimora
- fix(store): enforce overwrite(false) at publish time, not just before the save (#5781) @antimora
- feat(pack): make atomic writes the default (#5832) @antimora
Documentation & Examples
- burn 0.21 cite (#4969) @Redhawk18
- fix(doc): correct SgdConfig::init (#4986) @jaweed3
- fix(docs): improve error messages with shape/dimension context (#4996) @jaweed3
- fix(doc): inconsistent assertion for non-negative (#5016) @YichiZhang0613
- fix(example):
Device::enumerateconfigure (#5056) @laggui - chore(docs): update readme (#5080) @nathanielsimard @laggui
- docs: document per-device default dtypes; make AlreadyInitialized error actionable (#5209) @laurigates
- fix(examples): swapped lstm cell/hidden args (#5234) @ra1u
- fix(examples): fix lstm bias backprop (#5235) @ra1u
- fix(docs): correct LBFGSConfig::init doc comment (#5256) @Mikyx-1
- docs(book): remove backend generics, update device, learner, module and add optimizer section (#5276) @laggui
- fix(docs): correct tensor sort doc examples in orderable.rs (#5298) @4ktLuffy
- fix(docs): correct chained reduction doc examples in numeric.rs (#5301) @4ktLuffy
- docs(pack): clarify tensor thread-safety contract (#5500) @laggui
- docs(tensor): fix unfold window formula (#5559) @Mikyx-1
- docs(flex): fix stale statements in burn-flex docs and comments (#5619) @7487
- docs(book): update backend extension guides and include maintained example sources (#5731) @laggui
- docs(pytorch-reader): sync README with lib.rs (#5747) @superroket169
- docs: fix typos and grammar in books and API comments (#5748) @hikmetba-bit
- docs: clarify guidelines for minor documentation fixes (#5765) @laggui
- docs: update guides and examples for Burn 0.22 (#5762) @laggui
- docs: cover remaining 0.22 migration points and refresh stale pages (#5768) @antimora
- docs: select the package when running examples from the repo root (#5773) @Liberxue
- clarify 0.22 migration guidance and update API examples (#5800) @laggui
- fix(examples): correctly format web inference probability labels (#5816) @laggui
- fix(examples): correct WebGPU inference and keep live predictions responsive (#5823) @laggui
- docs(book): document burn.toml runtime configuration (#5845) @laggui
- docs: fix dead links in README and the Burn Book (#5863) @pratikgx
- feat(examples): train and run MNIST on another machine's GPU with remote-mnist (#5906) @ThierryCantin-Demers
- docs: fix license badge links in crate READMEs (#5923) @pratikgx
- docs: update device selection table (#5937) @laggui
- docs(book): sync the book with 0.22 API changes (#5954) @antimora
- fix(examples): repair examples that no longer build or run (#5956) @antimora
- docs: clarify burn-cpu vs burn-flex CPU backends (#5955) @antimora
- docs(book): update ONNX import chapter for burn-onnx 0.22 (#5972) @antimora
- docs: 0.22 release documentation pass (#5973) @antimora
- docs: fix broken DeepWiki badge in README (#5966) @antimora
- feat(example): bf16 training for ag-news text classification (#5984) @nathanielsimard
Maintenance & Dependencies
- Bump version to 0.22.0-pre.1 (#4933) @laggui
- fix(deps): update enumset to 1.1.13 (#4979) @laggui
- Update CubeCL & CubeK (#4992) @louisfd
- chore(cubek): update to interpolate refactor (#5003) @SamuelBelanger
- fix(dep): update aes to 0.9.1 (yanked) (#5047) @laggui
- fix(ci): use flex as
default_backendand split jobs into backends, crates, and examples shards with merged lcov report (#5041) @laggui - fix(ci): remove
fail_ci_if_errorfor code coverage (#5055) @laggui - chore(deps): bump actions/download-artifact from 7 to 8 (#5053) @dependabot[bot]
- chore(deps): bump codecov/codecov-action from 6 to 7 (#5054) @dependabot[bot]
- fix: fix inconsistent assertions (#5062) @YichiZhang0613
- chore: bump MSRV from 1.92 to 1.95 (#5077) @laggui
- chore(deps): update polars to 0.54 (#5072) @getong
- update cube @louisfd
- chore(deps): bump actions/checkout from 6 to 7 (#5093) @dependabot[bot]
- fix(audit): update quinn-proto version in Cargo.lock (#5099) @laggui
- chore(deps): update cube (#5098) @louisfd
- update dependency versions in Cargo.lock (#5165) @SamuelBelanger
- update cubek rev (#5169) @Charles23R
- update cubecl (#5181) @Charles23R
- chore(deps): bump actions/stale from 10 to 11 (#5247) @dependabot[bot]
- chore: release 0.22.0-pre.1 (#5249) @laggui
- fix: update publish deps (#5250) @laggui
- Update revs (#5265) @nathanielsimard
- update cube (#5275) @louisfd
- chore: update cubek rev w/ matmul fixes (#5291) @laggui
- update cube (#5299) @louisfd
- fix(xtask): apply the webgpu feature on the wgpu CI runner (#5300) @4ktLuffy
- chore(deps): update cubecl and cubek for
BlockLevel::BlockTensor(unimplemented) (#5303) @ThierryCantin-Demers - update cubecl/cubek (#5315) @ThierryCantin-Demers
- update cube (#5322) @louisfd
- fix(deps): update cubecl and use correct windows crate versions (#5327) @laggui
- fix: correct some copy-pasta mistakes (#5329) @laggui
- chore: bump to version 0.22.0-pre.2 (#5338) @laggui
- Bump cubecl and cubek revs for the pliron codegen fixes (#5355) @nathanielsimard
- chore: simplify cargo run-checks local validation (#5356) @laggui
- fix(deps): update h2 transitive dep (#5381) @laggui
- Update cube (#5389) @louisfd
- chore(ci): pin the shared actions to v11 (#5390) @ThierryCantin-Demers
- chore(ci): pin the publish workflow to v11 (#5391) @ThierryCantin-Demers
- chore: remove stale comment in root Cargo.toml (#5397) @ThierryCantin-Demers
- chore: tracel github actions v11 (#5398) @syl20bnr
- chore: fix clippy lints (#5401) @laggui
- chore(deps): update cubecl + cubek (#5403) @ThierryCantin-Demers
- chore: update to xtask v5 (#5404) @syl20bnr
- Update/cube (#5420) @louisfd
- chore: fix removed import (#5427) @laggui
- Chore/bump cubecl cubek revs (#5497) @nathanielsimard
- fix(deps): restore windows dependency compatibility in lockfile (#5499) @laggui
- chore(deps): update cubecl + cubek to fix fma sigsev (#5501) @ThierryCantin-Demers
- Update/cube (#5518) @louisfd
- chore(deps): follow cubecl's collapsed memory throughput modes (#5510) @ThierryCantin-Demers
- fix(no-std): use shared sync primitives across crates (#5543) @laggui
- chore(deps): reduce unnecessary dependencies (#5544) @nathanielsimard
- chore(deps): update cubecl and cubek for fallible throughput probes (#5561) @ThierryCantin-Demers
- chore: update cubecl and cubek revs (#5586) @nathanielsimard
- ci: reduce redundant compilation in macOS tests (#5669) @laggui
- fix(deps): update rustls for cargo audit (#5672) @laggui
- ci: reduce redundant compilation across test suites (#5671) @laggui
- ci: reduce examples overhead by skipping coverage setup and enabling caching (#5681) @laggui
- fix(deps): avoid enabling CubeCL through linalg defaults and respect vision defaults (#5694) @laggui
- chore: update cubek (#5701) @louisfd
- chore: update cubek (#5703) @Charles23R
- fix(ci): publish burn-einsum and include it in no-std checks (#5727) @laggui
- fix(deps): restore the windows crate versions the cube update changed (#5735) @ThierryCantin-Demers
- chore: bump version to 0.22.0-pre.4 (#5776) @laggui
- chore(deps): bump iroh to 1.2.0 (#5830) @ThierryCantin-Demers
- chore(deps): drop unused bincode workspace dependency (#5831) @antimora
- chore(deps): bump cubek to a066bdd (#5841) @louisfd
- chore(deps): bump cubek to b40b522 (#5848) @louisfd
- chore(deps): bump cubecl to a1bb768 and cubek to 9db95ba (#5854) @nathanielsimard
- fix(deps): restore gpu-allocator's windows version to match wgpu-hal (#5867) @ax1s-x1zz
- chore(deps): bump cubek to 7d60e30 (#5870) @louisfd
- chore(deps): bump cubek to 41a4ab0 and cubecl to 33b6dfb (#5887) @louisfd
- chore: bump cubecl to f0cf8383 and cubek to 9f35e2ab (#5895) @nathanielsimard
- chore: bump cubek to 6927503 (#5905) @louisfd
- update cube (#5907) @louisfd
- chore: bump cubek to bb461fb (#5913) @louisfd
- chore: bump cubek to 5753e85 (#5915) @louisfd
- chore: bump cubecl to 48301b5 and cubek to 2e07d52 (#5917) @nathanielsimard
- chore: bump cubek to e84f83d and cubecl to 48301b5 (#5918) @louisfd
- chore: bump cubek to 08f8f96 (#5926) @louisfd
- chore(lint): allow redundant_field_names pending upstream Clippy fix (#5927) @laggui
- chore(lint): allow redundant_field_names pending upstream Clippy fix (#5928) @laggui
- chore: bump cubek to cafd262 and cubecl to 3631455 (#5934) @louisfd
- chore: bump cubek to 645ead8 (#5936) @louisfd
- chore: bump cubek to fbed329 (#5940) @louisfd
- chore: bump cubek to 3d53983 and cubecl to 202a9bc (#5943) @louisfd
- chore: bump cubek to 94a2266 (#5946) @louisfd
- chore: bump cubek to a8207f8 (#5957) @louisfd
- fix(ci): let the scheduled valgrind and cargo-careful jobs install their tools (#5941) @glaziermag
- chore: upgrade dependencies (#5958) @antimora
- chore: update version to 0.22.0 (#5985) @laggui