parakeet-rs 0.4.0
Pure Rust backend via burn!
Every model can now run without ONNX Runtime, on burn, from the same .onnx files: TDT (v3 and Parakeet Ultra), CTC, Unified, EOU, Nemotron (English and multilingual), Multitalker, Sortformer diarization and Cohere Transcribe. ONNX Runtime stays the default and works as before. every model gives the same transcripts (and speaker segments for diarization) on burn as on ONNX Runtime, on both the CPU and the GPU. You can also check this with scripts/check_burn_parity.sh.
How to use it
1. Pick a feature in Cargo.toml. default-features = false leaves ONNX Runtime out entirely (no ONNX Runtime library to ship):
# Mac (Apple GPU)
parakeet-rs = { version = "0.4", default-features = false, features = ["metal"] }
# Any other GPU (Vulkan, DX12)
parakeet-rs = { version = "0.4", default-features = false, features = ["wgpu"] }
# CPU only
parakeet-rs = { version = "0.4", default-features = false, features = ["burn"] }To keep ONNX Runtime and add burn next to it, keep the default features: features = ["metal"]. Optional models are added as before: features = ["metal", "cohere"].
2. Pick the matching execution provider in code:
use parakeet_rs::{ExecutionConfig, ExecutionProvider, ParakeetTDT};
let config = ExecutionConfig::new().with_execution_provider(ExecutionProvider::BurnWgpu);
let mut parakeet = ParakeetTDT::from_pretrained("./tdt", Some(config))?;| Feature | Runs on | Execution provider |
|---|---|---|
burn
| CPU | BurnCpu
|
wgpu
| any GPU (Metal, Vulkan, DX12) | BurnWgpu
|
metal
| Apple GPUs (fastest on Mac) | BurnWgpu
|
vulkan
| Vulkan GPUs (Linux, Windows, Android). Also works on Mac, where it runs on Metal underneath; metal is faster there
| BurnWgpu
|
burn-cuda
| NVIDIA GPUs through burn's CUDA backend | BurnCuda
|
burn-rocm
| AMD GPUs through burn's ROCm backend | BurnRocm
|
apple-amx, x86-v4
| add these for faster burn on the CPU (Apple Silicon / AVX-512) | BurnCpu
|
BurnCpu is available with every burn feature. Without ONNX Runtime it is also the default, so from_pretrained(path, None) runs on burn's CPU.
3. Use the fp32 model files. burn reads the same .onnx files as ONNX Runtime, but not int8 exports (they are skipped). Cohere's fp32 export is about 8.4 GB.
- The first GPU run compiles and tunes kernels (seconds); they are cached on disk.
- Every model's output was checked against ONNX Runtime with
scripts/check_burn_parity.sh: identical transcripts and speaker segments on the test audio, on the CPU and on Metal.
Breaking changes
ortis now an optional (but still default) feature. Withdefault-features = false, enableburnor one of the GPU features above.ExecutionProvider,ErrorandExecutionConfigare#[non_exhaustive]: add a_ =>arm to matches on them, and buildExecutionConfigwithnew()and itswith_*methods.
Other changes
- Moondream's Parakeet TDT Ultra is supported (README, export script).
- Repeated words are kept in word and sentence timestamps.
- Sortformer diarization: fixed a failure on some audio lengths ("Array has a non-contiguous layout") and the mel feature layout.
- Cohere keeps the encoder cache from the first decoding step instead of recomputing it on every other token.
- Stricter ONNX loading: broken files give an error instead of a panic.
tokenizers0.23.2.
What's Changed
- bugfix: Keep repeated words in word and sentence timestamps by @altunenes in #131
- Add moondream's parakeet ultra by @altunenes in #132
- Pass Sortformer chunks to ort in standard layout by @umarrasydan in #135
- Burn wgpu by @altunenes in #136
- Add cohere to the burn backend by @altunenes in #137
- Prepare 0.4.0 burn hardening, fixes and version bump by @altunenes in #138
- expose apple-amx and x86-v4 features for burn on the cpu by @altunenes in #139
New Contributors
- @umarrasydan made their first contribution in #135
Full Changelog: v0.3.8...v0.4.0