github altunenes/parakeet-rs v0.4.0

6 hours ago

parakeet-rs 0.4.0

Pure Rust backend via burn!

Every model can now run without ONNX Runtime, on burn, from the same .onnx files: TDT (v3 and Parakeet Ultra), CTC, Unified, EOU, Nemotron (English and multilingual), Multitalker, Sortformer diarization and Cohere Transcribe. ONNX Runtime stays the default and works as before. every model gives the same transcripts (and speaker segments for diarization) on burn as on ONNX Runtime, on both the CPU and the GPU. You can also check this with scripts/check_burn_parity.sh.

How to use it

1. Pick a feature in Cargo.toml. default-features = false leaves ONNX Runtime out entirely (no ONNX Runtime library to ship):

# Mac (Apple GPU)
parakeet-rs = { version = "0.4", default-features = false, features = ["metal"] }

# Any other GPU (Vulkan, DX12)
parakeet-rs = { version = "0.4", default-features = false, features = ["wgpu"] }

# CPU only
parakeet-rs = { version = "0.4", default-features = false, features = ["burn"] }

To keep ONNX Runtime and add burn next to it, keep the default features: features = ["metal"]. Optional models are added as before: features = ["metal", "cohere"].

2. Pick the matching execution provider in code:

use parakeet_rs::{ExecutionConfig, ExecutionProvider, ParakeetTDT};

let config = ExecutionConfig::new().with_execution_provider(ExecutionProvider::BurnWgpu);
let mut parakeet = ParakeetTDT::from_pretrained("./tdt", Some(config))?;
Feature Runs on Execution provider
burn CPU BurnCpu
wgpu any GPU (Metal, Vulkan, DX12) BurnWgpu
metal Apple GPUs (fastest on Mac) BurnWgpu
vulkan Vulkan GPUs (Linux, Windows, Android). Also works on Mac, where it runs on Metal underneath; metal is faster there BurnWgpu
burn-cuda NVIDIA GPUs through burn's CUDA backend BurnCuda
burn-rocm AMD GPUs through burn's ROCm backend BurnRocm
apple-amx, x86-v4 add these for faster burn on the CPU (Apple Silicon / AVX-512) BurnCpu

BurnCpu is available with every burn feature. Without ONNX Runtime it is also the default, so from_pretrained(path, None) runs on burn's CPU.

3. Use the fp32 model files. burn reads the same .onnx files as ONNX Runtime, but not int8 exports (they are skipped). Cohere's fp32 export is about 8.4 GB.

  • The first GPU run compiles and tunes kernels (seconds); they are cached on disk.
  • Every model's output was checked against ONNX Runtime with scripts/check_burn_parity.sh: identical transcripts and speaker segments on the test audio, on the CPU and on Metal.

Breaking changes

  • ort is now an optional (but still default) feature. With default-features = false, enable burn or one of the GPU features above.
  • ExecutionProvider, Error and ExecutionConfig are #[non_exhaustive]: add a _ => arm to matches on them, and build ExecutionConfig with new() and its with_* methods.

Other changes

  • Moondream's Parakeet TDT Ultra is supported (README, export script).
  • Repeated words are kept in word and sentence timestamps.
  • Sortformer diarization: fixed a failure on some audio lengths ("Array has a non-contiguous layout") and the mel feature layout.
  • Cohere keeps the encoder cache from the first decoding step instead of recomputing it on every other token.
  • Stricter ONNX loading: broken files give an error instead of a panic.
  • tokenizers 0.23.2.

What's Changed

New Contributors

Full Changelog: v0.3.8...v0.4.0

Don't miss a new parakeet-rs release

NewReleases is sending notifications on new releases.