What's New in v0.14.0
Features
- Cohere Transcribe ASR engine (#487) — new 14-language encoder-decoder ASR backed by CohereLabs/cohere-transcribe-03-2026, converted to CoreML as a mixed-precision INT8 encoder + FP16 cache-external decoder hybrid.
Model
- INT8 hybrid (recommended): FluidInference/cohere-transcribe-q8-cache-external-coreml — 1.8 GB encoder, iOS 18+
- FP16 reference: FluidInference/cohere-transcribe-cache-external-coreml — 3.6 GB encoder, iOS 17+
Architecture
- 48-layer Conformer encoder, fixed input
[1, 128, 3500](35 s @ 10 ms hop) - 8-layer decoder with host-managed (Parakeet-style) external KV cache, 108-token window
- 16,384 SentencePiece tokens
- Max 35 s per call (matches upstream
max_audio_clip_s: 35)
Supported languages
English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Greek, Arabic, Japanese, Chinese, Korean, Vietnamese.
API
import CoreML
import FluidAudio
let models = try await CoherePipeline.loadModels(
encoderDir: encoderDir, decoderDir: decoderDir, vocabDir: decoderDir)
let result = try await CoherePipeline().transcribe(
audio: samples, models: models, language: .english)
print(result.text)Language must be specified explicitly — Cohere Transcribe uses a conditioned prompt that hard-codes the language; wrong language produces degenerate output.
Benchmarks
LibriSpeech test-clean, full split (2,620 utterances), Apple M2 (2022), Tahoe 26.0:
| Subset | Samples | WER | CER | RTFx (per-file mean) | RTFx (total audio/compute) |
|---|---|---|---|---|---|
| test-clean | 2,620 | 1.77% | 0.60% | 2.04× | 1.72× |
FLEURS, full splits, all 14 languages, M4 Pro, 48 GB:
| Code | Language | Samples | WER | CER | RTFx |
|---|---|---|---|---|---|
| en_us | English | 647 | 5.63% | 3.19% | 2.49× |
| fr_fr | French | 676 | 6.22% | 3.11% | 2.21× |
| de_de | German | 862 | 5.84% | 2.83% | 1.98× |
| es_419 | Spanish (LatAm) | 908 | 4.53% | 2.40% | 1.34× |
| it_it | Italian | 865 | 4.03% | 2.04% | 3.15× |
| pt_br | Portuguese (BR) | 919 | 6.44% | 3.38% | 2.79× |
| nl_nl | Dutch | 364 | 8.07% | 4.14% | 2.04× |
| pl_pl | Polish | 758 | 7.49% | 3.23% | 1.98× |
| el_gr | Greek | 650 | 11.50% | 5.45% | 2.00× |
| ar_eg | Arabic (EG) | 428 | 18.46% | 6.71% | 2.06× |
| ja_jp | Japanese | 650 | 60.13%† | 6.25% | 2.23× |
| cmn_hans_cn | Mandarin (Simp) | 945 | 98.52%† | 12.01% | 1.85× |
| ko_kr | Korean | 382 | 16.39% | 6.67% | 1.84× |
| vi_vn | Vietnamese | 857 | 9.55% | 6.87% | 1.55× |
†Japanese and Mandarin are written without word boundaries, so WER on the raw hypothesis is tokenization-artifact noise; CER is the accuracy metric for those languages.
CLI
# Single-precision (FP16 or INT8 in one dir)
swift run -c release fluidaudiocli cohere-transcribe audio.wav \
--model-dir /path/to/cohere-fp16 --language en
# Mixed precision (INT8 encoder + FP16 decoder)
swift run -c release fluidaudiocli cohere-transcribe audio.wav \
--encoder-dir /path/to/q8 --decoder-dir /path/to/f16 \
--vocab-dir /path/to/f16 --language en
# Benchmark (LibriSpeech or FLEURS) with per-100-file checkpointing
swift run -c release fluidaudiocli cohere-benchmark \
--dataset librispeech --subset test-clean \
--model-dir /path/to/q8 --auto-download \
--output cohere_testclean.jsonCaveats
- Language must be specified — no automatic language ID; pass
CohereAsrConfig.Languageon every call. - Single-chunk only — the CoreML pipeline processes one 35 s window per call. The upstream Python reference supports >35 s audio via sliding-window chunking with 5 s overlap; FluidAudio does not implement that wrapper yet — files exceeding 35 s are skipped by the benchmark CLI with a warning.
- Beta — first release of the Cohere integration.
Documentation
See Documentation/ASR/Cohere.md for full setup, upstream config provenance, and benchmark reproduction commands.
Model-conversion pipeline
The CoreML export + host-side feature/decode fixes live in FluidInference/mobius#41 — faithful numpy port of FilterbankFeatures (per-feature CMVN, n_fft=512, Slaney mel), cross-attention masking for padded encoder frames, SentencePiece byte-fallback detokenization, and INT8 W8A16 encoder quantization.
Full Changelog: v0.13.7...v0.14.0