github FluidInference/FluidAudio v0.14.0
v0.14.0 Cohere Transcribe

latest releases: v0.17.7, v0.17.6, v0.17.5...
5 months ago

What's New in v0.14.0

Features

  • Cohere Transcribe ASR engine (#487) — new 14-language encoder-decoder ASR backed by CohereLabs/cohere-transcribe-03-2026, converted to CoreML as a mixed-precision INT8 encoder + FP16 cache-external decoder hybrid.

Model

Architecture

  • 48-layer Conformer encoder, fixed input [1, 128, 3500] (35 s @ 10 ms hop)
  • 8-layer decoder with host-managed (Parakeet-style) external KV cache, 108-token window
  • 16,384 SentencePiece tokens
  • Max 35 s per call (matches upstream max_audio_clip_s: 35)

Supported languages

English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Greek, Arabic, Japanese, Chinese, Korean, Vietnamese.

API

import CoreML
import FluidAudio

let models = try await CoherePipeline.loadModels(
    encoderDir: encoderDir, decoderDir: decoderDir, vocabDir: decoderDir)
let result = try await CoherePipeline().transcribe(
    audio: samples, models: models, language: .english)
print(result.text)

Language must be specified explicitly — Cohere Transcribe uses a conditioned prompt that hard-codes the language; wrong language produces degenerate output.

Benchmarks

LibriSpeech test-clean, full split (2,620 utterances), Apple M2 (2022), Tahoe 26.0:

Subset Samples WER CER RTFx (per-file mean) RTFx (total audio/compute)
test-clean 2,620 1.77% 0.60% 2.04× 1.72×

FLEURS, full splits, all 14 languages, M4 Pro, 48 GB:

Code Language Samples WER CER RTFx
en_us English 647 5.63% 3.19% 2.49×
fr_fr French 676 6.22% 3.11% 2.21×
de_de German 862 5.84% 2.83% 1.98×
es_419 Spanish (LatAm) 908 4.53% 2.40% 1.34×
it_it Italian 865 4.03% 2.04% 3.15×
pt_br Portuguese (BR) 919 6.44% 3.38% 2.79×
nl_nl Dutch 364 8.07% 4.14% 2.04×
pl_pl Polish 758 7.49% 3.23% 1.98×
el_gr Greek 650 11.50% 5.45% 2.00×
ar_eg Arabic (EG) 428 18.46% 6.71% 2.06×
ja_jp Japanese 650 60.13%† 6.25% 2.23×
cmn_hans_cn Mandarin (Simp) 945 98.52%† 12.01% 1.85×
ko_kr Korean 382 16.39% 6.67% 1.84×
vi_vn Vietnamese 857 9.55% 6.87% 1.55×

†Japanese and Mandarin are written without word boundaries, so WER on the raw hypothesis is tokenization-artifact noise; CER is the accuracy metric for those languages.

CLI

# Single-precision (FP16 or INT8 in one dir)
swift run -c release fluidaudiocli cohere-transcribe audio.wav \
    --model-dir /path/to/cohere-fp16 --language en

# Mixed precision (INT8 encoder + FP16 decoder)
swift run -c release fluidaudiocli cohere-transcribe audio.wav \
    --encoder-dir /path/to/q8 --decoder-dir /path/to/f16 \
    --vocab-dir /path/to/f16 --language en

# Benchmark (LibriSpeech or FLEURS) with per-100-file checkpointing
swift run -c release fluidaudiocli cohere-benchmark \
    --dataset librispeech --subset test-clean \
    --model-dir /path/to/q8 --auto-download \
    --output cohere_testclean.json

Caveats

  • Language must be specified — no automatic language ID; pass CohereAsrConfig.Language on every call.
  • Single-chunk only — the CoreML pipeline processes one 35 s window per call. The upstream Python reference supports >35 s audio via sliding-window chunking with 5 s overlap; FluidAudio does not implement that wrapper yet — files exceeding 35 s are skipped by the benchmark CLI with a warning.
  • Beta — first release of the Cohere integration.

Documentation

See Documentation/ASR/Cohere.md for full setup, upstream config provenance, and benchmark reproduction commands.

Model-conversion pipeline

The CoreML export + host-side feature/decode fixes live in FluidInference/mobius#41 — faithful numpy port of FilterbankFeatures (per-feature CMVN, n_fft=512, Slaney mel), cross-attention masking for padded encoder frames, SentencePiece byte-fallback detokenization, and INT8 W8A16 encoder quantization.


Full Changelog: v0.13.7...v0.14.0

Don't miss a new FluidAudio release

NewReleases is sending notifications on new releases.