Laya-MLX v0.4.0 selectively ports upstream Laya v0.4.1 runtime features to native MLX inference.
Behavior change
Router() now sends language-undecided inputs to the multilingual checkpoint. Identified English still uses the English checkpoint. Set Router(default="english") to retain the previous fallback behavior.
Changes
- Keep resident-model requests and status reads available during cold checkpoint construction. Deduplicate builds and synchronize targeted unloads per checkpoint.
- Register custom checkpoints with
register,unregister,registered, ormodels=. Source replacement and unregistration cannot resurrect an outdated in-flight model. - Add
predict_tournamentto narrow large choice sets without an embedding model. Its probabilities and usage describe the final pass; metadata records finalists and elimination rounds. - Add scalar and per-option-count
min_confidencegates to Agent and Router. Gate results are reported without replacing the original answer or automatically invoking a fallback. - Use FP32 intermediate computation for CPU parallel-option attention to avoid native FP16 kernel crashes on hosted macOS 26; GPU precision and parameter storage precision are unchanged.
- Support checkpoints trained with
option_layout="parallel"throughout MLX inference, compilation, prefix caching, and conversion. Existing checkpoints remain sequential; do not enable parallel layout by editing an old checkpoint's configuration. - Fix language detection for capitalized acronyms, address fragments, emphasis capitals, and mixed scripts; fix French device-footer removal.
- Release MLX free-buffer caches on unload and eviction while preserving caller-owned and in-flight model references.
- Update English/Chinese documentation, attribution, and CI with separate pinned sequential and parallel-layout upstream references.
Validation
- 231 tests passed on CPU and 231 on GPU before release.
- Three published checkpoints in FP32 and FP16: 378/378 argmax agreements with the pinned sequential upstream reference.
- 100 deterministic finite calls per configuration, 600 total, with zero measured active-memory growth.
- Parallel-layout tests cover all 24 permutations of four options in FP32/FP16, numerical comparison with upstream v0.4.1, compilation, caching, and export/reload.
- Ruff, wheel/sdist strict metadata checks, and installed-wheel smoke tests passed.
Raw checkpoint validation · API and migration details
Install or upgrade:
python -m pip install --upgrade laya-mlx==0.4.0