This release selectively brings upstream Laya runtime fixes through 4aa6761 into the native MLX port. The neural architecture parity reference remains 573e5b6; this is not full upstream API parity.
Changes:
- Preserve the newest conversation turns under truncation, including prefix-cached inference; fix the zero-room left-truncation edge case.
- Support custom noul labels and boolean criteria keys; reject criteria keys that would otherwise be silently ignored.
- Preserve Unicode in structured instructions and improve named input-validation errors.
- Add
answer_confidencewhile retaining the existingconfidencesemantics. - Report state truncation and options whose token spans collapse under the question budget.
- Update language detection and email cleaning to preserve real requests and improve multilingual routing.
- Preserve resident checkpoints during incremental preload and handle empty or language-neutral hints correctly.
- Protect the prompt-prefix LRU under concurrent preparation and bound cosine similarity results.
Validation: 176 local tests passed, plus 162 upstream email checks and 38 upstream language-statistics checks adapted to the MLX implementation. Lint, formatting, wheel/source-distribution builds, and strict distribution checks passed. Real-checkpoint accuracy and long-running memory benchmarks were not rerun.
Maintenance direction: prioritize native MLX implementations of upstream behavior over independent model variants, service APIs, and additional demos.
Future vX.Y.Z tag pushes now run the test/build/publish workflow and create a GitHub release automatically.