- Add Gemma 4 E2B instruction-tuned models in
q4f16_1andq4f32_1, with text and audio input. The FP32 variant does not requireshader-f16. - Update the bundled TVM runtime to
@mlc-ai/web-runtime@0.28.0-dev0and switch the prebuilt catalog tov0_2_86/base. - Rebuild 224 WebGPU libraries: 112 each in
baseandsg32. The release includes compiler revisions, configuration snapshots, checksums, validation results, and a rebuild script. - Use 4K-prefill Phi-3.5 vision libraries so default square images fit. The previous 2K variants remain in the binary folder.
- Include manifest-based audio and image loading, crash-resumable generation, and fixes for interrupted conversations, cancelled streams, model locks, token counts, and audio cleanup. Manifest loading remains optional; existing records retain the legacy path.
Validation: 466 WebLLM tests, 105 runtime tests, and 334 model/library parameter-schema checks passed. Browser checks passed for Gemma text, streaming and audio; legacy text generation; embeddings; and Phi vision, across base and subgroup variants. Both new FP32 models were checked with shader-f16 disabled.
Artifacts:
Compiler sources: TVM be03dc94b0, based on Apache main 27576ce2d1 with a WebGPU FP32 prefill fix; MLC-LLM 42fa855891, based on main 8978ea9aec with compatibility updates for current TVM APIs. Compilation used explicit TVM and MLC source paths rather than MLC's pinned TVM submodule.