github mlc-ai/web-llm v0.2.86

3 hours ago
  • Add Gemma 4 E2B instruction-tuned models in q4f16_1 and q4f32_1, with text and audio input. The FP32 variant does not require shader-f16.
  • Update the bundled TVM runtime to @mlc-ai/web-runtime@0.28.0-dev0 and switch the prebuilt catalog to v0_2_86/base.
  • Rebuild 224 WebGPU libraries: 112 each in base and sg32. The release includes compiler revisions, configuration snapshots, checksums, validation results, and a rebuild script.
  • Use 4K-prefill Phi-3.5 vision libraries so default square images fit. The previous 2K variants remain in the binary folder.
  • Include manifest-based audio and image loading, crash-resumable generation, and fixes for interrupted conversations, cancelled streams, model locks, token counts, and audio cleanup. Manifest loading remains optional; existing records retain the legacy path.

Validation: 466 WebLLM tests, 105 runtime tests, and 334 model/library parameter-schema checks passed. Browser checks passed for Gemma text, streaming and audio; legacy text generation; embeddings; and Phi vision, across base and subgroup variants. Both new FP32 models were checked with shader-f16 disabled.

Artifacts:

Compiler sources: TVM be03dc94b0, based on Apache main 27576ce2d1 with a WebGPU FP32 prefill fix; MLC-LLM 42fa855891, based on main 8978ea9aec with compatibility updates for current TVM APIs. Compilation used explicit TVM and MLC source paths rather than MLC's pinned TVM submodule.

Don't miss a new web-llm release

NewReleases is sending notifications on new releases.