github JustVugg/colibri v1.6.0
colibri v1.6.0

4 hours ago

colibrì v1.6.0 — the regression fixed, and prefill learns to read each expert once

If you are on v1.5.0, update now

v1.5.0 shipped a performance regression that hit GLM-5.2 (#856), left up to
~60 GB of RAM unused with a 13-point expert hit-rate loss (#885), and broke
Kimi K3 outright on some machines (#888). Two independent defects, both fixed:

  • #869 — the planner priced every expert row at the container's widest
    width, so mixed-width containers (int4 experts + int8 MTP) undercounted how
    many rows fit: the cache was silently halved.
  • #914 (bherald) — pin budgets now account real row widths; the reporter
    of #885 confirmed RSS (~105 GB on his box) and 93–97% hit rates restored.

If your symptom was "reinstalled, same commands, can't get above 0.2 tok/s"
(#939) — this release is the fix. No parameter changes needed.

DeepSeek V4: prefill I/O cut nearly in half

  • #941 — expert-major prefill. The batched prefill ran the FFN position
    by position: measured 4.37 disk reads per distinct expert — 42% of
    prefill bytes were re-reads. The MoE now routes the whole chunk first and
    reads each distinct expert once. On a 113-token prompt: 12,576 → 7,303
    disk reads, 168 → 98 GB moved. Token-exact by construction and verified
    against the per-position path on the real checkpoint (26/26 teacher-forced
    positions, 8/8 greedy). V4_EXPERT_UNION=0 restores the old path.
  • #927 — the expert-loader pool is now on. The dual loader existed but
    was never compiled in: every expert load ran at queue depth 1 against
    disks that scale nearly linearly to QD8. Pool on by default at depth 3;
    V4_LOADER_LANES=<1..16> raises it. Measured ladder on the same prompt,
    same cache-controlled harness: QD1 157 s → 3 lanes 129 s → 10 lanes 114 s
    to first token.
  • #897 (Blakeolson21) — the engine can size its OpenMP team around the
    loader workers instead of scheduling compute onto their CPUs. Caveat found
    while preparing this release, stated rather than buried: both launchers set
    OMP_NUM_THREADS themselves, which the engine reads as a deliberate
    operator choice, so this policy currently applies only when deepseek_v4
    is invoked directly. #958 addresses it and is deliberately held back — its
    first form reserved from logical CPUs, which would over-subscribe SMT
    hosts in exactly the way #805 fixed everywhere else.
  • #877 (ZacharyZcR) — the VRAM expert tier can be fed from RAM, not only
    from disk reads.

Formats and kernels

  • #926 (steve-m) — AVX2 kernel for the int3-g64 (fmt=5) matmul. It was
    scalar on every x86 CPU without AVX-512 — which is most consumer hardware.
    Cherry-picked from #601 with authorship intact.
  • The Vulkan compute path now decodes fmt=5 alongside int8/int4/MXFP4, and
    CI builds and runs the Vulkan backend headless under Lavapipe (#895),
    so that backend is no longer untested territory.

CLI and web

  • #928 — UP/DOWN recall previous prompts in chat (readline was never
    activated; arrow keys printed escape codes — #922).
  • #942 (dcutugno) — the last line of a streamed reply no longer vanishes
    when the token-count footer appears (#910): the markdown stream now closes
    before the spinner erases the line.
  • #909 — pasted multiline prompts survive intact.
  • #896 — the web chat renders markdown (matching the TUI) and no longer
    clamps replies to 1024 tokens.
  • #871 — the web dashboard waits for the engine instead of opening on a
    dead port; #872 — Kimi K3 gained a real RAM budget and --ram now
    means something; #850 — SIGTERM is handled.
  • #880 — the launcher is tested to dispatch every engine it ships,
    not merely bundle it (#879: OLMoE was named in the banner and routed to
    the GLM engine).

Tools

  • #902/#930 (terrizoaguimor) — an offline residency simulator (CNRE),
    with its cost calibration and capacity allocation corrected in review:
    physical misses are no longer conflated with routed requests, and the
    dynamic allocator evaluates the full feasible frontier.

Fixed in the final day, because contributors kept finding real bugs

  • #962 (Blakeolson21) — the FP8 converter ignored os.pwrite's return
    value: a short write silently truncated a shard, on every platform. Now it
    drains the buffer, and runs on native Windows too (O_BINARY, no pwrite).
  • #960/#961 (Blakeolson21) — the Inkling converter could not even start on
    native Windows (import fcntl), and both coli bench/coli convert looked
    for Scripts\python3, a file Windows venvs never create.
  • #956 (gouravkargwal) — a client disconnecting before the first token
    never cancelled: the turn ran to its token limit and the server thread
    stayed blocked (#908).
  • #957 (gouravkargwal) — the OpenMP spin-wait knobs are skipped on macOS:
    measured +122% decode time on M1 Max, reproduced on M3 (#707).
  • #968 (terrizoaguimor) — OLMoE: temperature now travels over COLI_TEMP
    (on Windows %TEMP% is a directory, and atof of a path silently forced
    greedy decoding), and the expert-cache last resort no longer steals a slot
    with a read in flight — two writers on one slab was silent corruption.
  • #944 (benmaster82) — a config with num_experts_per_tok larger than
    n_routed_experts is rejected at load instead of reading out of bounds.

Credits

bherald, mohamedmastouri2000-boop (cross-session methodology that kept our
own claims honest), steve-m, Blakeolson21, ZacharyZcR, terrizoaguimor,
dcutugno, Zach and everyone who measured, reported, and re-measured.

Full change list: v1.5.0...v1.6.0

Don't miss a new colibri release

NewReleases is sending notifications on new releases.