github ggml-org/llama.cpp b11454

latest release: b11455
pre-releaseone hour ago
Details

model : add K2 Horizon dense and MoVA support (#29535)

  • model: K2 Horizon gguf conversion code

  • model: loading hparams and tensors in k2-horizon.cpp

  • model: K2 Horizon compute graph

  • model: K2 Horizon compute graph adjustment and registering tokenizers

  • model: K2 Horizon chat template and accomodate safetensors naming

  • unicode : add the K2-Horizon pre-tokenizer splitter

The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the
general std::regex fallback, which fails two ways.

On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all:
llama-quantize, llama-imatrix and llama-perplexity all abort with
regex_error(error_escape) before a token is produced.

Where the fallback does compile it is still wrong. unicode_regex_split collapses each
codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D
are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback
byte. The literal ‌ and ‍ alternatives in K2's regex can then never match and
every ZWNJ or ZWJ ends a letter run.

The splitter is the existing llama3 one with a single rule widened, since K2's regex
differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ.

tests/test-unicode.cpp gains a case for this: it fails before the change with
[Amy] [ZWNJ khaham] and passes after with the run intact.

  • tests: expand K2 Horizon unicode splitter coverage

  • unicode: handle K2 Horizon case folding and empty input

Assisted-by: Codex

  • jinja : support sequence indices in selectattr and rejectattr

Assisted-by: Codex

  • model : add K2 Horizon dense and MoVA support

Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes.

Assisted-by: Codex

  • chat : support K2 Horizon reasoning and tool calls

Assisted-by: Codex

  • conversion: remove obsolete K2 Aurora alias

Assisted-by: Codex

  • k2-horizon: enforce response schemas and load YaRN betas

Constrain final JSON after reasoning, accept flexible JSON tool envelopes,
enforce XML dialects, and handle repeated or alternate thinking markers.
Load YaRN beta metadata instead of retaining the default values.

Add schema, streaming, continuation, and model reload regressions. Validate
CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips.

Assisted-by: Codex

  • renaming template fixture

  • adressing cisc follows ups

  • desloppify the parser / adress aldehir comments

  • clean test-chat

  • remove fallback : model trained mostly on high anyway

  • fix k2 attn_v_exp tn splitting and metal fusion baseline

  • k2-horizon : forward expand views before sums

  • k2-horizon: copy embds before group norm to fix TP

  • disable tesnor parallelism


Co-authored-by: Ryandito Diandaru ryandito.diandaru@mbzuai.ac.ae
Co-authored-by: WestWaters mario.papaleo2013@gmail.com
Co-authored-by: Natani L. Mayday 71436458+TaskPuppyNatani@users.noreply.github.com
Co-authored-by: West 100190545+WestWaters@users.noreply.github.com
Co-authored-by: aaryamonvikram aaryamonvikram@gmail.com
Co-authored-by: aaryamonvikram 96529820+aaryamonvikram@users.noreply.github.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.