github ggml-org/llama.cpp b11443

latest release: b11445
pre-release3 hours ago
Details

models : consolidate nextn row cropping into shared helpers (#30017)

  • mimo2 : always emit h_nextn

the other nextn-capable models set it unconditionally

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

  • models : consolidate nextn row cropping into shared helpers
  • replace the duplicated crop conditions and the per-model flags (narrow_early,
    crop_before_ffn, crop_last_layer, emit_h_nextn) with two helpers on llm_graph_context:
    crop_before_nextn() / crop_after_nextn()
  • models that only tested embeddings_nextn_masked now share the same condition, so they
    crop the last layer before the nextn capture whenever extraction is off
  • t_h_nextn is now set unconditionally in mimo2, qwen4exp and deepseek4 (as in the other
    nextn-capable models); host-side reads stay gated by cparams.embeddings_nextn

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.