github ggml-org/llama.cpp b10796

latest release: b10797
pre-release4 hours ago
Details

src : add n_expert_used_max function (#28323)

  • src : add n_expert_used_max function

With Commit c61b98b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:

llama_model_load: error loading model: model has expert layers but no expert layers are used

And later:

/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: #25444 (comment)

  • src : use hparams.n_expert_used_max in llama_model_base::load_hparams

  • src : use 0 as initial value for n_expert_used_max

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.