github ggml-org/llama.cpp b10413

latest release: b10414
one hour ago
Details

common : auto-detect spec type from draft GGUF metadata (#26814)

  • common : auto-detect spec type from draft GGUF metadata

When -md loads a local draft model without --spec-type, the sidecar
inference in common_models_handler_apply only checks HF repo sidecars
and misses local files. The draft model loads into VRAM but speculative
decoding never activates (types stays NONE).

Read general.architecture from the draft GGUF header and map:
dflash + markov_w1.weight tensor -> draft-dspark
dflash without markov head -> draft-dflash

Assisted-by: opencode

  • common : address review feedback on spec-type auto-detect PR
  • Fix comment spacing to match surrounding style (/* .x = / not /.x =*/)
  • Add LOG_INF when auto-detection fires so users can see why spec decoding enabled
  • Document single-file assumption for split-GGUF edge case

Addresses bot review feedback on #26814.

  • common : move spec-type GGUF auto-detect into speculative module
  • add common_speculative_types_from_gguf() in speculative.cpp/.h
  • use gguf_context_ptr (RAII) from ggml-cpp.h
  • reduce comments to a single line per AGENTS.md style

Addresses review feedback on #26814

  • common : add doc note and join SPC_INF line in spec-type auto-detect

Assisted-by: opencode

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.