github ggml-org/llama.cpp b10254

latest release: b10255
2 hours ago
Details

chat : add new template for DeepSeek V4 Flash 0731 (#26398)

  • common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

  • Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
  • Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
  • Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
  • Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

  • Fix deepseek v4 0731 template selection

  • remove unneeded lower normalization

  • Fix DSML parser to consume the tool call separator

  • address aldehir requests

  • address aldehir comment

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.