github ggml-org/llama.cpp b11557

latest release: b11558
pre-releaseone hour ago
Details

rpc : turn GGML_RPC_DEBUG into a verbosity level and add logs (#29544)

  • rpc : turn GGML_RPC_DEBUG into a verbosity level and add logs

GGML_RPC_DEBUG is now parsed as a number: 0/unset disables debug logs,
1-3 emit increasingly detailed output (events, per-command trace,
transport detail). Non-numeric values fall back to 1.

The duplicated env/macro blocks in ggml-rpc.cpp and transport.cpp are
replaced by a shared log.h, which becomes the single choke point for
all logging of the RPC backend: LOG_ERROR/LOG_WARN/LOG_INFO for
unconditional severity logs and LOG_DBG/LOG_DBG2/LOG_DBG3 for the
verbosity-gated ones. The transport files no longer need ggml-impl.h,
and the server banner now also goes through the ggml logger (stderr).

Missing logs are added on both the client (handshake, buffer ops,
tensor transfers, graph computes, cache decisions) and the server
(per-command dispatch, graph nodes), including the negotiated
transport via the new socket_t::transport_name().

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

  • rpc : align log levels with the documented verbosity semantics

The first pass introduced GGML_RPC_DEBUG levels, but many call sites did not
follow the documented semantics and some logs bypassed the macros entirely:

  • Route the remaining raw GGML_LOG_* call sites through the local LOG_* macros
    so that log.h stays the single choke point for RPC logging
  • Emit per-tensor transfers and per-command handler traces at level 2, and
    one-shot lifecycle events (backend and buffer type creation, buffer
    allocation, graph compute, graph cache eviction) at level 1
  • Drop the duplicated enqueue trace in rpc_dispatcher::send, work() already
    logs every dispatched command together with its round-trip timing
  • Promote degraded-operation events to LOG_WARN: remote allocation failure,
    RDMA falling back to TCP, unexpected peer disconnect
  • Log graph cache eviction on both sides, since client and server must stay
    in sync for the incremental graph update to be valid
  • Distinguish Unknown command (opcode out of range) from Unhandled command
    (valid opcode with no handler)
  • Fix format specifiers (%zu for size_t, %u for device ids, 0x for hex
    pointers) and a missing newline in an error log
  • Document that llama.cpp applications drop ggml debug records unless the
    verbosity threshold is raised with -lv 5

Assisted-by: pi:llama.cpp/Qwen3.8-Flash-Next

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.