AI Models tab — see every server, and what's driving them
- Detect every recognised AI server regardless of live GPU usage — servers no longer vanish when their model unloads; each shows Loaded or Idle.
- ~20 server types now recognised: Ollama, vLLM, llama.cpp, LocalAI, HF TGI/TEI, faster-whisper/Speaches, koboldcpp, tabbyAPI, text-generation-webui, LM Studio, xinference, Aphrodite, Infinity, InvokeAI, Stable Diffusion (A1111/Forge/SD.Next), ComfyUI.
- Caller attribution — a new "Driven by" breakdown shows which services are calling each model server (e.g. Ollama), split by connection-time. Sampled, so long LLM streams are tracked reliably; sub-second calls are approximate.
- ComfyUI lists real checkpoint names.
- Oversized idle catalogues (faster-whisper exposes 400+ registry models) collapse to a single summary row.
- Fix: Peak·Avg average bar was invisible (invalid
hsl()+alpha) — now renders.
🤖 Generated with Claude Code