github HKUDS/DeepTutor v1.6.13

2 hours ago

DeepTutor v1.6.13 Release Notes

Release Date: 2026.10.04

Building on v1.6.12, v1.6.13 expands reading, listening, and question-bank practice, improves document parsing and speech configuration, and brings Polish plus more model providers. Chat attachments now follow the selected workspace, with existing uploads and URLs preserved during the transition.

What's New

Read, listen, and revisit

Immersive Reading gains natural read-aloud audio and quiz star rewards. Immersive Watching gains timestamped marks, notes, and suggested review points. PDF resizing and EPUB chapter navigation preserve reading position more reliably; learners can open explicitly assigned books and reading materials within their granted access.

Practice from your question bank

Start practice directly from the Question Bank and select multiple saved questions for a chat. Guided Learning shuffles multiple-choice options and avoids exposing correct answers through option letters or retrieval prompts. Co-Writer lets you choose the model for an edit.

More reliable document retrieval

Choose an independent vision model for document image descriptions; successful descriptions are cached per workspace, with optional bounded batching. MinerU cloud parsing slices oversized PDFs and resumes completed parts, while local failures report actionable causes. Optional tiny-scan normalization, numbered-exercise lookup, and better figure retention improve textbook retrieval. deeptutor kb eval measures retrieval quality against a local evaluation dataset.

Speech settings that explain failures

Speech models have a configurable synthesis timeout used by previews and playback. DashScope Qwen-Audio and Qwen3 TTS use their respective endpoints, voices, formats, and regional requirements. Preview errors distinguish authentication, permissions, parameters, quota, network failures, and timeouts. Reply playback preserves math, cancels stale requests, and prevents overlapping audio.

Workspaces and connections

New chat uploads live in the selected workspace's chat/attachments directory so workspace tools can find them. Earlier uploads move there when their conversation is used or its data is migrated; existing attachment URLs remain readable. SearXNG connection tests explain Docker addressing, disabled JSON output, and empty search results. Web sources can pair bilingual content, and knowledge-base checks reuse probes without blocking the API.

Partners, languages, and providers

Export Partner group conversations as Markdown and configure Telegram response policies per chat. Feishu delivery and Telegram code blocks are more reliable. Polish joins the interface languages; Requesty, API Route, and FutureInfra join LLM providers, with native MiniMax speech and Xiaomi MiMo v2.5 preset speech support.

Chat recovery and everyday fixes

History search checks full conversation transcripts and excludes recycled sessions. Long conversation branches, reconnecting submissions, LaTeX selections, and learner routes recover more reliably. Image payloads no longer inflate token estimates, GLM reasoning options match supported settings, and deeptutor doctor honors the configured provider and token budget.

Updated setup and feature documentation accompanies the release at deeptutor.info.

Upgrade Notes

Run pip install -U deeptutor; Docker users pull ghcr.io/hkuds/deeptutor:latest.

  • Attachments: no manual move is required. Earlier uploads are relocated into the selected workspace when needed; keep existing application data mounted during the upgrade.
  • Speech: set Request timeout (seconds) in the speech model when synthesis needs more time; the default is 60 seconds, with a supported range of 5–600 seconds. Use Preview voice to test synthesis; fetching a model list alone does not validate a voice.
  • DashScope Qwen-Audio TTS: use a Beijing API key and a voice supported by the selected model. A provider API base URL selects the matching speech endpoint automatically; Qwen3 TTS uses its own voices and WAV output.
  • Document images: select an image-description model in Document Parsing settings if the main model cannot inspect images. Batching and tiny-scan normalization are optional; older indexed content needs re-parsing and reindexing to gain newly generated descriptions or figure records.
  • Learner books: administrators must grant the Books surface and explicitly assign a book; learner access stays scoped to those assignments.
  • systemd installs: stop the service, upgrade DeepTutor in its Python environment, and restart the service.

Full Changelog: v1.6.12...v1.6.13

Don't miss a new DeepTutor release

NewReleases is sending notifications on new releases.