github kizuna-ai-lab/sokuji v0.42.1

3 hours ago

Highlight: rebuilt audio

Sokuji's audio was rebuilt in this release. Every sound now goes only where it belongs: what the meeting should hear goes into the meeting, and what is meant for you plays only on your own speakers or headphones.

  • Replay stays with you: replaying a message plays it only on your speakers or headphones. Before, the meeting heard it a second time.
  • No more delayed echo of your own voice: with Original Audio Passthrough and the speaker monitor both on, you used to hear your own voice in your headphones a moment late. Your original voice now goes only into the meeting.
  • Voice previews play on the speakers or headphones you chose, not on the system default (KizunaAI, Soniox and Doubao AST 2.0).
  • 0% really is silent: Original Audio Passthrough at 0% sends nothing into the meeting. It used to play at 30%.

New

  • Doubao AST 2.0 voices: the translation can now be spoken in a voice from Doubao's library of 506 voices — filter by gender, age and style, and preview one before you use it — or, where the language pair supports it, in a clone of the speaker's own voice as before. The voice is chosen per target language (#577, #586).
  • Doubao AST 2.0 understands more languages: 20 source languages plus two Chinese dialects, and spoken translation from Chinese or English now also reaches Korean (#586).
  • Doubao AST 2.0 accepts an API key as well as an App ID and Access Token.
  • Push-to-Talk and Push-to-Translate on every provider: Speech Mode is now one setting for the whole app, and Soniox, KizunaAI, Palabra AI, OpenAI Translate and OpenAI Live gain both modes.
  • Hold to talk in the meeting overlay: the extension's in-meeting subtitle overlay has a Hold button for Push-to-Talk.
  • Max Speech Duration for Free (on-device) translation: a slider (10–40 s, default 30 s) sets how long one utterance can run before it is split. Some speech models split sooner, at the longest stretch they transcribe correctly.

Improvements

  • Languages: language names are shown in your UI language, one language pair is shared by every provider and kept when you switch provider, and your current pair is listed first (#582).
  • Free (on-device) translation covers 74 languages, each with both a speech-recognition model and a voice; 22 languages are new. 24 of them are voiced through Edge TTS and need a network connection (#583).
  • Chinese, Japanese and Korean text uses a matching sans-serif font and never falls back to SimSun; the meeting overlay no longer shows a serif font (#559, #581).
  • Changing the language pair or the audio mode no longer re-checks your API key, so it sends no request to the provider; a key you type is checked shortly after you stop typing (#585).
  • OpenAI Translate and Gemini Live Translate: each translation is shown beside the speech it translates.
  • Gemini: Gemini 3.x dialogue models keep what you say while Gemini is still answering instead of dropping it, a dropped connection resumes, and new setups start on Live Translate.
  • Clearer error messages, in every UI language.
  • Voice library rows show each voice's gender, age and style in your UI language, Soniox's included (#586).

Bug fixes

  • Whisper (WebGPU) no longer fails on a language it has no entry for (#582).

Removed

  • Zoom AI Services and Volcengine Speech Translate (#553).
  • OpenAI Compatible (Legacy Realtime), including the CometAPI preset (desktop app only).
  • OpenAI Realtime's WebRTC connection and its temperature setting; OpenAI Realtime always connects over WebSocket.

Note — after updating

  • Pick your languages again: the language pair is reset once to the selected provider's default.
  • If you used a provider that was removed, the app opens on KizunaAI. Your old settings stay on disk but are no longer used.
  • Speech Mode is taken once from the provider you had selected. OpenAI Realtime's former "Disabled" option (its push-to-talk) becomes Auto; choose Push-to-Talk again if you used it.
  • OpenAI Realtime: a saved model that is no longer listed runs as the default model, gpt-realtime-2.1-mini.
  • Palabra AI: a setup with app credentials may open in API-key mode; switch it back once. Timbre detection is off.
  • Exports: the .txt file has one block per exchange (the source, then → translation), the models line reads asr= / translation= / tts= for every provider, and the .json file uses a new format, sokuji-conversation/2.
  • Word-by-word highlighting during playback is drawn only where the provider reports timing.
  • Free (on-device) translation no longer offers Punjabi, Xhosa and Zulu, which lack a matching voice or speech model.

Full Changelog: v0.41.1...v0.42.1

Don't miss a new sokuji release

NewReleases is sending notifications on new releases.