github VocaHQ/vocalinux v0.18.0

4 hours ago

Vocalinux v0.18.0

Vocalinux 0.18.0 is a minor on the stable line. Default engine is still whisper.cpp. This release leans hard into Wayland: an in-app Dictation Pad that never touches text injection, a PipeWire capture path for microphone and system audio, and a RemoteDesktop portal injection backend that works without ydotoold.

The headline fix: the dictation hotkey is now grabbed and suppressed, so the shortcut key stops leaking into the app under it, on every desktop. A heavy round of community issues and pull requests lands here as well.

AppImage, Flatpak, and Snap still ship whisper.cpp only.

Highlights

Area Description
Dictation Pad In-app window receives dictation; copy text out by hand; nothing fragile about Wayland can break it
Hotkey suppression evdev grabs the dictation shortcut and forwards everything else through a uinput clone; also works when /dev/uinput is write-only
Wayland RemoteDesktop portal injection backend; Alt+Shift / Win+Space layout switching keeps working on GNOME Wayland
Audio PipeWire capture for mic and system-audio sources; optional lowering of other audio while dictating
Dictation A shortcut per language, follow the active keyboard layout, tray history of recent dictations, floating glow overlay
Transcription --transcribe-file with per-speaker labels via TinyDiarize, plus a transcript viewer
Extensibility Opt-in D-Bus activation for compositor global shortcuts and scripts, custom dictionary with corrections, postprocessing script hook
Packaging vocalinux-bin AUR package ships the AppImage; release workflow can now publish a signed self-hosted Flatpak remote; gated snap-promote workflow

New Features

  • Dictation Pad: dictate into the app's own window, review, and copy the text out yourself. The bulletproof path for compositors where injection is fragile. Enable it in Settings or open it from the tray menu. (#887; fixes #726 reported by @hopsayer)
  • Per-language dictation shortcuts, push-to-talk race fixes, and dictation history: hold a different hotkey per language, and reopen recent dictations from the tray menu. (#880; fixes #805 reported by @krisek)
  • Follow the active keyboard layout while dictating, so recognition language tracks the layout you are typing in. (#837 by @viveknadig; fixes #821)
  • Recent dictation snippets in the tray menu. (#487 by @LuigiKraken)
  • Floating glowing dictation overlay that shows listening/processing state. (#516)
  • Lower other audio while dictating so the mic hears you, not your music. (#861)
  • PipeWire capture path: native PipeWire streams for microphone and system-audio sources, with the capture source factored out of the recognition loop. (#889, #883; fixes #751 and #760 reported by @kacperpaczos)
  • RemoteDesktop portal text-injection backend for Wayland, no ydotool daemon required. (#885; fixes #750 reported by @kacperpaczos)
  • File transcription with TinyDiarize per-speaker labels: vocalinux --transcribe-file audio.wav or from the tray, with a speaker-tagged transcript viewer and export. (#884; fixes #756 reported by @kacperpaczos)
  • Opt-in D-Bus activation so compositor-level global shortcuts (and your own scripts) can start and stop dictation. (#568 by @webenefits; fixes #761 reported by @kacperpaczos)
  • Combined custom dictionary: terms bias the recognizer and corrections fix the transcript, from one file. (#890)
  • Postprocessing script hook: pipe transcriptions through any user command before they are injected. (#479 by @karottenreibe)
  • Bilingual dictation: configurable second-language Whisper candidates, with deferred settings edits so unsaved changes do not leak. (#424 by @juanfradb)
  • Optional verified Orukeet model for the Parakeet engine. (#840 by @Nathan-Roll1)
  • Keep recording while an idle-unloaded model reloads instead of dropping the start of the utterance. (#851 by @mre31)
  • config.json can pin the text-injection backend (force wtype, ydotool, portal, etc.). (#649 by @HashimAbdulaziz; fixes #476 reported by @waldemar-p)
  • Run a local VocaGateway from Settings -> Advanced (podman-first). (#774)
  • Settings category list scrolls, so every page stays reachable on small windows. (#886; fixes #678 reported by @hopsayer)

Bug Fixes

Hotkeys and injection

  • The dictation shortcut no longer leaks into the focused application: evdev grabs the keyboard and forwards all other keys through a uinput clone. (#873; fixes #871 reported by @bernt-matthias)
  • Clone creation now retries on a write-only /dev/uinput, which is what the 0.17.0 installer's udev rule produced; without this the grab silently failed and the key leaked again. (#893)
  • Alt+Shift and Win+Space layout switching keeps working on GNOME Wayland instead of being swallowed. (#876; fixes #848 reported by @hopsayer)
  • The IBus guard no longer flips GNOME/X11 sessions to a US layout. (#827 by @AmirF194)
  • Shortcut recorder learns unmapped F19/F24 and XF86-aliased F13-F23 keys. (#844; fixes #843 reported by @bisgardo)

Settings, downloads, and models

  • Cancelling a model download aborts blocked fetches instead of waiting on a stalled server. (#888; fixes #679 reported by @hopsayer)
  • The model download dialog shows "verifying" after the last byte instead of reading finished while it is still checking the file. (#864 by @guilhermefeitosa66; fixes #863 reported by @guilhermefeitosa66)
  • Vosk refuses to load a model that would exceed available memory instead of tripping the OOM killer. (#850 by @AmirF194; fixes #676 reported by @hopsayer)
  • Test Dictation textbox: the scrolled-window stripe artifact is gone, and the box shows recognized text again. (#853; fixes #847 and #720 reported by @hopsayer)
  • Dictation Tone grays out while Enable Sound Effects is off. (#877; fixes #849 reported by @hopsayer)
  • Speech Model follow-ups from post-0.17.0 QA. (#836; fixes #834)
  • Stop passing an explicit device index when zero devices enumerate. (#891)
  • The View Logs dialog closes via the titlebar X. (#892)
  • Settings dialog internals consolidated: guard flags and widget construction each live in one place. (#878; fixes #793 reported by @kacperpaczos)

Updates and site

  • The update checker falls back gracefully when the GitHub API rate-limits instead of showing an error. (#846; fixes #845 reported by @lmstud)
  • vocalinux.com: Phone chip, "Free forever" Cost copy, and a privacy-page GA disclosure fix. (#868; fixes #867)

Installer, packaging, and CI

  • The installer now installs hash-pinned dependencies and build tools, and is split into sourced modules with a generated distro package map. (#856, #862, #872 by @sesav)
  • Snap gains hardware-observe so hotkeys can read input devices. (#858; fixes #857 reported by @RhysU)
  • New vocalinux-bin AUR package ships the AppImage with per-arch digests, published by the release workflow. (#879; fixes #817 reported by @hopsayer)
  • Release workflow can now publish a signed self-hosted Flatpak OSTree remote, so installs can update in place once the remote is online. (#875; fixes #785)
  • A snap-promote workflow gates stable promotion behind candidate QA. (#881; fixes #783)
  • Release verification now works even when a run fails after publishing; the wtype fallback tests pin the ydotoold probe. (#842, #841 by @sesav)
  • Dependency advisories cleared and workflow tokens scoped to repository reads; GitHub Actions group bumped. (#869, #865, #866 by @Mr-Sunglasses; #870)
  • Docs: VocaMac marked available (v1.0.0); README ecosystem/SmartScreen notes cleaned up. (#859 by @Mr-Sunglasses; #835)

Thanks

This release exists because people filed careful issues and sent real patches. Thank you all.

Reporters: @hopsayer for nine issues that shaped the release: Dictation Pad (#726), the scrollable settings sidebar (#678), cancellable downloads (#679), the Vosk OOM (#676), Test Dictation fixes (#720, #847), GNOME Wayland layout switching (#848), the Dictation Tone gray-out (#849), and the vocalinux-bin AUR package (#817). @kacperpaczos for the Wayland and PipeWire foundations (#750, #751, #756, #760, #761) and the settings-consolidation nudge (#793). @bernt-matthias for #871, which is why the hotkey finally stays out of your editor. @krisek for per-language shortcuts (#805), @waldemar-p for the injection-backend pin (#476), @bisgardo for F19 recording (#843), @lmstud for the update-checker rate limit (#845), @RhysU for snap hardware-observe (#857), and @guilhermefeitosa66 for both reporting and fixing the download-dialog finish state (#863, #864).

Patch authors: @karottenreibe for the postprocessing script hook (#479), @webenefits for D-Bus activation (#568), @juanfradb for bilingual dictation (#424), @LuigiKraken for tray dictation history (#487), @viveknadig for layout-following language (#837), @Nathan-Roll1 for the Orukeet model (#840), @mre31 for recording through a model reload (#851), @HashimAbdulaziz for the injection-backend config (#649), @AmirF194 for the IBus layout fix and the Vosk OOM guard (#827, #850), @guilhermefeitosa66 for the download dialog fix (#864), and @sesav for the installer module split, hash-pinned deps, package-map generation, release-verify resilience, and the ydotoold test pin (#856, #862, #872, #842, #841), and @Mr-Sunglasses for dependency advisories, workflow token scoping, and the VocaMac docs (#869, #865, #866, #859).


Install / Upgrade

Recommended (install.sh)

curl -fsSL https://raw.githubusercontent.com/VocaHQ/vocalinux/main/install.sh -o /tmp/vl.sh && bash /tmp/vl.sh

Optional extras: --engine=faster_whisper or --engine=parakeet. Default remains whisper.cpp. The installer needs Python 3.11 or newer and distro python3-gi / python3-gobject.

AppImage

Download Vocalinux-0.18.0-x86_64.AppImage or Vocalinux-0.18.0-aarch64.AppImage from this release (whisper.cpp path; does not bundle Faster Whisper or Parakeet).

chmod +x Vocalinux-0.18.0-x86_64.AppImage
./Vocalinux-0.18.0-x86_64.AppImage

Host text-injection tools (xdotool, wtype, or ydotool) are still required for app-into-app typing; the Dictation Pad needs none of them. FUSE is needed to run the AppImage on most hosts.

Arch Linux (AUR)

yay -Syu vocalinux        # source build
yay -Syu vocalinux-bin    # the AppImage, packaged

Both packages are published from this tag. Give the Publish to AUR jobs a few minutes after the GitHub Release appears, then rebuild.

PyPI

pip install -U vocalinux

PyPI is the Python package only. For system deps, desktop integration, and models, use install.sh.

Snap

vocalinux_0.18.0_amd64.snap is attached to this release. The Store still has not listed it: review-tools stopped the upload with allow-installation on the uinput plug. latest/edge is still 0.16.2 rev 7, which has no uinput plug, so sudo snap connect vocalinux:uinput fails on the Store snap.

Sideload the amd64 file from this release:

sudo snap install --dangerous ./vocalinux_0.18.0_amd64.snap
sudo snap connect vocalinux:audio-record   # if mic is not auto-connected
sudo snap connect vocalinux:raw-input      # global keyboard shortcuts
sudo snap connect vocalinux:uinput         # native Wayland typing

--dangerous is required because this file is not a Store revision. It will not refresh from the Store. After Canonical grants uinput, sudo snap install vocalinux --edge (or snap refresh) is the updating path. stable is still a human promote.

Store page: https://snapcraft.io/vocalinux

Flatpak

Download Vocalinux-0.18.0-x86_64.flatpak or Vocalinux-0.18.0-aarch64.flatpak from this release (whisper.cpp only; not on Flathub).

flatpak remote-add --if-not-exists flathub https://dl.flathub.org/repo/flathub.flatpakrepo
flatpak install flathub org.gnome.Platform//50
flatpak install --user ./Vocalinux-0.18.0-x86_64.flatpak
flatpak run com.vocalinux.Vocalinux

The release workflow can now publish a signed self-hosted OSTree remote alongside the bundles. Once the remote is online, Flatpak installs can update in place instead of re-downloading a .flatpak each tag.

Upgrade details: docs/UPDATE.md

Full changelog: v0.17.0...v0.18.0


Verifying what you downloaded

SHA256SUMS covers every file in this release (wheel, sdist, both AppImages, both Flatpaks, the amd64 snap):

sha256sum -c --ignore-missing SHA256SUMS

The wheel and the sdist are the same bytes published to PyPI.

Don't miss a new vocalinux release

NewReleases is sending notifications on new releases.