Speakr v0.10.7-alpha
This release changes what "Archived" means. Archiving is now something you do: it takes a recording out of the main list and deletes nothing. The state that used to carry that name, a recording whose audio was deleted and whose transcript was kept, is now called Audio removed, and you can put a recording into it yourself. The release also lets several files be joined into one recording at upload, reworks how voice profiles are built and matched, rebuilds the Identify Speakers dialog, makes the AI title instructions configurable at every level the summary prompt is, and fixes a dependency change that would have broken every PostgreSQL installation on the next image build.
Archive, and "Audio removed"
Archive hides a recording from the main list and keeps everything: audio, transcript, summary, notes, tags and shares (#394, requested by @setup562). The archive icon sits next to inbox and star at the top of an open recording, and the selection bar archives several recordings at once. Search still finds archived recordings, marked with a small archive badge, and the Archived toggle below the search bar shows only archived recordings, where the same icon moves one back. Like inbox and star, archiving is personal: archiving a recording someone shared with you hides it from your list only.
Audio removed is the new name for what Speakr previously labelled "Archived": the state audio-only retention (DELETION_MODE=audio_only) leaves behind, with the media file deleted and the transcript, summary and notes kept. Nothing about that state has changed except its name. It is found through a new Audio removed quick filter next to Starred and Inbox, which replaces the old "Archived" toggle.
You can now also remove the audio of a single recording by hand. Delete audio, keep transcript, in the recording's new ⋯ menu, deletes the stored media and leaves the recording in the Audio removed state. It needs the same permission as deleting the recording. For a recording kept as video, the video file is its only stored media, so the action reads Delete video, keep transcript and removes the video together with its sound.
If you relied on the old "Archived" toggle to find recordings whose audio retention had removed, use the Audio removed filter instead. API v1 keeps returning archived and unarchived recordings by default, so existing integrations see no change; is_archived is now part of each recording, settable through PATCH and batch updates and personal to the caller as in the web app, archived=true|false filters the list, and POST /api/v1/recordings/{id}/delete-audio removes the media. The internal endpoint /api/recordings/archived keeps its old meaning, recordings with audio removed, and is also available as /api/recordings/audio-removed.
A reorganised recording header
With two more actions, the icons at the top of a recording needed grouping. They now run: inbox, star and archive; folder and tags; a Reprocess menu (transcription, summary, and reset for a stuck recording) and Identify Speakers; then share and a ⋯ menu holding Delete audio and Delete recording, so the destructive actions are one click further away. Reprocessing the transcription is no longer offered once a recording's audio is removed, since there is nothing left to transcribe. On phones, where the toolbar scrolls, the same actions stay as individual buttons in the same order.
The same header on every page
The Account, Admin, Group Management and Inquire pages each had their own header, with different menus: some had notifications, some a separate dark mode button, and team administrators reached different pages depending on where they opened the menu. Every page now uses the main view's header: the token budget, Inquire, New Recording and the same user menu with notifications, settings, shared transcripts, administration, help, colour scheme and sign out. Entries that open a dialog of the main view, such as New Recording or Shared Transcripts, take you there with the dialog open.
Pages also no longer flash a loading screen. The loading overlay used a fixed dark colour and appeared on every page load, so a page that loaded in a quarter of a second flashed a differently coloured spinner first. It now uses your colour scheme and appears only when a page takes longer than 0.4 seconds to load.
The Identify Speakers dialog, rebuilt
The dialog for naming speakers was rebuilt around recordings with many detected speakers, following suggestions from @checkmeck (#395).
- Speakers are sorted by speaking time by default, each with its time and share of the recording. Sorting by first appearance or by name is one click away, and the choice is remembered.
- Speakers with under 30 seconds of speech, typical of diarization in a large room, are grouped at the end and can be merged into another speaker individually or all at once. A merge is applied when you save and undone by Cancel.
- Selecting a speaker shows only their segments. "Show all segments" keeps them highlighted with Previous and Next buttons to step through their turns.
- The name field suggests your saved speakers as soon as you click into it, voice matches first, supports the arrow keys, Enter and Escape, and points out when a name differs from a saved one only in capitalisation.
- Each speaker has a button that plays their longest segment, segments show their timestamps, and the one playing is highlighted.
- The volume slider no longer spills out of its popover.
Two problems behind the dialog are fixed as well. Saving with a line edit pending used a second code path that renamed the speakers but skipped updating voice profiles and speaker snippets, so those recordings never trained anyone's profile. Both paths now share one implementation. And names that differ only in case, such as "john" and "John", created two separate speaker profiles; they now resolve to the saved spelling.
Configurable AI titles
The instructions for AI-generated titles were fixed in code: at most eight words, no filler, only the main topic. They are now configurable the same way the summary prompt is, with the same precedence (#400, requested by @stanjourdan):
- a tag's title prompt (several tags with title prompts are combined, in the order the tags were added)
- the folder's title prompt
- the user's own title prompt, under Account, Prompt Options
- the administrator's default, in the admin Default Prompts tab
- the built-in default, unchanged from before
Group tags and folders carry a title prompt too, and API v1 exposes title_prompt on tags and folders. Only the instructions are editable: the transcript, the output language and the rule that the model replies with the title alone are still added by Speakr, after the transcript, so a custom title prompt does not affect prefix caching. A naming template's {{ai_title}} placeholder receives whatever these instructions produce.
Several files into one recording
A recorder or phone often splits a long meeting into several files. Once two or more files are in the upload list, a switch above it now offers One recording next to the usual Separate recordings. The files are numbered in the order they will be joined, sorted by modified time and then by name, so part2 comes before part10, and they can be reordered by dragging or with the arrows. The dialog shows the total length and takes an optional title. Each file is uploaded and checked as usual; when the last one arrives, Speakr joins the audio and transcribes it once, with the folder, tags and transcription options chosen in the dialog. Until now this required uploading the files as separate recordings, waiting for each to be transcribed, and merging them, which transcribed the same audio twice.
Files uploaded for a join are held apart from recordings until the join happens, so they never appear in lists or search. If one file fails to upload, the others are discarded, and parts of a join that never completes are removed after a day. Up to 20 files can be joined.
Voice matching, rebuilt
A voice profile used to be a single average that every named recording nudged by 30%. That had three costs: one wrongly named speaker shifted the profile for good, the same person on a phone and in a meeting room blurred into one vector that matched neither well, and nothing tied a profile to the embedding model that produced it.
Profiles are now built from samples. Each time you name someone, Speakr keeps that recording's voice embedding for them, up to 20 per person. Samples that resemble each other form a voice variant, and a new recording is compared with every variant, the closest one counting. Because the variants are rebuilt from the samples every time, corrections are complete: renaming a speaker later moves their sample to the right person, even after the transcript already shows names, and the new Voice profile panel under Account, Speakers Management lists every sample with its recording and lets you remove one.
Profiles are kept per embedding model. Speakr already identified the backend's embedding model at startup by embedding a bundled clip. Each model it recognises now gets its own voice space, and embeddings are only compared within one. Switching from WhisperX to OpenASR and back keeps each backend's profiles, instead of mixing vectors that cannot be compared. Profiles from before the upgrade belong to the model they were built with and keep matching exactly as before.
Matching is more careful.
- A speaker with less than 15 seconds of speech in a recording neither trains a profile nor gets auto-labelled (
VOICE_PROFILE_MIN_SPEECH_SECONDS). - A sample unlike everything a well-established profile holds is kept out and logged as a likely wrong name.
- Auto-labelling gives each person to at most one speaker per recording.
- Auto-labelled samples count half as much as confirmed names until you confirm them.
- Embeddings of any size are accepted everywhere. Suggestions and profile updates had required 256 dimensions, so a backend returning another size would have stopped learning without any message.
Thresholds calibrate themselves. Once a voice space has enough samples, Speakr compares each sample with the same person's other samples and with other people's, and sets the match threshold between the two. Suggestions and the low, medium and high auto-labelling settings follow it; the admin Voice embedding compatibility card lists each voice space with its sample count and threshold.
Existing profiles need nothing: they are read as one sample weighted by the recordings that built them, and written out as a sample the first time the person is named again.
Voice profiles through OpenASR
The OpenASR connector can now request per-speaker embeddings, so voice profiles work with it the same way they do with WhisperX. Set ASR_RETURN_SPEAKER_EMBEDDINGS=true with an OpenASR server from openasr#379 onward; it is off by default because older servers reject the extra field (#380).
Also fixed
- PostgreSQL connections after the next build. SQLAlchemy 2.1, released on 24 September, made psycopg 3 the default driver for
postgresql://URLs. Speakr ships psycopg2, so any fresh install or image build would have failed every database connection. SQLAlchemy is now pinned below 2.1. Reported and fixed by @jjsmackay (#401). - Installed desktop app. The web app manifest requested a window-controls overlay that Speakr never implemented, so an installed desktop Chromium app could not be dragged and its window controls covered the header icons. Contributed by @jjsmackay (#402).
- Actions failing behind some reverse proxies. Over HTTPS, Speakr checks that the browser's Referer header matches the host it receives. A proxy that strips the Referer, or forwards a different Host, made every action fail. The page then refreshed its security token and retried, which could never help, and showed a generic error while the server logged nothing. The server now replies with the actual reason, naming the two hosts when they differ, and logs a warning. The troubleshooting guide covers the two proxy settings involved (#388).
- Deleting a recording through API v1 left its media behind. The endpoint built a file path by hand that never matches how recordings are stored now, locally or in S3. Single and bulk deletes in the web app, API v1 single and batch deletes, merging with the originals removed, and retention now share one deletion routine, so each removes the media, snippets and processing jobs and sends the
recording.deletedwebhook. - Long transcriptions on PostgreSQL. The worker kept a database transaction open, but idle, for the whole transcription call and for the summary and title requests, so on a multi-hour recording PostgreSQL's
idle_in_transaction_session_timeoutcould drop the connection and stall the queue. The transaction now ends before each of those calls and a new one stores the result. Reported by @JoeBoll (#405). - Empty transcripts counted as success. A transcription service that finished without returning any text left the recording marked Completed, so a failed job looked successful. Such a recording is now marked Failed with an explanation, and the Mossland connector no longer reports an empty result as finished. Reported by @JoeBoll (#406).
- Error messages with a literal
{status}. Several error and confirmation messages used a placeholder syntax the interface never fills in, so users saw "(HTTP {status})" instead of the status code. A test now checks every translation for it. - The Account page could stay on "Loading Speakr...". A timer that shows the loading overlay could fire after the page had already finished loading and removed it, covering the page for good. Once the overlay is hidden it now stays hidden.
Upgrading
Pull the new image and start as usual. New tables (speaker_voice_sample, voice_embedding_space, upload_join_part) and columns (title_prompt on users, tags and folders, is_archived on recordings and shared recording states, and two voice-matching columns on recordings) are created automatically, and the admin default title prompt is seeded with the built-in instructions. The images now bundle FFmpeg 8.1.3, up from 8.1.2. No configuration changes are required. I would still recommend taking a copy of transcriptions.db before upgrading, as with any release.
Full Changelog: v0.10.6-alpha...v0.10.7-alpha