github murtaza-nasir/speakr v0.10.13-alpha

2 hours ago

Speakr v0.10.13-alpha

In this release, an administrator can limit how much storage each user keeps, the Shared Transcripts list includes recordings shared with other users and with you, and right-to-left text is shown in its own direction. Speaker counts set in the environment are sent to WhisperX, and a local language model can be used without an API key.

Added

  • Storage quotas (discussion #413). An administrator can set a storage quota in MB for each user in Add User and Edit User, and a default for new accounts in System Settings (default_storage_quota_mb). Users without a quota have no limit; existing accounts have none after the upgrade. Storage is the size of the audio and video a user keeps: archived recordings and recordings protected from auto-deletion count, but recordings with their audio removed do not.
    • An upload that would exceed the quota is refused with the space used and the size of the file. The same check is made for uploads from the browser, the API, the share target, the ASR Voice Recorder app and sliced uploads.
    • A recording cannot be started in the app once the quota is reached. A recording in progress is always kept, even if the quota is exceeded.
    • A file in the watch folder that does not fit is left in place, and the user receives one notification. The file is added once there is room.
    • Usage appears in the admin users table, in Edit User and in the System Statistics users table. Users see it under Account and, when a quota applies, as a second meter in the header. Thanks to @devkit415 for the suggestion.
  • Keeping the originals of a merge without their audio. When merging, you choose whether to keep the original recordings, keep them without their audio, or delete them. Without their audio, their transcripts, summaries and notes are kept and the space is freed. The audio is removed only after the merged recording is saved. When there is not enough storage for a second copy, this is explained in the merge dialog and the option is preselected.
  • Usage and Limits on the Account page. AI tokens and transcription minutes used this month, and storage, are each shown against their limit when one is set.

Changed

  • Shared Transcripts (#416). The list was created when Speakr had only public links. It now has three sections: your public links, the recordings you have shared with other users and groups with each person's permissions, and the recordings shared with you. A manual share can be revoked from the list; a share made through a group tag or folder lasts until the tag or folder is removed. On wide screens, each entry is a single row. Thanks to @devkit415 for the report.
  • Identify from conversation. The "Auto Identify" button in the speaker dialog is now called "Identify from conversation". Names mentioned in the conversation are found with the text model; voice matches are suggested without it. The button is disabled when no text model is configured.
  • Account page layout. The account statistics are one row across the page, with personal details on the left and usage and sign-in on the right.

Fixed

  • Right-to-left text (#414). Arabic, Hebrew, Persian and Urdu were left-aligned. The direction of each paragraph is now set by its own text, in transcripts, summaries, notes, chat, the share page and titles, so in a recording that switches language the direction is set line by line. Thanks to @devkit415 for the report.
  • Speaker counts from the environment (#415). ASR_MIN_SPEAKERS and ASR_MAX_SPEAKERS were only read with the older USE_ASR_ENDPOINT=true setting, so they were never sent to WhisperX with ASR_BASE_URL alone, the documented setup, including for the watch folder. They are now read whenever the ASR connector is active, as the default below the upload form, tags and folders. When a tag minimum is above the environment maximum, the tag value is used for both. Thanks to @shinyh99 for the report.
  • Local language models without an API key. Every model call was refused when TEXT_MODEL_API_KEY was empty, also for a local server that needs no key, and chat was not available at all. The key is now optional when TEXT_MODEL_BASE_URL is set. Without either, no requests are sent, and "No text model is configured" is reported for each model step.

Upgrading

Pull the new image and start as usual. The database is updated at startup with one new column (user.storage_quota_mb, empty) and one new setting (default_storage_quota_mb, 0); nothing is limited until an administrator sets a quota.

  • With ASR_MIN_SPEAKERS or ASR_MAX_SPEAKERS in the environment and ASR_BASE_URL set, these values are now sent to WhisperX.
  • A chat model set with CHAT_MODEL_NAME and CHAT_MODEL_BASE_URL but no CHAT_MODEL_API_KEY was ignored, and the text model was used for chat. It is now used, without a key; the text model's key is never sent to another server.

API: GET /api/v1/users/me includes storage (used, quota and available bytes), /api/v1/capabilities lists storage_quota, and an upload over the quota is refused with 507 and the code storage_quota_exceeded. GET /api/shares/overview returns the Shared Transcripts list.

Full Changelog: v0.10.12-alpha...v0.10.13-alpha

Don't miss a new speakr release

NewReleases is sending notifications on new releases.