github coconautilus17/LibraForge v0.2.5

3 hours ago

Upgrade notes (please read before updating)

  • Add the new libraforge-settings volume once. If you run the published
    image or docker-compose.dist.yml, everything you saved in Settings
    (publisher and title-noise patterns, the Audiobookshelf / abs-agg /
    abs-tract connections, report retention) lived inside the container and was
    lost every time you updated. Settings now live in a named volume. Download
    the new docker-compose.dist.yml, or add
    -v libraforge-settings:/app/settings -e LIBRAFORGE_SETTINGS_DIR=/app/settings
    to your docker run command. Settings saved before you add the volume are
    already gone on recreate and need to be entered once more. (Issue #261)
  • New author-name scheme: off after you update, on for new installs.
    0.2.5 adds a universal way of writing author initials (V.A. Lewis,
    J.K. Rowling: a dot after each letter, no spaces between, one space before
    the rest of the name). Because you updated an existing install it is
    switched off, so your organized library is not touched. You will see a
    one-time notice on your first page load after updating; you can turn the
    scheme on or off, and add your own exceptions, under Settings, Author
    names
    . If you turn it on, folders and files already in your library are
    renamed only when you choose to run the opt-in migration script (dry run by
    default, revertible). Back up your Audiobookshelf database first: renamed
    authors show up as new author entries there. (Issue #262)

New Features

  • Author names: universal initials scheme (opt-in): author initials are
    always written X. per letter with no spaces between them and one space
    before the rest of the name, so V A Lewis, J.K.Rowling and JK Rowling
    become V.A. Lewis and J.K. Rowling, and a lone initial gets its dot
    (Kevin J. Anderson). It applies to the author tags and metadata.json
    that the fixer and Manual Review write; Folder Forge does not apply it, it
    names author folders from that metadata. On for fresh installs, off when an existing install is updated
    (decided once, on the first start of 0.2.5), with a one-time notice
    explaining the change. Settings, Author names has the switch, a written
    explanation of the rule, Patterns in use (shipped exceptions such as
    names where the letters are not initials, each can be switched off) and
    Custom patterns (your own, matched regardless of spacing and
    punctuation). An opt-in scripts/normalize-author-names.py brings existing
    folders, sidecars and (with --tags) embedded tags in line: dry run by
    default, change log, --revert, and it refuses to run while the scheme is
    off. (Issue #262)

  • Settings: one Patterns in use / Custom patterns layout: Publishers, Title noise
    and Author names now share one renderer and the same two lists: known
    patterns ship with LibraForge and can each be switched off, private
    patterns are yours. Every change is saved immediately (no Save button).
    The author card has one field for a private name. Publishers that earlier
    runs learned (Aethon Audio, Mountaindale Press, Hidden Infinity and 15
    more) now ship as known publishers, and names with a digit, such as
    Comedian0 L, are never read as initials.

  • Settings: Folder Forge scan root: Settings, Library has a new
    Folder Forge scan root field. Folder Forge starts with it instead of a
    hard-coded folder, and the field's note shows the built-in default it uses
    when the field is empty (/audiobooks/_unorganized). It can still be changed
    for a single run on the Folder Forge page.

  • Match report: Initials Fixed: when a run rewrites an author only
    because the author-name scheme unified its initials, the book gets an
    Initials Fixed badge in the full match report, a summary tile, and an
    Author Initials Fixed status filter that lists them all.

  • Chapter Forge — a new feature that detects spoken chapter markers in
    single-file audiobooks (e.g. one long MP3 with no embedded chapter table),
    so they can get real chapters the way a multi-file book already would.
    Three detection backends: Hybrid (silence detection + speech-to-text
    transcription, with an optional LLM review pass), Full transcription
    (pure ASR), and Audible chapters (pulls the official chapter list
    straight from Audible when available). Detected chapters are reviewable
    and editable against their transcript/silence evidence before saving, with
    per-chapter confidence, sequence-gap detection, and a run log. Any two
    detection methods can be run back-to-back and diffed side by side, or
    compared directly against Audible's official chapter list. Saved chapters
    live in the same libraforge.json sidecar Fixer/Manual Review already
    use, and M4B Tool can now embed them when merging a book into an M4B.
    Hybrid/Full transcription's ASR models are opt-in (heavy, ~825MB) and
    require building with Dockerfile.unified instead of the default image;
    Audible chapters works out of the box with no extra dependencies.
    When the image doesn't include them, Hybrid and Full transcription are
    shown as unavailable with a notice explaining how to enable them.

  • Folder Forge: {subtitle} naming token — subtitles are now extracted
    from sidecar/marker/tags and available as a naming-template token, treated
    as decoration (not core) alongside narrator/publisher/year/asin/edition.

  • Ebook (epub/pdf) discovery and matching — LibraForge now knows about
    ebooks at all, not just audiobooks. A standalone .epub/.pdf, or a
    matching pair sharing an epub/pdf sibling bucket-folder layout, is
    discovered independently of the audio walker and auto-matched against
    Open Library (primary) with Goodreads backfilling a blank cover/summary.
    Matched metadata writes to the same libraforge.json sidecar mechanism
    Fixer/Manual Review already use (media_type: "ebook"), never touching
    the epub/pdf file itself. Surfaced in Manual Review's search results with
    an EPUB/PDF/EPUB+PDF format badge and an Ebooks Match
    Report filter, with its own compare/apply flow.

  • Folder Forge organizes ebooks — every discovered .epub/.pdf now
    gets the same folder-per-book treatment as an audiobook:
    Author/Series/Book N - Title/, with " - EBOOK" appended to the series
    (or, for a series-less book, the title) so it never collides with the
    audiobook edition's own folder. This is unconditional — an ebook that
    shares an audiobook's exact filename is no longer bundled into the
    audiobook's folder as a companion; it always gets its own - EBOOK
    folder, whether or not a matching audiobook exists. Reuses the
    organizer's existing metadata cleanup, naming templates, structure cache
    and conflict detection rather than a separate mover. Scope: new/staged
    ebooks only — an existing EPUB//PDF/ bucket-folder library is left
    exactly as it is, not reorganized.

  • Enrichment Forge: Goodreads genres from reader shelves: each book's Goodreads
    shelves are read directly with their vote counts, so LitRPG, progression fantasy,
    cultivation, harem and young adult come through when enough readers agree (abs-tract's
    Goodreads output is limited to three generic genres). Paced like Metadata Forge's
    Goodreads calls, pauses automatically when rate-limited, and no longer needs abs-tract.
    Goodreads readers shelving a book as erotica/smut/nsfw is shown as explicit evidence.

  • Enrichment Forge: more sources, voted genres: books are also looked up in AudioSilo
    and Open Library, series in progressionfantasy.co.uk's catalogue and the HaremLit wiki,
    and summaries/descriptions are scanned for LitRPG, cultivation, harem and similar
    keywords. Every label is mapped onto one controlled genre list and voted on, so the
    result is a few main genres (the ones to build collections from) plus subgenres, not a
    pile of every label any source used. Audible's "Dragons & Mythical Creatures" shelf
    counts as Fantasy only; the Dragons subgenre needs dragon-specific evidence. Hover a
    genre to see which sources support it; other suggestions are one click away.

  • Enrichment Forge: standalone books: books with no series can be enriched too.

  • Enrichment Forge: your genres come first: genres you set by hand in Audiobookshelf
    are kept and marked "yours"; the genre tag inside the audio file counts as a source;
    store-style merged genres are split ("Action & Adventure" becomes Action and
    Adventure) and Science Fiction is written as Sci-Fi. LitRPG and Cultivation are main
    genres with Progression Fantasy kept as their subgenre.

  • Enrichment Forge: explicit, book by book: per-book explicit evidence and choice;
    the HaremLit wiki's rating pre-selects, everything else is shown as evidence.

  • Enrichment Forge: whole library: compile every series and standalone book in one
    resumable run, review the proposed genres in a table, and apply the rows you choose.

  • Enrichment Forge: collections from genres: create and refresh one Audiobookshelf
    collection per genre; your own collections are never touched.

  • Retire old metadata.json files: scripts/migrate-legacy-metadata-json.py compares
    every old metadata.json in your library with Audiobookshelf and removes it (the more
    recently edited side wins; a newer file's values are merged into Audiobookshelf first).
    Dry run by default, --apply to perform (it backs up every file first). These files made Audiobookshelf undo
    LibraForge's direct edits on its next scan. (#298)

  • Audiobookshelf: edits go straight through its API: when Audiobookshelf already
    knows a book, Metadata Forge, Manual Review, Fix Series, Enrichment Forge and ebook
    apply write it with Audiobookshelf's own API instead of a metadata.json it might
    re-read stale on a later scan and use to undo your edits. A book Audiobookshelf hasn't
    scanned yet still gets a metadata.json, which a background sweep removes once
    Audiobookshelf picks the book up. A book is only ever written by its exact folder, never
    matched by a shared placeholder ASIN. Two opt-in modes use Audiobookshelf's current data
    as input (Metadata Forge --weight-abs-metadata, Folder Forge --trust-abs-metadata),
    and each book shows which channel it was written through. (#288, #291)

  • Manual Review: Audiobookshelf only: apply, edit and Fix Series can skip the audio
    file's own tags and write only to Audiobookshelf, like Metadata Forge's
    --metadata-json-only. (#293)

  • Pocket FM as a manual source: Pocket FM audio series (Supreme Magus, Shadow Slave)
    aren't on Audible. Paste a show's link into the Manual Review card ("Fill from Pocket
    FM") for a normal result card, or in Fix Series ("Fill from Pocket FM") to fill
    series, author, genre and language for every book at once; each book keeps its own
    number, since Pocket FM has no volumes. Read from the public show page; Pocket FM has no
    public search, so it can't match automatically. (#324, #331)

  • Metadata Forge writes Audible's genres: an Audible match now carries its store
    categories, named the way Enrichment Forge names genres (Fantasy, Epic Fantasy, Urban
    Fantasy), so a file's "Audiobook" or foreign-store genre tag gets replaced. In
    Audiobookshelf a genre is only set when the book has no real genre yet (about 2,700
    books held nothing but "Audiobook"), so genres you or Enrichment Forge chose are never
    replaced, and Enrichment Forge no longer counts a genre tag LibraForge wrote as a second
    Audible vote. (#318)

Fixes

  • Folder Forge: a real series name ending in a genre word was wiped even
    from trusted metadata
    — Street Cultivation (a real, Audible-confirmed
    series) was being planned as a standalone book, identically to genuine
    junk, because the marketing-cleanup filter had no trust exemption for
    series the way it already did for titles. Fixed with a narrower rule than
    a blind copy of the title exemption: a series kept when trusted must still
    be dropped if it is nothing but a genre/marketing word or a coupling of
    them (LitRPG, Fantasy Cultivation), so Arthur Stone's genuinely-generic
    LitRPG pseudo-series still correctly disappears. (Issue #276)

  • Folder Forge: series/title marketing cleanup is now visible — a book
    whose series or title was dropped or trimmed as generic marketing/genre
    text gets a review reason and a run-summary count instead of silently
    losing the field. A censored in-word asterisk no longer splits a title, a
    series that is really the author credit list is dropped, a series can be
    read from a "Series, Book N" subtitle when Audible has none, and a fixer
    digit-in-hyphenated-compound (The 3-Day Effect) is no longer read as a
    book number. (Issues #263-#267)

  • Folder Forge: a merged single file next to its own chapter files is now
    called out as a possible duplicate
    , with the folder and file paths shown
    on the card, instead of a generic "skipped conflict". The "folder name
    matches the template" skip now only applies to a library-wide scan, not to
    the normal _unorganized import folder, so books actually waiting to be
    organized are no longer mistaken for already-organized ones. (Issues
    #268-#269)

  • Folder Forge: the review-reason filter no longer explodes into one
    option per book
    — a review reason that varies per book (a path, a
    chapter count, the source series text) stayed out of the grouped/filtered
    reason text; the per-book detail now shows on the book's own card
    instead. (Issues #273-#274)

  • Folder Forge: the Suspicion Report was full of false positives — every
    ordinary standalone book (no series, no problem) was flagged as
    "missing_series", and a purely informational review reason such as "title
    matches series name" was promoted to a suspect on its own. Verified on a
    real 391-book run: 57 suspects dropped to 22, with every real issue still
    caught. (Issue #275)

  • Settings saved in the UI were lost when the published image was
    updated
    : the published image and docker-compose.dist.yml only kept
    /auth and /app/reports across updates, so saved patterns, the
    Audiobookshelf / abs-agg / abs-tract connections and report retention were
    discarded whenever the container was recreated. They now live in a
    libraforge-settings volume (LIBRAFORGE_SETTINGS_DIR); see the upgrade
    notes above for the one-time step. (Issue #261)

  • Folder Forge: Chapter Forge's companion files could get stranded on
    move/rename
    — its loose artifact files (.libraforge-chapters.srt,
    .libraforge-chapters.cue, .libraforge-ai-review.md) are now tracked as
    companions alongside the sidecar, so they move with the book instead of
    being left behind at the old path.

  • Organizer: custom naming template ignored for already-indexed series —
    a book routed into a series folder the organizer had already placed a book
    into used to silently fall back to the default "Book N - Title" leaf and
    original filename, regardless of an active custom naming template — so a
    custom template only ever applied to the first book found in a given
    series. Every subsequent book in that series now renders through the same
    template as a brand-new series would. (Issue #251)

  • Match Report: series-drift badge over-flagged clean series tagging —
    series groups are now split into "clean tagging" (informational "series"
    badge, no longer shown in the Suspicion Report widget) and "actual drift"
    (raw tag variance or an author mismatch — keeps the "series drift" flag).

  • Author names: a broadcaster/publisher acronym in a credit was misread as
    initials
    — BBC - Andrew Marshall & John Lloyd was being rewritten to
    B.B.C. - ..., since the unspaced-capitals rule (the one that correctly
    turns JD into J.D.) couldn't tell "BBC" apart from real initials. It
    now checks the token against the existing publisher catalog first. (Issue
    #278)

  • Audible series text with a redundant book number baked in — Audible
    sometimes returns a series title with the sequence written into the text
    itself instead of (or alongside) the separate sequence field, e.g.
    "Blight, Book 1" rather than series "Blight" + sequence "1". Left
    alone, that text was written verbatim to every sidecar/tag a book ever
    got, and the redundant wording defeated series grouping (siblings never
    shared the same clean base text). Every future Audible match is now
    cleaned automatically; a standalone backfill tool
    (scripts/fix-series-book-number-suffix.py) handles data written before
    this existed, dry run by default, with --apply/--log/--revert and
    --tags for embedded MP4/ID3 tags too. It also keeps metadata.json's
    own series field in sync with the fix — real books stayed on stale (some
    of it doubled, "Name, Book #N #N") text otherwise, since nothing
    regenerated it once the underlying data was corrected, including a book
    with no marker/tag data of its own to derive the correction from (a
    whole-library sweep catches those too). A bare "Book #" with no number
    ("Dean Koontz: From the Vault, Book #") is cleaned too. (Issue #281, PR #297)

  • Folder Forge: a loose_file move left libraforge.json/metadata.json/cover
    behind
    — moving a single audio file only ever brought the file itself
    (and any sidecar sharing its exact filename) to the new location; a
    generic same-folder sidecar or cover with no relationship to the audio's
    filename was silently left at the old path. Verified against a real
    391-book run: 366 books had lost their sidecar at the destination.
    Recovered the orphaned files and fixed the mover to bring them along
    going forward. (Issue #282)

  • abs-agg: fabricated publisher, wrong provider badges, and unit bugs
    across several providers
    — a batch verification pass against real
    abs-agg results found the normalizer was fabricating a publisher value
    instead of leaving it blank when the provider didn't supply one, and
    mislabeling which provider a match actually came from (Issue #254); the
    batch normalizer wasn't threading a provider's real language field
    through for LibriVox specifically, forcing every match to the same
    default (Issue #255); and Die drei ??? durations were being treated as
    seconds when abs-agg already returns them in minutes, inflating runtime
    by 60x (Issue #256). ARD Audiothek, LibriVox, and Die drei ??? are now
    live-verified working correctly. Hardcover is reachable now that its
    token is configured, but abs-agg's own backend search for it returns
    empty results for known titles — an upstream abs-agg bug, not something
    fixable on our side. Storytel, Audioteka, and BookBeat remain blocked on
    a separate upstream abs-agg bug rejecting the market parameter those
    providers require.

  • Changes no longer revert on the next Audiobookshelf scan: an old metadata.json
    left in a book's folder is compared with Audiobookshelf and removed whenever Metadata
    Forge, Manual Review or Enrichment Forge writes the book through the API. (#298)

  • Enrichment Forge no longer overwrites every book's narrator: the narrator is empty
    by default and only applied when you opt in. (#299)

  • Enrichment Forge explicit can now be cleared: Don't change / Explicit / Not explicit. (#303)

  • Enrichment Forge shows the books' real current genres (tags are listed separately). (#300)

  • Enrichment Forge skips items with no audio (e.g. an ebook checklist in a series
    folder) instead of matching them to unrelated books. (#301)

  • Enrichment Forge keeps parent genres from Audible's categories (a Space Opera book
    also gets Science Fiction; children's/teen audiences are kept). (#302)

  • Every series Enrichment Forge lists can be compiled (odd series names gave
    "Series not found"). (#304)

  • Same-named series by different authors are no longer merged in Enrichment Forge. (#305)

  • "Audiobook" is never written as a genre: M4B Tool no longer stamps it on the files it
    builds, and the "Audio Book" spelling is filtered too. (#306)

  • Security: the Audiobookshelf API key showed in plain text in run details: a run's
    command, as shown in the UI and stored in its log and report, included the
    --abs-api-key value. It is now redacted everywhere. (#292)

  • Stale series tags stayed in MP4/ID3 files: rewriting a book's series, subtitle,
    ASIN, ISBN or publisher left the old value behind as a duplicate tag that
    Audiobookshelf could read instead of the new one; MP4 files also now get the
    ©mvn/©mvi series atoms. (#289)

  • Metadata Forge: many more books match (a forced run over an _unorganized folder
    went from 137 to 145 of 151):

    • A book split into numbered part files no longer takes a part number as its book
      number, its title as its series, or one chapter's name as its title (48 Laws of
      Power, 1177 B.C., Supreme Magus).
    • A folder holding separate full-length books (every file book-sized, each album naming
      a different book) is no longer grouped as one book, and a folder with two numbered
      copies of the same book is reported instead of turning every file of one copy into a
      fake book.
    • Titles keep a leading number that belongs to them ("1177 B.C."); "By Eric H. Cline" in
      a title is read as the author; "(Read by X)" in a folder names the narrator, not the
      author; "Title (Series, Book N)" in a title tag gives the series and number; any
      credit tag (album artist, artist, composer) can confirm the author when a rip puts an
      uploader or narrator in the author tag.
    • An album that just repeats the title is no longer taken as the series (this made
      Sapiens a series-only match and let a different Harari book score high), and a book
      titled like its own unnumbered series matches in full.
    • Release-group tags ("[PZG]"), a bracketed publisher ("[Yen Audio]"), bare
      "Audiobook"/"Unabridged" words and an "Audiobook:" prefix are stripped from titles
      and authors, and the AUDIBLE_ASIN tag is read.
    • A folder named "Series, Book 05 - Title" no longer has its title read as the author.
    • When two results tie, the one read by your narrator wins, then the one whose SKU or
      publisher matches the file; the same recording listed under two ASINs (a US and a UK
      edition) takes Audible's first listing instead of going to manual review.
    • A title with a different edition's subtitle ("Troy: The Siege of Troy Retold" vs
      Audible's "Troy") matches.
      (#311, #312, #313, #315, #317, #323)
  • Metadata Forge: Opus/OGG tags were never read, and a slow file read over the network
    could be saved as the file's original-tags backup with no tags in it, after which every
    run matched that book blind (and restoring the backup would have erased its tags). A
    timed-out read is retried with more time and a failed one is never saved. (#323)

  • Metadata Forge: phantom series number 1: a rip's track number ("1/1") was reported
    as the sequence of books that have no series; a track number now only counts next to a
    real series tag. (#314)

  • Never a series: a genre ("LitRPG"), a marketing phrase ("A LitRPG Series"), a format
    word, a publisher or a bare book number is no longer read or written as a series, even
    when Audible supplies it; a descriptor such as "(light novel)", "(publication order)" or
    "(Full-Cast Editions)" is dropped from a series name, and other parentheses stay part of
    it ("Detroit Free Zone (DFZ)" used to become "DFZ"). (#315)

  • Goodreads fallback: a longer Goodreads title now counts only when the extra words
    are a separated subtitle (a "48 Laws of Power ..." notebook listing was accepted as the
    book); a Goodreads match without a series keeps the file's series instead of erasing it;
    a "series" that is just the book's own title and subtitle ("Postwar ... Since 1945"
    #1945) is dropped. (#316, #323)

  • Smart mode fixes a dirty series: a series such as "Dean Koontz: From the Vault,
    Book #" from a rip's tags is now replaced in smart mode too, and replaced rather than
    kept beside the clean one in Audiobookshelf. (#320)

  • Overwrite mode could clear a book's Audiobookshelf genres when the match had no
    genre. (#318)

  • Stopping a run now stops it: a cancelled run went on to scan the whole library for
    duplicate ASINs before exiting (a dry run now exits at once), and two starts in the same
    moment (a double click, a second tab) launched two identical runs; a second run over the
    same or an overlapping folder is now refused while one is active. (#321, #322)

  • Metadata Forge: a match to another recording no longer replaces your narrator: when
    Audible only sells a different edition of a book (Tunnels: your Recorded Books copy read
    by Steven Crossley, Audible's Audible Studios edition read by Paul Chequer), the match's
    details are still used but the narrator from your files is kept. The match report marks
    these books with a Different edition badge, a note in the card naming both readers
    and lengths, a summary tile and a filter. (#332)

  • Match report: series groups: the search box now filters series groups too (they
    stayed on screen whatever you typed). Group names no longer keep a book label ("...,
    Vol.", a bare "Volume"); ebooks are no longer mixed into audiobook groups (an audiobook
    and its own epub showed up as a two-book "series"); a copy of the same unnumbered title
    is not a series; and titles like "Series, Vol. 6: Subtitle" or "Title (Series, Book 2)"
    now group, so series such as The Saga of Tanya the Evil and The Furyck Saga are found. A
    dry-run report explains that its "no series" groups are expected until you apply. (#329)

Docs

  • README and docs/features.md describe the author names scheme, the shared
    pattern settings and the new settings volume (including the one-time step
    when updating).
  • README: fixed the Uvicorn credit link (uvicorn.org → uvicorn.dev).
    (Issue #260)

Don't miss a new LibraForge release

NewReleases is sending notifications on new releases.