github 54yyyu/zotero-mcp v0.14.1

5 hours ago

Added

  • semantic_search.persist_directory moves the ChromaDB index (#617). The index was pinned to ~/.config/zotero-mcp/chroma_db, and on Windows a user folder with non-ASCII characters (for example C:\Users\王林澜) makes ChromaDB write part of the index to the wrong place once it holds about 1,000 items, after which every search fails with Error loading hnsw index. Set the key in config.json to an ASCII path; the server, db-status and zotero-mcp setup all follow it, and the default location is unchanged. A non-ASCII path on Windows now logs a warning that names the setting. Thanks @lots-o. The pre-update backup follows the configured directory.

Fixed

  • On Windows, two servers no longer run the same semantic-index update at once (#267). The update lock used fcntl, which Windows lacks, so there it quietly did nothing: a client that starts the server twice (Codex does) ran two full-text indexes of the same library side by side, each holding gigabytes. The lock now uses msvcrt.locking on Windows, on a byte past the stored pid so the process that loses can still read who holds it. Reported by @lwz20210407.
  • zotero-mcp setup and zotero-cli no longer crash on a Windows console (#26). Under the default cp1252 or cp936 console encoding, printing an emoji (setup-info, setup) or an item title in another script raised UnicodeEncodeError and aborted the command, so setup died before it wrote its config. Both entry points now keep the console's encoding and replace what it cannot hold.
  • zotero-mcp setup finds the Claude Desktop config of the Microsoft Store build on Windows (#26). That build reads %APPDATA%\Claude through a package redirect, so the real file is under %LOCALAPPDATA%\Packages\Claude_<id>\LocalCache\Roaming\Claude\; setup wrote to %APPDATA%\Claude Desktop and reported success while Claude never saw the server. The packaged path is now probed, and written to when it exists, like the other known locations.
  • Tables without closing tags no longer flood HTML extraction with empty cells (#687). html.parser nests an unterminated <tr>/<td>, and markdownify then printed a blank header row and a | --- | row as wide as the whole table for every nested row: zotero_get_item_fulltext on a 74-row datosmacro statistics snapshot returned 218,832 characters, 92% | | | and | --- | --- | runs (one line of 125,237 characters), around 17K of real data. The scaffolding (a header row of nothing but empty cells, the separator under it, and the runs of both inside the nested rows) now collapses to one short row, so that page extracts to 17,931 characters with every word kept; 70 of 355 snapshots in one library shrink, 1.3M characters in all. A table with a real header keeps its separator and its empty cells.
  • Deeply nested web page snapshots are extracted instead of silently replaced by Zotero's flat index text (#688). html.parser leaves unterminated tags nested, so pages with many unclosed <li>/<td> tags exceeded Python's recursion limit inside markdownify ("Extraction failed ... maximum recursion depth exceeded"), and zotero_get_item_fulltext fell back to the index text with no structure: 7,061 characters for a datosmacro page whose Markdown is 129,328, and the same failure kept a 1.1 MB article out of the semantic index. Extraction now retries once on a worker thread with a raised recursion limit (5000, restored afterwards) and a 64 MB stack; pages nested more than about 2,000 levels deep (one converts to 29M characters) still use the fallback.
  • Voyage embeddings say whether they embed a document or a query (#684). Through the OpenAI provider with a Voyage base_url, indexing now sends input_type: "document" and search sends input_type: "query", so Voyage can tune the pair for retrieval. Other base URLs are unchanged. An existing index is not rebuilt. Thanks @sghng.
  • Annotations written by zotero-cli and the MCP tools sort in reading order (#608). The sort index was page|y|x in PDF coordinates, where y grows upward, so a page's annotations came out bottom to top in Zotero's list. It is now Zotero's own page|character offset|top: the offset is the nearest page character (spaces and control characters not counted) from PyMuPDF's text, so it can differ from Zotero's by a few characters, and the top is the page height minus the rect's top edge. A two-column page now orders left column first. Thanks @dnlfis42.
  • zotero-cli get collection-items fails for a collection that is not in the library (#606). A key from another library came back as ok: true with zero items, which reads as an empty collection. It now exits 1 with a not_found (or tool_error) envelope, or the message on stderr. A genuinely empty collection is still a success. Thanks @MohammadHijjawi97.
  • update-db without --fulltext no longer indexes PDF annotations (#604). The API scan skipped attachments and notes but not annotations, so each highlight became its own document, while the incremental fetch and the SQLite scan already skipped them. All three now share one list of item types. Annotations already in an index stay until update-db --force-rebuild. Thanks @RizgarOzan.

Don't miss a new zotero-mcp release

NewReleases is sending notifications on new releases.