github 54yyyu/zotero-mcp v0.14.0
v0.14.0: Zotero Agent

4 hours ago

Zotero Agent, a new Zotero plugin: an AI chat panel inside Zotero for your own Claude Code, Codex or pi. Install: zotero-cli plugin, or download zotero-agent.xpi below and use Tools > Plugins > Install Plugin From File. It updates itself. Guide: docs/chat-plugin.md · stevenyuyy.com/zotero-mcp/agent

Added

  • Zotero Agent, a Zotero plugin. An AI chat panel inside Zotero for your own Claude Code, Codex or pi, with the open item, page and selection as context and zotero-cli for the library. The built zotero-agent.xpi ships in the wheel and is attached to each GitHub release; zotero-cli plugin prints where it is and how to install it (--path prints only the path). See docs/chat-plugin.md.
  • zotero-cli read --find "phrase" locates a passage and returns the matching pages with short snippets, without reading whole pages. --start-page is optional with --find.
  • zotero-cli open shows a PDF page or an annotation in the Zotero reader (open KEY --page N, open --annotation KEY, annotations create --open).
  • zotero-cli plugin --reveal shows the plugin file in the file manager; the plugin docs have a prompt that lets any agent do the install.
  • Notes take the note editor's full formatting. notes create and notes update accept Markdown with $math$ and a small safe tag subset (<u> <s> <sub> <sup> <mark>, colour and highlight colour spans) and convert it to Zotero's note HTML; notes update --append adds to an existing note. Existing citation, highlight and image markup survives an edit.

Fixed

  • Note search no longer misses real matches behind markup-only hits (#635). search_notes_local applied its SQL LIMIT before discarding notes whose match was only in HTML markup, so a query like zotero or version (present in Zotero 7 citation attributes) could fill the limit with discarded rows and return few or no notes. The limit now applies after the filter.
  • Web page snapshots no longer carry their images as base64 (#670). The Zotero Connector saves a page with every image inlined as a data: URI, and HTML extraction copied each one into the Markdown, so zotero_get_item_fulltext on an 8 MB snapshot returned 6.1 million characters (about 1.5M tokens) around 33K characters of text and failed over HTTP with "Server-sent event exceeded the 1048576 byte limit"; semantic search embedded the same base64. Images, video posters and sources, and links whose URL is a data: URI are now dropped: an image leaves [image: <alt text>] (or [image]), a link leaves its text. Remote images and links are unchanged.
  • SQLite free-text search ignores case and accents, like Zotero's own search (#663). With ZOTERO_BACKEND=sqlite, zotero_search_items compared raw values with SQLite's LIKE, which folds ASCII case only, so índice and indice missed a title with Índice, and SUCESIÓN missed sucesión, while the API path found them. Title, creator, abstract, tag and note matches now go through zsearch_norm on both sides, as zotero_advanced_search already does since #417. The folding function itself is about seven times faster, because a dash regex that could never match after unidecode is gone.
  • A failed note update is a tool error, not a success message (#669). zotero_manage_note(action="update") returned its failure as text, so MCP clients saw isError=false even when Zotero answered HTTP 500, and any read failure (a refused connection, a timeout) was reported as "No item found". Failures now raise a tool error that keeps the underlying cause; only a real 404 says the note is missing. Thanks @lipaopao000.
  • zotero-cli reports a failed fulltext read and a rejected library switch as failures (#599, #595). A fulltext attachment error came after the metadata, and --json wrapped it as data, so it arrived as ok: true with exit status 0; the error now leads the text and get fulltext no longer treats it as data. switch-library with an unsupported library type or an unknown group, feed or personal id now starts with Error: like other refusals. Thanks @feiiiiii5.
  • zotero_export_bibliography(item_keys=...) renders every requested item (#662). Keys were fetched from /items, which also returns each item's notes and attachments, so those filled the 100-row page: 60 keys gave 52 entries and 80 gave 53, with no warning, in every format. Keys now go to /items/top in batches of 50, Zotero's limit for one itemKey filter.
  • zotero_export_bibliography exports a whole collection or library, not the first 100 rows. A 191-reference collection came back with 51 entries and no notice, because the fetch included each item's child attachments, which render as empty entries, used up the 100-row cap and were then dropped; the library-wide export stopped at 68 of 470 the same way, and BibTeX took one 100-row page. It now reads the top-level item endpoints and pages until exhausted for bib, citation and bibtex (191 of 191 and 470 of 470 live). A large export starts with the usual response-size warning.
  • Plain-text notes keep <, > and & (#664). zotero_create_note wrapped plain text in <p> without escaping it, so x<y and y>z was stored as markup, Zotero read <y and y> as a tag, and the text inside it was lost (reading the note back gave xz). Plain text is now HTML-escaped. Text is treated as HTML when it contains a tag such as <h2>, <ul> or <p class=...>, not only the exact <p> or <div>.
  • A parent collection named ambiguously is an error, not a guess (#665). zotero_create_collection and zotero_update_collection resolved a parent_collection name to the first collection with that name, so with a Readings subcollection under two courses the new or moved collection could land under the wrong one, and the tool reported success. Parent names now go through the same resolver as zotero_set_item_collections (#233): an ambiguous name returns an error listing the candidates, a Parent/Child path picks one, and an eight-capital name such as PROJECTS is read as a name unless a collection has that key.
  • zotero_advanced_search answers any Zotero field from SQLite. Only seven fields (title, abstract, DOI, publication, dates, item type) were translated to SQL, so a condition on extra, publisher, volume, ISBN, url, language or any other field sent the whole query on a walk of the library over the API: about 2 s, 11 requests and 2.5 MB on a 2,500-item library, against 2 ms in SQL. Any field in the database's field table is now queried directly, resolved per item type like title, with the name bound as a parameter. Names the database does not know behave as before, and so do ordering operators (isGreaterThan, isBefore, ...) on these fields, which stay on the API path because it orders numeric values by magnitude where SQL would order the stored text (volume > 9 must hold for 12).
  • zotero_search_items applies tag and item_type to its semantic fallback. When every text search came up empty, the last-resort semantic search ignored both filters: tag="a" returned items that do not carry the tag and item_type="book" returned articles, each presented as a match for the filtered query. The hits are now filtered with the same syntax as the text searches (OR, -tag, a || b types, case and accents folded for tags), and the index is asked for more candidates so the limit survives the filter. A tag containing a wildcard, which cannot be checked here, returns no fallback hits. The fallback note also no longer sits between the title and the with tags: line.
  • zotero_find_related_papers and the Scite tools no longer block every other tool while they wait on the network (#431). They held the shared Zotero API lock across each OpenAlex or Scite request, so with 1.5 s of OpenAlex latency a concurrent zotero_get_recent waited 6.3 s (2.7 s behind scite_enrich_item). The lock is now taken only around the reads from Zotero itself; the same test waits 0.01 s.

Changed

  • License: the Zotero Agent plugin (plugin/) is AGPL-3.0-or-later, the license Zotero itself uses. The MCP server and zotero-cli stay MIT, so the package is now MIT AND AGPL-3.0-or-later and carries both license files.
  • A new README and logo. The red Z in a chat bubble is now the project's logo; the README shows the three ways in (MCP server, zotero-cli + skill, Zotero Agent) with screenshots, and the PyPI summary says what the package is for.
  • Reading a PDF in page ranges no longer re-parses it every time. zotero_read_pdf_pages and zotero_get_item_fulltext ran pdf-inspector's whole-document Markdown pass on every call, and it costs about the same for one page as for the whole file (3.7-6 s on some 28-page papers, with the GIL held so other tools stalled), so reading a paper in chunks paid it again for each chunk. The parse of an unchanged file is now kept in a small in-process memo (3 files, 8M characters in all), so repeat and different-range reads of the same PDF return in milliseconds. Indexing is unchanged.
  • update-db no longer probes every item for fulltext. For each item the indexer asked for the parent's /fulltext (always empty for a regular item) and then its /children, so a 552-item library cost 1041 requests to find text for 99 items. It now asks only for each item's PDF attachments' text, and a whole-library run lists those attachments in one paged pass first: 106 requests, same text. The docs now say plain update-db also indexes the text Zotero already extracted, in local mode too.

Don't miss a new zotero-mcp release

NewReleases is sending notifications on new releases.