Added
- Added a new AI chunking mode to semchunk that leverages Isaacus enrichment models to hierarchically segment texts.
- Made it possible to chunk Isaacus Legal Graph Schema (ILGS) Documents instead of just strings.
- Added a new
tokenizer_kwargsargument tochunkerify()allowing users to specify custom keyword arguments to their tokenizers and token counters.tokenizer_kwargscan be used to override the default behavior of treating any encountered special tokens as if they are normal text when using atiktokenortransformerstokenzier. - Where a
tiktokenortransformerstokenizer is used, started treating special tokens as normal text instead of, in the case oftiktoken, raising an error and, in the case oftransformers, treating them as special tokens. - Added support for Python 3.14.
Changed
- Demoted asterisks in the hierarchy of splitters from sentence terminators to clause separators to better reflect their typical syntactic function.
- Dramatically improved performance when handling extremely long sequences of punctuation characters.
- All arguments to
chunkerify()except for the first two arguments,tokenizer_or_token_counterandchunk_size, are now keyword-only arguments. - All arguments to
chunk()except for the first three,text,chunk_size, andtoken_counter, are now keyword-only arguments. - Significantly improved performance in cases where
merge_splits()was the biggest bottleneck by switching from joining splits with splitters to indexing into the original text. - Slightly sped up
merge_splits()by switching to the standard library'sbisect_left()function which is now faster than the previous implementation.
Removed
- Dropped support for Python 3.9.