Changed
- all normalized string_metrics can now be used as scorer for process.extract/extractOne
- Implementation of the C++ Wrapper completely refactored to make it easier to add more scorers, processors and string matching algorithms in the future.
- increased test coverage, that already helped to fix some bugs and help to prevent regressions in the future
- improved docstrings of functions
Performance
- Added bit-parallel implementation of the Levenshtein distance for the weights (1,1,1) and (1,1,2).
- Added specialized implementation of the Levenshtein distance for cases with a small maximum edit distance, that is even faster, than the bit-parallel implementation.
- Improved performance of
fuzz.partial_ratio
-> Sincefuzz.ratioandfuzz.partial_ratioare used in most scorers, this improves the overall performance. - Improved performance of
process.extractandprocess.extractOne
Deprecated
- the
rapidfuzz.levenshteinmodule is now deprecated and will be removed in v2.0.0
These functions are now placed inrapidfuzz.string_metric.distance,normalized_distance,weighted_distanceandweighted_normalized_distanceare combined intolevenshteinandnormalized_levenshtein.
Added
- added normalized version of the hamming distance in
string_metric.normalized_hamming - process.extract_iter as a generator, that yields the similarity of all elements, that have a similarity >= score_cutoff
Fixed
- multiple bugs in extractOne when used with a scorer, that's not from RapidFuzz
- fixed bug in
token_ratio - fixed bug in result normalization causing zero division