The headline changes for v0.3.0 are a major upgrade of the native LZ engine, and the Compression Transformer, a neural selector that builds numeric compression graphs on the fly.
LZ Improvements
OpenZL shipped a new LZ engine in v0.2.0, and since then we've been busy expanding the feature set and improving decompression speed.
First and foremost, OpenZL shipped an implementation of PivCo Huffman, a new Huffman layout by Marcin Żukowski that allows for significantly faster decoding speeds by leveraging SIMD code. The LZ graph uses PivCo Huffman as the entropy backend to accelerate decompression. Compared to v0.2.0 at compression level 1 with a 64KB window size decompression speed has improved by another 33%, which leaves decoding speed 144% faster than Zstandard at equivalent settings.
| Compressor | Compression Ratio | Compression Speed | Decompression Speed |
|---|---|---|---|
| OpenZL LZ v0.3.0 | 2.73 | 467 MB/s | 3062 MB/s |
| OpenZL LZ v0.2.0 | 2.74 | 466 MB/s | 2288 MB/s |
| Zstd | 2.74 | 419 MB/s | 1254 MB/s |
Additionally, we've expanded the supported compression levels to the equivalent of Zstandard levels 1 through 7, and support window sizes up to 256 MiB. We've also added support for configurations that offer faster decompression speed than LZ4 at similar compression ratios.
Finally, we've added an LZ trainer in the zli that can be invoked with zli train --profile lz $SAMPLE_DIR --pareto-frontier --output $OUTPUT_DIR. It leverages OpenZL's graph-based model to tune the LZ parameters and backend compression graphs to your data and outputs the Pareto Frontier of configurations that offer varying speed vs. compression ratio tradeoffs.
See the blog post for more details!
The Compression Transformer
Choosing the right combination of codecs for a given input has always been OpenZL's greatest challenge. The Compression Transformer now makes that choice automatically: it builds the compression graph on the fly, for each input, with no per-source training and no change on the decompression side.
Evaluated on 868 families of numeric streams (34,737 files, 17.9 GB), it compresses on average 35% better than zstd -19, and lands within 1.1% of OpenZL graphs trained.
| numeric width | zstd -19 | xz -9 | OpenZL untrained | OpenZL trained | Transformer | vs zstd -19 | vs xz -9 |
|---|---|---|---|---|---|---|---|
num8
| 12.26x | 10.19x | 9.867x | 12.23x | 12.42x | +1.3% | +21.9% |
num16
| 6.772x | 6.467x | 6.515x | 8.415x | 8.429x | +24.5% | +30.3% |
num32
| 3.056x | 3.483x | 3.516x | 4.390x | 4.255x | +39.2% | +22.2% |
num64
| 4.578x | 5.586x | 6.114x | 8.723x | 8.477x | +85.2% | +51.8% |
| all | 5.752x | 5.920x | 6.036x | 7.846x | 7.760x | +34.9% | +31.1% |
Under the hood, a small neural network scores the candidate codecs for each stream, one decision at a time, and the streams they produce go back through the same loop until the graph is complete. The details are in the blog post.
The Transformer is enabled at compression level 7 and above, through ZL_GRAPH_NUMERIC and the numeric profiles of zli (e.g. zli compress -l 7); the default level (6) is unchanged. It is also exposed as a standalone graph, ZL_GRAPH_TRANSFORMER_NUMERIC, which ACE and the LZ trainer can now select as a backend.
Miscellaneous
Several codecs and graphs join the catalog. sparse_num compresses numeric streams dominated by zeros. The zstd codec can now use trained dictionaries: zli train produces them alongside the graph as a dictionary bundle, which compress, decompress and benchmark load with --dict-bundle, and which library users load through the new ZL_DictLoader API. The brute-force selector is now a standard, serializable graph (ZL_GRAPH_BRUTE_FORCE), and an opt-in codec output cache avoids repeating identical work across its trials.
Training is more predictable. --max-time-secs now bounds the whole ACE training instead of each graph separately (#930), --max-num-candidates limits the number of trained candidates, and training respects the requested format version, so a trained compressor only uses codecs its target decoders support.
On the platform side, ARM gains NEON and SVE2 kernels and is now tested in CI, the vendored xxHash moves to v0.8.4 (#964), and binaries shrink by about 447 KiB with the removal of the legacy built-in GBT numeric model. zli compress accepts -l/--level, big-endian integer profiles (be-*) join the numeric ones, and --version now reports the correct version. SDDL gains a between() built-in, and an @instant_parse record annotation, which makes the compiler reject a record whose layout would require scanning.
A handful of changes are worth checking before upgrading:
the frame format max version is now 27 (was 24), and new frames use it by default: set ZL_CParam_formatVersion to 24 to produce frames that v0.2.0 can read. Serialized compressors using ZL_GRAPH_NUMERIC require v0.3.0. ZL_MaterializerDesc2 is renamed ZL_MaterializerDesc, replacing the previous descriptor of that name; ZL_Compressor_registerBruteForceSelectorGraph() is replaced by ZL_Compressor_buildBruteForceSelectorGraph(); and signed integer profiles no longer force a ZigZag transform.
Full Changelog: v0.2.0...v0.3.0