github fcorbelli/zpaqfranz 65.3
Just about everything

4 hours ago

zpaqfranz 65.3 BEWARE: this is a "big bang" build

This release introduces a substantial number of new features and changes, most of which have undergone only limited testing.
Please be careful when upgrading: try it on a copy first, and keep using -m1 (or whatever you use now) for the backups that matter until you trust the new features.

On the positive side, processing performance for large files has been improved dramatically, resulting in significantly faster processing times.

Prebuilt binaries are also provided for a range of “universal” environments, including both Intel and non-Intel NAS platforms, as well as ESXi.

As always, the more feedback and real-world testing we receive, the more we can improve and refine the software.


In a nutshell

  • on big files, up to about three times faster than 7-Zip (see the short answer);
  • a faster compression overall;
  • a MUCH faster new experimental -turbo mode for parallel processing;
  • two back-compatible brand new compression methods, -m6 (LZ4) and -m7 (LZAV);
  • mount support on FreeBSD, OpenBSD and macOS;
  • a completely redesigned mount cache, making large-file reads much faster;
  • improved hash healing and -touch;
  • new NAS, ESXi and macOS builds;
  • numerous portability fixes and additional autotests.

The headline: a is faster, and you do not have to do anything for it.
With the default method (-m1), 65.3 writes the very same archive as
65.2, to the byte, in three quarters of the time. Add -turbo and it takes
half the time, and still writes the very same archive.
(Byte for byte means checked with cmp, using -timestamp so that the date of the version is the same too.)

This document has four parts:

  • Part one: what is new, in short, for everybody
  • Part two: the same things in more detail, for power users
  • Part three: "lo spiegone", how it works inside, for developers
  • Appendix: a test on real data, zpaqfranz against 7-Zip, WinRAR and zstd

How fast is it? The short answer

For compressing big files (virtual machines, disk images, databases), depending on the hardware, zpaqfranz 65.3 is up to about three times faster than 7-Zip in its fastest mode, and it uses a fraction of its CPU.

In the test at the end, on an 11.8 GB virtual machine:

  • -m7 -turbo: 5.7 s, against 21.9 s of 7-Zip (-mx1) and 40.4 s of WinRAR (-m1): 3.8 and 7 times faster (the archive is 4% bigger than 7-Zip's);
  • -m1 -turbo, which writes the smallest archive of the test: 13.7 s, 1.6 times faster than 7-Zip.

Please read those numbers with the right eyes:

  • The environment matters. They depend on the hardware and on the operating system (here a Ryzen 9 with 16 cores, Windows 11, a RAM disk). On your machine they will be different. The purpose is not to verify a difference of a few percent, which would need a specific laboratory, but to show how wide the difference is. When a program goes from 500 MB/s to 1.5 GB/s, a wall clock is enough to notice.
  • It is an unequal comparison. zpaqfranz cuts every file into fragments, computes their SHA-1, deduplicates, hashes every file and keeps an index of all the versions. tar | zstd, 7-Zip and WinRAR do much less. zstd practically does nothing but compress, and it is fast also for this reason: on this data it is faster than zpaqfranz, and that is fine.
  • With many small files the gap is small: 1.25 times faster than 7-Zip, and tar | zstd is the fastest.

What runs where

system -turbo -m6 / -m7 mount
Windows yes yes yes (WinFsp)
Linux yes yes yes (libfuse3)
FreeBSD yes yes yes (fusefs-libs3)
OpenBSD yes yes yes (the FUSE of the base system)
macOS yes yes yes (FUSE-T or macFUSE), not in the universal binary
Solaris yes yes no
NAS (static, musl) yes yes no
ESXi no no no
  • -turbo needs more than one thread: with -t1 a plain a runs.
  • The JIT exists only on Intel/AMD processors. The ARM builds (ARM NAS, Apple Silicon) run the interpreter.
  • The Apple Silicon half of the macOS universal binary is built, but it was never run: there was no ARM Mac to try it on. If somebody wants to give me an Apple M-series machine for testing, I will be happy to try it there too.

Part one, what is new, for the user

  1. a is faster, same archive
  2. -turbo: a even faster
  3. -m6 and -m7: fast compression that any zpaq can extract
  4. mount on FreeBSD, OpenBSD and macOS
  5. mount reads big files much faster
  6. mount -test: does the mount work here?
  7. Hashes that heal themselves, a better -touch
  8. Smaller things, and how to read the banner

1.1 a is faster, same archive

Nothing to type. The same a, with the same default method, is faster:

zpaqfranz a z:\backup.zpaq c:\vm

Measured on Windows (Ryzen 9 7950X3D, 32 threads), with a Debian virtual
machine (11.8 GB) plus 1.1 GB of mixed files (1,238 of them). Sources and archives are on a RAM disk, so
the disk is never the limit: this measures the programs, not the storage. As said in
the short answer, your numbers will depend on your hardware and
operating system: the point is the size of the difference, not a few percent.

method 65.2k 65.3 65.3 -turbo
-m1 34.4 s 25.3 s 17.1 s
-m0 31.7 s 25.9 s 9.6 s

Why is -m0 (no compression at all) not faster than -m1? Because the limit is not the
compression: it is the cutting of the files into fragments, done by one single thread.
That is exactly what -turbo attacks.

And, as said at the top, the archives are identical: the same file, not just the same size.
That also means that deduplication against your old archives keeps working exactly as before.

Details: Part two, 2.1, Part three, 3.3.

1.2 -turbo: a even faster

zpaqfranz a z:\backup.zpaq c:\vm -turbo

-turbo spreads the slowest part of a over all the cores: cutting the
files into pieces (the deduplication fragments) and computing their SHA-1.
It only does it on files of 4 MB or more, where it pays. The small ones go the
usual way.

The archive is the same one with or without -turbo. It is experimental because it is new code, not because it writes something different:
the autotest checks it every time (work fc of the two archives).

How much it gives. The gain comes from working inside one file, so it grows with the size of the file:

  • huge files (virtual machines, disk images, databases) on machines with many cores: the best case. About 3 times faster with -m0, -m6, -m7; 1.5 to 1.7 times with -m1, which then becomes limited by the compression itself;
  • medium files (tens of MB): some gain;
  • many small files: little or nothing;
  • slow methods (-m2 and up): little, the compressors are the bottleneck, not the fragmenter;
  • slow disks: the disk is the limit, and -turbo cannot help.

The price is more total CPU time (about +50%: with -m1 on the virtual machine 154 s of CPU became 225 s): many threads
cutting at the same time. Wall time goes down a lot, CPU time goes up somewhat.

With one thread (-t1) there is nothing to spread, and the usual code runs. The ESXi build does not have
-turbo (it says so, and does a plain a).

Details: Part two, 2.2, Part three, 3.2.

1.3 -m6 and -m7: fast compression that any zpaq can extract

Two new methods, meant for speed, for the big backups where -m1 takes too
long or too much CPU.

zpaqfranz a z:\vm.zpaq c:\vm -m7        :: LZAV
zpaqfranz a z:\vm.zpaq c:\vm -m7h       :: LZAV "hi": smaller, slower
zpaqfranz a z:\vm.zpaq c:\vm -m6        :: LZ4
zpaqfranz a z:\vm.zpaq c:\vm -m6h9      :: LZ4 HC, level 9

The one to try is -m7: about as fast as -m1, a few percent bigger, and
it uses a small fraction of the CPU. With -turbo it becomes the fastest way
to put a big virtual machine into an archive:

(same data as 1.1) time archive
-m1 25.3 s 6.10 GB
-m7 28.8 s 6.46 GB
-m7 -turbo 9.6 s 6.46 GB
-m7h -turbo 14.8 s 6.23 GB

The important thing: any zpaq can extract them, and quickly. Not just this
zpaqfranz: older zpaqfranz versions, and the original zpaq 7.15 from 2016, too.
The decoder travels inside the archive. This zpaqfranz simply recognizes it and
uses a native decoder instead, which is faster still.

Both are marked experimental: try them, and keep -m1 (or whatever you
use now) for the backups that matter until you trust them.

Details: Part two, 2.3, Part three, 3.4.

1.4 mount on FreeBSD, OpenBSD and macOS

mount (an archive as a read-only drive or folder) was Windows and Linux.
Now:

  • FreeBSD (fusefs-libs3)
  • OpenBSD (the FUSE that comes with the system, nothing to install)
  • macOS, with FUSE-T (no kernel extension) or macFUSE
zpaqfranz mount /backup/nas.zpaq /mnt/zpaq
zpaqfranz mount ~/nas.zpaq                   :: macOS: makes the folder, opens the Finder

Ctrl+C now reliably unmounts on every system. Previously, on Linux, it left
behind a dead mountpoint ("Transport endpoint is not connected").

The mount is an option you build in (-DZPAQMOUNT); how to tell whether your executable has it is
explained in 1.8. How to build it on each system: "Building with Mount
Support", at the top of the source.

Details: Part two, 2.4, Part three, 3.5.

1.5 mount reads big files much faster

The cache of mount has been rewritten. Previously, with many cores, it threw a
block away the moment it had decompressed it, and decompressed it again at
the next read. A 531 MB file took 17 seconds and 158 decompressions of just
24 blocks, and one packed with -m2 was still being read after 13 minutes. Now
every block is decompressed once.

A program that jumps around inside a file (a video player, a database)
no longer makes the mount decompress blocks that nobody will read.

1.6 mount -test: does the mount work here?

zpaqfranz mount z:\backup.zpaq z:\mnt -test

Mounts, reads every file through the mountpoint, checks it against the
hash stored in the archive, reads pieces at random positions, unmounts, and
says Mount test: OK (exit code 0) or what is wrong. Handy before trusting a
new machine, a new FUSE, a new container.

1.7 Hashes that heal themselves, a better -touch

When only the attribute of a file changed (read-only, archive bit,
permissions), a stores a new version without reading the file again, and
carries the hash over from the previous version. But if that previous hash
was not usable (an archive made by zpaq 7.15, or with another hash
algorithm, or a zero hash written by an old version), v would say FAILED for
that file forever. Now that file is read again and its hash stored. Same
content, so no space is used: one read, once.

-touch does the same thing on request, on files that did not change at all:

zpaqfranz a z:\old715.zpaq c:\data -touch

Previously, it faked a date change on every file, re-read and re-stored
everything, and needed a second run to put the real dates back. Now it reads only
the files without a usable hash, stores them once, with their real date, in
one run. Also handy to switch to another hash:

zpaqfranz a z:\archive.zpaq c:\data -touch -sha256

Details: Part two, 2.5, Part three, 3.7.

1.8 Smaller things

  • Windows: c -checksum, r and others no longer crash or quit when the
    output is redirected to a file (a scheduled task, > log.txt).
  • Windows: extracting to -to ./somewhere no longer creates a folder
    called _.
  • OpenBSD: the JIT works there too (it was off: everything was
    interpreted). Extracting a -m5 archive went from 93 to 56 seconds.
  • A build without optimizations (no -O) says so at startup, because it is
    5 to 10 times slower.
  • -touch -summary no longer takes -summary as the date of -touch.
  • The "open" version (no network, no external components) no longer contains
    anything about mount (FUSE, WinFsp).
  • The banner tells how your executable was built. Look at it carefully:
    in zpaqfranz v65.3-JIT67+M,... the JIT means that the JIT is there (so an Intel/AMD
    64-bit build), the 6 and the 7 mean that both -m6 and -m7 are built in, and the +M
    means that mount is active. Please always include the banner line when you report something.

Part two, the details, for power users

  1. Faster a without -turbo
  2. -turbo
  3. -m6 (LZ4) and -m7 (LZAV)
  4. mount
  5. Hash healing, -touch
  6. Fixes
  7. The autotest
  8. Building

2.1 Faster a without -turbo

Four changes in the compression library. None of them changes a single
byte
that is written:

  • The LZ77 hash table of -m1/-m2 (64 MB) is now zeroed on purpose
    when it is allocated. Previously, it was read before it was ever written: every
    first touch of a page mapped the shared zero page, and the first write then
    had to copy it and flush the TLB of every core (copy-on-write). On Linux,
    -m1 on 4 GB with 4 threads spent 40 of its 160 CPU seconds in the kernel
    just for that.
  • Prefetch: the hash of the position 32 bytes ahead is computed as well,
    and its cache line is requested from RAM before it is needed. -m1 on
    4 GB, 4 threads, Linux: 42.4 s and 163 CPU seconds became 15.9 s and 62.
  • Stored blocks (-m0, and the new -m6/-m7) are copied a buffer at a
    time instead of one byte per call (a quarter of the CPU of -m0 was there),
    and the archive writer does the same.
  • lg() is a single instruction (__builtin_clz).

What is left is the main thread of add(): cutting the files into fragments,
one byte after the other. That is what -turbo is about.

More on this in Part three, 3.3.

2.2 -turbo

  • Only for plain files of 4 MB or more (ZPAQ_TURBOMIN, can be changed
    at build time with -DZPAQ_TURBOMIN=...). Measured: below 2 MB it is slower
    (up to -25% on 256 KB files), from 4 MB on it is always faster (+24% -m1,
    +48% -m7 at 4 MB already).
  • Not for stdin, -image, memory files, -debug -zero, or with one thread
    (-t1): there the usual loop runs.
  • Memory: two batches of 64 MB (+ one fragment), allocated at the first big
    file only.
  • -verbose tells how it went:
    -turbo 3.362 pieces, 12.407.896 bytes cut on the main thread (on the data of the first table of
    1.1: only 12 MB had to be cut serially).
  • Not in the ESXi (ANCIENT) build: there -turbo says so and runs a plain a.
    The NAS builds have it.
  • The price is more total CPU (about +50%: -m1 on the virtual machine, 154 s of CPU became 225 s):
    many threads cutting at the same time, often on SMT siblings, which are slower per
    thread than one thread alone, plus a small overlap where two chains cut the same bytes.
    Wall time goes down a lot, CPU time goes up somewhat.

Today -turbo is an option, and add() is untouched: -turbo runs
add2(), a copy of it. The plan is to make it the default, and remove the
old path, once it has been proven enough.

How it works: Part three, 3.2.

2.3 -m6 (LZ4) and -m7 (LZAV)

Syntax:

switch what
-m6 LZ4 default, 16 MB blocks
-m6h9 LZ4 HC, level 1..12 (9 is the sweet spot)
-m6a8 LZ4 fast, acceleration N (bigger = faster, larger)
-m66h9 the digit after the 6 is the block size: 2^6 = 64 MB
-m7 LZAV default, 16 MB blocks
-m7h LZAV "hi": smaller, slower
-m76 LZAV with 64 MB blocks
  • Why LZAV beats LZ4 here: LZ4 looks back 64 KB. LZAV looks back 2 MB, and
    farther (up to 256 MB, within the block) when the offset carry allows it.
    So -m7 is smaller than -m6, and just as fast.
  • A block that does not shrink is stored, as -m1 does.
  • Extraction: zpaqfranz 65.3 decodes natively (faster than -m1). Any other
    zpaq runs the decoder written in ZPAQL that is inside the block: fast too,
    only the native one is faster still. It is also how p -verify double checks (see
    Part three, 3.4).
  • -m6/-m7 can be switched off at build time: comment out #define ZPAQLZ4
    / #define ZPAQLZAV at the top of the source (then they are -m5 again,
    as in older versions). They are off by themselves in the ANCIENT/ESX builds.

2.4 mount

Where it builds:

system FUSE build (see the top of the source)
Windows WinFsp (DLL loaded at run time) -DZPAQMOUNT -I".../WinFsp/inc"
Linux libfuse3 -DZPAQMOUNT $(pkg-config fuse3 --cflags --libs)
FreeBSD fusefs-libs3 + kldload fusefs same, the -I of pkg-config is NOT optional
OpenBSD FUSE 2.6 of the base system -DZPAQMOUNT -lfuse, root only
macOS FUSE-T or macFUSE -std=c++11 -DZPAQMOUNT -I/usr/local/include/fuse3 ... -lfuse3

Unmounting: Ctrl+C, kill (SIGTERM), SIGHUP, and umount from another
terminal all end with zpaqfranz quitting by itself and nothing left mounted.
If something still is (a kill -9, an exit from somewhere else), an atexit
hook tries fusermount3 -u -z, fusermount -u -z, umount -l, umount,
umount -f, in this order, and checks after each one. The hint printed is the
right one for the system (fusermount3 -u on Linux, umount elsewhere).

macOS: no mountpoint (or *) = one is made (/Volumes/<archive> as root,
otherwise in $TMPDIR), the Finder opens on it, and it is removed at the
end. The Finder shows the name of the archive (-fuseopt volname=X for
another one).

Speed:

  • one cache for the whole budget (ZPAQFUSE_CACHE_MB, default 256), the
    least recently used block goes first, and never the one just made.
    Decompression happens outside the lock, on ZPAQFUSE_SHARDS readers. A block
    wanted by two threads is decompressed once.
  • read-ahead (4 blocks, never more than half of the cache) only for somebody
    reading the file from one end to the other.
  • the position inside a file is found with a binary search, set up once
    at open(). Previously, the fragment list was walked from the start at every
    read, which cost minutes on a 100 GB file.
  • macOS: files declare 1 MB blocks (st_blksize): macFUSE reads in chunks of
    that size, 4 KB meant 12 s instead of 1.2 s for a 673 MB file.

Measured: FreeBSD 14.2, a 337 MB file in 2.0 s with one decompression per
block. macOS 12.7, 673 MB in 1.1 s (macFUSE) and 2.5 s (FUSE-T). OpenBSD 7.9,
248 MB in 1 s (it was 12, before the JIT worked there).

mount -test checks: every file of the mounted version is there, with its
size and date (2 s tolerance), nothing more is there (the tree is walked), the
content matches the stored hash (XXHASH64, SHA-256, whatever the archive has;
size and date only for 7.15 files), and 8 random pieces per file are
identical to a sequential read. With -all, every version folder can be walked. The
list of what is expected comes from the index read the way x and v read
it, not from the mount. So the two readers check each other.

More on this in Part three, 3.5.

2.5 Hash healing, -touch

  • a: a file with the same date and size, a different attribute, and no
    usable previous hash (all zeros, another algorithm, 7.15) is read again.
    -debug3 prints no usable hash (attribute-only change), reading again.
  • -touch works on the same condition, without the attribute change. Only
    files without a usable hash are read, the real dates are kept, one run.
    -touch followed by another switch (-touch -summary) is no longer taken
    as -touch <date>.
  • Autotest: out51..out55 (attribute change, v, -715 then
    -touch -sha3, v, -touch again = nothing to do).

More on this in Part three, 3.7.

2.6 Fixes

  • Windows, output redirected: setupConsole() used to exit(6) when stdout
    was not a console, and terminalwidth()/terminalheight() read an
    uninitialized structure (garbage, then a crash in c -checksum). Both fixed.
  • franzsanitizepath(): a . component followed by something (./locale/x)
    is dropped instead of becoming _. .. stays neutralized.
  • myaddfile(): no more negative row count turned unsigned on tiny consoles.
  • OpenBSD: JIT via mprotect (W^X), see Part three, 3.6.
  • -DDEBUG and -DNOLM build again (they were broken in 65.2 too).
    With -DNOLM, custom_log2() no longer calls log2().
  • ESX/NAS/ANCIENT: zip -deflate is not available there (clear message
    instead of a build error), -huge checks under #ifndef ANCIENT.
  • License list: entry 26 is WinFsp (it said ascii-art).
  • ESXi build (gcc 3.4.6, C++98): it did not compile any more. The members
    of Jidac had in-class initializers (now set in a constructor), and a few
    newer bits (std::unique_ptr, vector::data(), >> closing two templates)
    are now written the C++98 way. Same code, same behaviour elsewhere.
  • Solaris: the system headers define _LITTLE_ENDIAN as an empty macro,
    so the endianness test inside lzav.h did not even parse ("operator '=='
    has no right operand"). The three macros are pushed away while lzav.h is
    read and given back right after (#pragma push_macro/pop_macro, only on
    __sun); lzav.h itself is untouched.
  • ARM, unaligned CRC-32: crc32_16bytes() read the buffer through a
    uint32_t*. On ARMv7 gcc may merge those loads into one LDM, which does
    not tolerate an unaligned address: t and x died with "Bus error" on a
    Cortex-A9 NAS (and on ARMv5 an unaligned load silently returns rotated
    data). Now memcpy: the same single load on x86/aarch64, same values,
    same speed.
  • Thread stack: threads were created with the libc default stack: 8 MB
    with glibc, but 128 KB with musl, 512 KB on macOS, 1-2 MB on FreeBSD.
    -m6/-m7 (hash tables on the stack) segfaulted with musl. Now the
    compression and extraction threads get at least 8 MB (virtual memory).
  • Autotest with -franzen: out30/out38 compared the decoded .franzen
    with a cleartext .zpaq that, since the Franzen-only archives of 65.1, is
    no longer written, so they always failed. Now they test the decoded archive
    (t, with -key for the AES one).
  • a -franzen when the .franzen cannot be made (the typical case:
    libsodium not installed; with libsodium nothing changes): it printed
    "cannot create franzen" and went on. With a Franzen-only archive (the
    default since 65.1) nothing at all was written, yet a ended with "all OK"
    and exit code 0. Appending to an existing .zpaq updated it and left the
    .franzen behind. Now a stops with an error (exit code 2) before writing
    anything.

2.7 The autotest

autotest -all -to <folder> now also covers:

out what
51-55 hash healing, -touch
64 names with UTF-8, CJK, emoji, %, %3A, spaces, empty ones
56-63, 65-67 mount -test of the archives above (only if a real mount works on this machine)
68-76 -m6, -m6h9, -m66h9: a, t, p -verify (ZPAQL), w, mount
77-82 -m7, -m7h: the same
83-86 -turbo: the archive must be identical to the one of plain a

The script does a cd into the test folder as its first line (some steps use
relative paths), so it can be started from anywhere.

2.8 Building

  • New defines, on by default: ZPAQLZ4 (-m6) and ZPAQLZAV (-m7). Comment
    them out at the top of the source to build without. They stay in the -DNAS
    builds (made with modern compilers), not in ESX/ANCIENT. -turbo follows the same rule:
    present in the NAS builds, not in the ESXi one.
  • The PCOMP are frozen: the SHA-1 of the bytecode of LZ4_PCOMP and
    LZAV_PCOMP is written in the source. autotest checks it first of all
    (and says "Do NOT release this build" if it changed), and -m6/-m7 refuse
    to write a block if it does not match.
  • NAS binaries: static, musl (they start on 2.6.x kernels too), -DNAS,
    for x86_64, i686, ARMv5, ARMv7 (generic, Cortex-A9, Cortex-A15) and ARMv8
    (generic, A53, A55, A57, A72), each one tested under qemu on its own core.
    They include -turbo, -m6 and -m7. See NAS_BINARI.txt.
  • Tuning knobs: -DZPAQ_TURBOMIN=<bytes> (default 4 MB), -DZPAQ_LZAHEAD=<n>
    (default 32, the prefetch distance of -m1).
  • Checked with -Wall: Windows (mingw g++), Fedora 44 (g++ 16.2) with plain,
    -DHWSHA2, -DZPAQMOUNT, -DNAS, -DESX, -DANCIENT, -DBIG,
    -DALIGNMALLOC, -DNOJIT, -DDEBUG, -DNOLM, -DIPV6, -DSFTP, -DOPEN,
    -DOPEN -DZPAQMOUNT, and the purged "open" source with and without
    -DZPAQMOUNT.
  • Also built and tested (autotest, then a/t/x and compare with every
    method and hash): ESXi (-DESX, gcc 3.4.6 on CentOS 3.9, static) and
    Solaris 11.4 (-DSOLARIS, gcc 7.3, amd64:
    g++ -O3 -DSOLARIS zpaqfranz.cpp -o zpaqfranz -pthread -static-libgcc -lkstat).
    On openSUSE Tumbleweed (g++ 13.2, libfuse 3.16.1) the full
    autotest -all passes, mount and Franzen included.
  • macOS universal binary (Intel + Apple Silicon in one file), without
    mount, so it starts on any Mac with no FUSE installed. The two halves are
    built apart, because only the Intel one can have the JIT, then joined:
  clang++ -arch x86_64 -mmacosx-version-min=10.15 -O3 -std=c++11 -Dunix -DHWSHA2 zpaqfranz.cpp -o zf_x86 -pthread
  clang++ -arch arm64  -mmacosx-version-min=11.0  -O3 -std=c++11 -Dunix -DNOJIT  zpaqfranz.cpp -o zf_arm -pthread
  lipo -create zf_x86 zf_arm -output zpaqfranz && strip zpaqfranz

No warnings with -Wall on either half. It runs on macOS 10.15 or later
(Intel) and 11.0 or later (Apple Silicon). The Intel half was tested on
macOS 12.7 (autotest, then a/t/x and compare with every method and
hash). The Apple Silicon half is built but not run here: there was no ARM
Mac to try it on.

  • The open version: everything about mount (FUSE/WinFsp headers, the
    command, its help, its autotest, the build notes, the libfuse and WinFsp
    licenses) is now between ///NOSFTPSTART and ///NOSFTPEND, so the purge
    removes it. -DOPEN without the purge does #undef ZPAQMOUNT too.

Part three, lo spiegone, for developers

  1. How a cuts a file, and why that matters
  2. -turbo: the same cuts, in parallel
  3. The libzpaq changes (default path, no byte changed)
  4. -m6 and -m7: a decoder inside the archive
  5. mount, inside
  6. The JIT on OpenBSD (W^X)
  7. Hash healing, inside
  8. The open version

3.1 How a cuts a file, and why that matters

zpaq does not deduplicate files, it deduplicates fragments. Every file is
cut into pieces of variable size, each piece is identified by its SHA-1, and a
piece already in the archive is not stored again, just referenced.

Where to cut is decided by the content, not by the position. A rolling hash
runs over the bytes, together with an order-1 predictor (o1[c1]: "after
byte c1, last time came byte o1[c1]"). In pseudocode:

for every byte c:
    hit = (c == o1[c1])
    h   = (h + c + 1) * (hit ? 314159265 : 271828182)
    o1[c1] = c;  c1 = c
    cut if  (h < 2^(22-fragment) and size >= MIN_FRAGMENT)  or  size >= MAX_FRAGMENT

With the default -fragment 6: a cut every 64 KB on average, never under
4 KB, never over 520,192 bytes. Because the cut depends on the content, inserting
a byte at the start of a file moves the cuts near the insertion, and after a
few fragments they fall exactly where they were before. The rest of the file
deduplicates.

This is also why the algorithm cannot be changed: other cuts mean other
fragments, other SHA-1, and no deduplication at all against every archive ever
written. Anything faster has to produce the same cuts.

The predictor also feeds the type analysis (hits, the o1 table: text?
executable? already compressed?) that picks the model for the block. So that
has to be identical too.

3.2 -turbo: the same cuts, in parallel

Why it is needed

Compression runs on many threads. The fragmenter does not: it runs on the main
thread, one byte at a time, together with the SHA-1 of the fragments and the
hash of the whole file. With the fast methods (-m0, -m1, -m6, -m7) the
compressors are waiting for it. perf, -m7 on a VM disk, before -turbo:

33.93%  Jidac::add()          the fragmenter: rolling hash, one byte at a time
23.72%  lzav_compress         worker threads, NOT on the critical path
19.88%  sha1_transform_x86    mostly the SHA-1 of every fragment

So a 16/32 core machine spends most of its time waiting for a single core, at
about 500 MB/s.

Guess, then prove

The fragmenter looks serial: h carries from one byte to the next. But
its state starts fresh at every cut: after a cut, h = 0, c1 = 0, o1
all zeros. That is the whole trick: everything after a cut depends only on the
bytes after that cut.

add2() reads a file in batches of 64 MB. Each batch is split into
N pieces (one per thread, at least 1 MB each). Every thread cuts its piece
as if a fragment began exactly at its start: that is a guess. The first
piece is not a guess, because the batch begins on a real cut. Each thread also
computes the SHA-1 of each fragment while it is still in the cache, so the
SHA-1 goes parallel too, for free.

Then the true chain is followed. It starts with the fragments of piece 0.
When the true chain ends a fragment exactly where the chain of piece 1 also
has a cut, from that byte on both start fresh on the same bytes: they are
the same chain
, and the guess of piece 1 is proven right. The chains of a
rolling hash synchronize by themselves, and usually within a fragment or two.

piece 0     |--a--|---b---|--c--|-d-|--e--|--f--|
piece 1               |--x--|---y---|--e--|--f--|    (starts from a guess)
                                    ^
                     both chains cut here, at the start of e:
                     from this byte on, piece 1 IS the true chain

If they do not meet (rare), the true chain goes on by itself (add2_scan
with dopo=1, one fragment at a time on the main thread) until they do. Each
piece goes 3 cuts past the start of the next one, to leave room for the
meeting. Nothing is trusted that was not proven. The -verbose counter says
how much was serial: 12 MB, on the data of the first table of
1.1, in 3,362 pieces.

The fragment still open at the end of a batch (less than MAX_FRAGMENT) is
copied in front of the next batch (TURBOROOM), so every batch starts
on a cut.

The rest of the pipeline

reader thread     reads batch N+1 (64 MB), updates the hash of the whole file
cutter threads    cut batch N into fragments + SHA-1 of every fragment
main thread       merges the chains, then for every fragment, IN ORDER:
                  lookup, type analysis, new block decision, sb.write
compressors       as always

Every fragment, in order, goes through elabora: a lambda holding the
same code as the loop of add() (lookup, type analysis, new block
decision, sb.write), with the data taken from the batch instead of
fragbuf. It was compared line by line with the original. Same fragments,
same order, same decisions: same archive.

Around it:

  • double buffering: a reader thread (add2_readthread) reads the next
    batch while the current one is cut, and updates the hash of the file
    (XXHASH64, SHA-256...) on it. One batch at a time, in order, so the file hash
    sees the bytes in the same order as add().
  • branchless fragmenter: (uc == o1[c1]) as a value, not as an if. The
    first version, with the if, was slower than the original: on compressed
    or random data the branch is a coin toss, and a mispredicted jump on every
    byte costs more than it saves.
  • files under ZPAQ_TURBOMIN (4 MB), and one thread, take the old loop:
    starting threads for a small file costs more than it saves.

Why a separate add2() and not a patch to add()

add() is the most important function of the program, full of special cases
(stdin, -image, memory files, NTFS/VHD imaging, -zero...). add2() is a
copy with a different loop, for plain files only, selected by -turbo.
add() is not touched at all, so if anything is wrong you just do not use
-turbo. Stdin, memory files, images and -zero go through the original loop
even inside add2().

Verified

Identical archives on Windows for -m0, -m1, -m7, -m7h
(11.8 GB VM + 1,238 mixed files, two versions), identical to 65.2k for
-m0/-m1. A matrix of 23 cases, two versions each, -turbo against plain
a, all identical to the byte: every hash (-sha256, -blake3, -xxh3,
-sha3, -md5, -whirlpool, -crc32, -nochecksum, -715), -verify,
-t1, -t2, -m0, -m2, -m3, -m6, -m66h9, -m7h, -m76,
-fragment 4, -fragment 9, -nodedup, -store. And in the autotest
(Windows and Fedora, 86/86), every time.

3.3 The libzpaq changes (default path, no byte changed)

The COW trap. LZBuffer allocates the LZ77 hash table (ht, up to
64 MB for -m1) with calloc(). For an allocation that big, calloc returns
fresh pages that were never touched: all mapped to the shared zero page.
fill() reads ht[h] before it writes it. The first read of each page
maps the zero page, and the first write then has to copy it (copy-on-write)
and invalidate that mapping on every core running the process (a TLB
shootdown, an inter-processor interrupt to each core). With many threads
each doing it on its own table, that was 25% of the CPU of -m1, in the
kernel. A memset right after the allocation writes every page once. After
that they are ordinary pages.

Prefetch. Every lookup of ht is a cache miss (a 64 MB table, random
index): perf showed 40% of fill() on the single line p=ht[h1^k]. The hash
of a position is known well before it is needed: a second rolling hash,
hpre, runs ZPAQ_LZAHEAD (32) bytes ahead and calls
__builtin_prefetch(&ht[hpre]). It needs no care: after minMatch steps the
old bytes are shifted out of the mask, so it is right by itself, and a wrong
prefetch only costs a little bandwidth. A bucket is 16 words = 64 bytes = one
cache line: one prefetch is enough. Measured distances: 1 → 5%, 8 → 18 s,
16 → 16.2 s, 32 → 15.9 s, 64 → 16.2 s.

Stored blocks. Encoder::compress(c) with no model appends one byte per
call to a buffer, and flushes it (with a 4 byte length prefix) when full.
compressstored(p, n) does the same with memcpy, flushing at exactly the
same points
, so the output is the same. Compressor::compress() uses it
when there is no model and no verify. OutputArchive::write() does the same
for put(): the flush points (and so what encryption and -chunk see) are
the same.

3.4 -m6 and -m7: a decoder inside the archive

A zpaq block starts with its own description (comp / hcomp / pcomp):

comp hh hm ph pm n     sizes of H and M arrays, n = number of models
hcomp ...              ZPAQL program computing contexts for the models
pcomp ...              ZPAQL program run on the decoded bytes (post-processor)

Every zpaq decoder, 7.15 included, must run the pcomp program on the
decoded bytes, and what it outputs is the file. That is the whole idea:

  • n = 0: no model at all. The arithmetic coder then stores the bytes as
    they are (like -m0). Decoding costs nothing.
  • The stored bytes are an LZ4 (or LZAV) stream, compressed by zpaqfranz
    with the real library.
  • The pcomp is an LZ4 (LZAV) decoder written in ZPAQL, a few dozen
    instructions. Any zpaq runs it and gets the original data back.

So the format of the archive does not change at all: it is a valid zpaq
block, with a post-processor, which zpaq has always supported (that is how
-m3/-m4 do BWT and E8E9).

The native shortcut. Running the ZPAQL decoder is already fast (any zpaq
does it, better with the JIT), but a native decoder is faster still. So when
the pcomp of a block arrives, its compiled bytecode is
compared
with the one of LZ4_PCOMP/LZAV_PCOMP (lz4_is_canonical,
compiled once, thread-safe static). If identical, the stored bytes are
collected and handed to ZLZ4_decompress_safe() / lzav_decompress_3()
instead. The output is the same: the ZPAQL program defines what is correct.
A block that is not a valid stream is an error natively, and would be
garbage through ZPAQL. That is why not a single byte of those programs may
ever change
: old archives carry the old bytecode. (It happened once, in
65.3w, when four redundant instructions were removed from the LZ4 decoder:
archives made by earlier 65.3 test builds still extract, through ZPAQL.)

-m6, LZ4. comp 0 0 0 16 0: pm = 16, M is a 64 KB ring buffer, exactly
the LZ4 window (offsets are 16 bits). The decoder is a small state machine in
r1: token → literal length (15 = more bytes follow, 255 = go on) →
literals → 2-byte offset → match length → copy from M[b - offset].

-m7, LZAV (format 3). LZ77 with a 2 MB window that the offset carry
(below) stretches up to 2 MB << 7 = 256 MB: an offset can point anywhere in the
block. So comp 0 0 0 (20+B) 0: M holds the whole block (16 MB with the
default), not a ring buffer as in -m6. Harder
to decode in ZPAQL:

  • 4 block types (literals, references with 10, 15, 21 bit offsets);
  • lengths in 7-bit groups;
  • offset carry: the low bits of an offset come from blocks before it
    (the top bits of the last offset byte of the previous reference, the CC
    bits of a literal block in between). A decoder that reads in order always
    has them: r3 (carry value) and r4 (carry shift), then
    offset = (own bits << csh) | cv.
  • Trap: ZPAQL if is a short jump (±127 bytes). The body of the offset state
    is longer, so it needs ifl (long if). Otherwise: "IF too big".

The ZPAQL decoders are tested every time by the autotest: p -verify is
unzpaq, a separate decoder that runs the ZPAQL, and compares with the
files on disk (out70, out73, out75, out79, out81).

LZ4 1.10 and LZAV 5.17 are included unmodified, each in a namespace of its
own (zlz4, zlzav). zpaqfranz had its own LZ4 1.9 (decompression only,
Windows only, for the file lists in NTFS ADS): it is gone, the 1.10 does
that too, and ADS written by the 1.9 are read by the 1.10 and vice versa
(checked). The LZ4 names were renamed ZLZ4_ mechanically, and 36 generic
macros are saved and restored with push_macro/pop_macro.

3.5 mount, inside

Layers

Jidac::mount()           arguments, backend, FUSE loop, -test
franzmount FUSE layer    getattr / readdir / open / release / read / statfs / init
franzmount generic       paths, names, dates, MountCache, mount_read_range()
MountBackend "jidac"     versions, children, entries, fragment -> block, decompression

The cache, before and after

Before: N shards (N = cores), each with its own reader, mutex and LRU,
block → shard by hash of the offset. 256 MB on 16 cores = 16 MB per shard,
less than one block (~20 MB with -m1, 64 MB and more from -m2 up).
Each block was evicted the moment it was inserted. On top of that, the shard
mutex was held during the decompression, so a read of a cached block waited
behind the decompression of another block that happened to land in the same
shard.

Now (MountCache):

  • one unordered_map + one LRU list for the whole budget, under one mutex;
  • an entry can be loading: a second thread wanting the same block waits on
    a condition variable instead of decompressing it again;
  • decompression runs outside the lock, on a pool of readers (each with its
    own archive handle). A prefetch leaves the last free reader to a
    foreground read;
  • the block just made is never evicted, even if bigger than the budget;
  • data handed out as shared_ptr<const string>: a 4 KB read of a 16 MB block
    costs no copy, and eviction during a read is harmless.

A file, once opened

open() resolves the path once and builds starts[], the offset of every
fragment. fh points to a MountFile. read() finds the first fragment with
upper_bound (binary search) instead of walking the list. release() frees it.
The size of the file is starts.back().

Read-ahead only for a sequential reader: the first read at offset 0, or a
read starting within MOUNT_SEQ_SLACK (8 MB) of where the previous one ended.
The slack is needed: kernels read ahead by themselves, from several threads,
and requests do not arrive in order (macFUSE asks a 673 MB file as 1 MB
requests from a handful of threads). At most 4 blocks, at most half the cache.

Portability, one system at a time

  • Headers: __has_include(<fuse3/fuse.h>) first, then <fuse.h>.
    FreeBSD installs into /usr/local/include/fuse3, which the base clang does
    not search. A FUSE 2 header picked up by mistake is an #error with the
    install hint, not a flood of errors.
  • OpenBSD: FUSE 2.6 in the base system, no libfuse3 at all
    (ZPAQMOUNT_FUSE2). Three differences: getattr without fuse_file_info,
    filler without flags (the MOUNT_FILL macro), init without
    fuse_config. Directories report st_nlink = 1: with 2, find (fts "leaf
    optimization") concluded there was nothing below and did not descend.
  • macFUSE: its "Darwin extensions" change the prototypes of getattr,
    readdir, statfs: FUSE_DARWIN_ENABLE_EXTENSIONS 0 before the include.
    st_blksize 1 MB on macOS only.
  • Signals: libfuse installs its SIGINT/SIGTERM/SIGHUP handlers only if
    the current ones are SIG_DFL
    (set_one_signal_handler). zpaqfranz had its
    own SIGINT handler, so libfuse left it alone, Ctrl+C called exit(1), and
    the mountpoint stayed dead. Now, for the time of the mount, the three
    signals go back to SIG_DFL, then are restored.
  • SIGPIPE on FUSE-T: the helper (go-nfsv4) goes away first, libfuse
    resets SIGPIPE to SIG_DFL, then fuse_unmount() writes to the closed
    socket: death by SIGPIPE (exit 141) at every unmount. SIG_IGN is not enough
    (libfuse takes it as its own and resets it). A handler that does nothing is
    neither, and libfuse leaves it alone.
  • OpenBSD Ctrl+C: the loop sits in a read of /dev/fuse that signals do
    not interrupt. So there the handler runs umount on the mountpoint
    (fork + exec are async-signal-safe, the path is in a plain char[]), and the
    loop ends by itself.
  • Relative archive names: fuse_daemonize() does chdir("/") even in
    foreground mode, and the readers open the archive lazily, at the first read.
    So mount pippo.zpaq /mnt showed the tree and gave EIO on every byte. The
    name is made absolute before mounting (not with realpath(): it may hold
    wildcards).
  • Birth time: impossible outside Windows. struct fuse_attr has no field
    for it, so the kernel reports "unknown".

mount -test and the WinFsp trap

mount -test runs the check on a thread while fuse_main() serves. To end
the loop from inside, the book says fuse_exit(fuse_get_context()->fuse). In
WinFsp the context of a callback is the FUSE 2 one (fsp_fuse_op_enter:
context->fuse = f), while fuse_exit is fuse3_exit, which takes it for a
struct fuse3 and follows a field of it: a crash, or a random handle
signalled. The FUSE 2 fsp_fuse_exit() fits that pointer and only does
SetEvent() on the loop. It is not in the fuse3 headers, so it is bound by hand,
like the rest of the DLL. The check thread raises g_mount_exit and asks for
a name that does not exist (a different one each time, so no cache answers).
The getattr that comes calls fsp_fuse_exit.

On *nix it simply unmounts, and mount_still_mounted() tells whether it
worked: st_dev of the mountpoint versus st_dev of its .. (no /proc,
so it works on BSD and macOS too). ENOTCONN counts as "still mounted".

Never kill a process with I/O pending on its own WinFsp volume: it
becomes a zombie in the kernel, and the volume stays. The autotest runs the
trial mount in a child process with a timeout and leaves it alone if it is
too slow, instead of killing it.

3.6 The JIT on OpenBSD (W^X)

OpenBSD refuses memory that is writable and executable at the same
time
, which is what the JIT asked for (mmap(PROT_READ|PROT_WRITE|PROT_EXEC)).
So it turned itself off and everything ran interpreted. The rule is about the two
rights together: written first, then turned executable with mprotect(), is
allowed, even without wxallowed. Now allocx() tries the usual way first
and, if refused, maps PROT_READ|PROT_WRITE and sets g_jit_mprotect.
protectx() flips the code to PROT_READ|PROT_EXEC once it is written, right
before the first run. If even that fails, the interpreter takes over, as
before. On the other systems nothing changes. OpenBSD 7.9: -m5 extraction
93 → 56 s, -m1 3 → under 1 s.

3.7 Hash healing, inside

The loop of add() that decides what to read looks at date and size. The
loop that writes the index (what to store) also looks at the attribute.
So an attribute-only change stored a new version of a file that was never
read, with the hash carried over by carryoverhash(). If that hash was all
zeros (older versions), of another algorithm, or missing (7.15), the zero
placeholder was carried on forever.

Now, in the first loop: same date, same size, and (attribute changed, or
-touch): carryoverhash() is asked. If it returns nothing usable,
rereadforhash makes the file be read, and DT::rehash tells the second
loop to store it even if nothing else changed. Same content → same fragments
→ no new data, one read.

3.8 The open version

The "open" source is produced by removing every line from a
///NOSFTPSTART to the next ///NOSFTPEND (inclusive). #define ZPAQFULL is
on such a line, so it disappears too. The purge is line based and does not
nest: a START inside a region that is already being removed would be
swallowed, and its END would stop the purge early. So the mount blocks
were marked where they are not already inside a purged region (the
mount part of the autotest script already was).

Marked now: the FUSE/WinFsp headers, the mount declaration, help_mount,
its entry in the help map, +M in the banner, the mount command and its
switches (-fuseopt, -backend), the whole franzmount namespace and
Jidac::mount(), the dispatch, autotest_spawn(), the expected mount results
in checkautotest(), the "Building with Mount Support" chapter and the mount
build lines, the libfuse and WinFsp licenses. Result: 147,895 → 117,866
lines, and no reference left to FUSE or WinFsp except in plain text (the zip
help says it works "without FUSE/WinFsp"). The purged source builds with and
without -DZPAQMOUNT (which then has nothing to switch on).


Appendix: a test on real data, zpaqfranz, 7-Zip, WinRAR and zstd

Just an example, for information: one machine, two folders, the fastest
mode of each program. It says nothing about extraction, and nothing about
what each format can do (versions, deduplication, recovery records...).

As said in the short answer, the result depends on the hardware and on
the operating system, and the aim is not to measure a difference of a few percent
(that would need a specific laboratory), but to show how wide the difference is.

This is not the same data as in 1.1 and 1.3.
There the data is a virtual machine plus 1.1 GB of mixed files. Here there are two dedicated
folders, because zstd compresses one file (or one stream), it cannot compress a folder: it needs
a tar in front. That is why -m1 takes 23.1 s here and 25.3 s there: different data, not a discrepancy.

This is not an equal fight. zpaqfranz cuts every file into fragments, computes the SHA-1 of each
one, deduplicates, hashes every file and keeps the index of all the versions. tar | zstd, 7-Zip and
WinRAR do much less. zstd in particular practically does nothing but compress, and it is fast also for this reason.

The machine: Windows 11, AMD Ryzen 9 7950X3D (16 cores, 32 threads),
127 GB RAM, idle (no virtual machine running). Sources and archives on a RAM
disk (see 1.1).

The data, the two opposite ends:

  • vmtest: a Debian 12 virtual machine (VMware), 16 files, 11.76 GB,
    almost all of it in a single 11.7 GB .vmdk.
  • vmtest2: an MSYS2 installation (compilers, headers, Python, man pages...),
    116,703 files in 5,381 folders, 6.62 GB, median file size 2.8 KB.

The programs, each in its fastest compression mode (store-only modes left
out):

program command
zpaqfranz 65.3 a z:\t.zpaq z:\vmtest -m1 (then -m1 -turbo, -m7 -turbo), -noeta
7-Zip 26.03 (x64) 7z a -mx1 -mmt=on -bso0 -bsp0 z:\t.7z z:\vmtest
RAR 7.23 (x64) rar a -m1 -mt32 -r -idq z:\t.rar z:\vmtest
zstd 1.5.7 (x64) tar -cf - -C z:\ vmtest | zstd -1 -T0 -q -o z:\t.tar.zst (also --fast=1)

For zstd the tar is the one that comes with Windows (bsdtar 3.7.7). The tar alone, with no
compression at all, is measured too, to see how much of the time is its own.

How it was measured: three runs in a row for each, the archive deleted before every run; the tables show the
median. "CPU" is the processor time of all the processes of the command (cmd, tar, zstd...),
summed over all threads: how much of the machine it takes, that is, how much it leaves to everything else.
Everybody can write their own measuring scripts: reproducing the milliseconds is not the point.
When a program goes from 500 MB/s to 1.5 GB/s, a wall clock is enough to notice.

vmtest, the virtual machine (11.76 GB, 16 files)

program time speed CPU archive size
zpaqfranz -m1 23.1 s 509 MB/s 154 s 5.64 GB 48.0%
zpaqfranz -m1 -turbo 13.7 s 861 MB/s 225 s 5.64 GB 48.0%
zpaqfranz -m7 -turbo 5.7 s 2,049 MB/s 53 s 5.97 GB 50.7%
7-Zip -mx1 21.9 s 537 MB/s 579 s 5.74 GB 48.8%
WinRAR -m1 40.4 s 291 MB/s 588 s 6.01 GB 51.1%
tar | zstd -1 8.1 s 1,450 MB/s 22 s 6.22 GB 52.9%
tar | zstd --fast=1 8.1 s 1,461 MB/s 21 s 6.57 GB 55.8%
tar alone (no compression) 8.1 s 1,446 MB/s 8 s 11.76 GB 100%

And zstd without tar, straight on the 11.7 GB .vmdk:
zstd -1 -T0 2.95 s (3,981 MB/s, 17 s of CPU, 52.9%), --fast=1
2.97 s (55.9%).

vmtest2, the small files (6.62 GB, 116,703 files)

program time speed CPU archive size
zpaqfranz -m1 18.2 s 365 MB/s 87 s 1.41 GB 21.2%
zpaqfranz -m1 -turbo 15.8 s 420 MB/s 108 s 1.41 GB 21.2%
zpaqfranz -m7 -turbo 13.8 s 481 MB/s 30 s 1.72 GB 25.9%
7-Zip -mx1 17.2 s 384 MB/s 174 s 1.47 GB 22.1%
WinRAR -m1 56.2 s 118 MB/s 381 s 1.98 GB 29.9%
tar | zstd -1 8.6 s 775 MB/s 18 s 1.88 GB 28.4%
tar | zstd --fast=1 8.5 s 781 MB/s 17 s 2.11 GB 31.8%
tar alone (no compression) 8.7 s 763 MB/s 9 s 6.72 GB 101%

What it says

  • zstd is the fastest compressor here, by far. On the .vmdk alone it
    runs at almost 4 GB/s: twice zpaqfranz -m7 -turbo, seven times 7-Zip,
    with very little CPU. Its fame is deserved. But remember that it is also the program doing the
    least: no fragments, no deduplication, no hashes, no index, no versions.
  • zstd alone is not a backup of a folder. Put tar in front, and
    tar becomes the limit: tar | zstd takes exactly the time of tar alone
    (8.1 s here, 8.6 s on the small files), since tar reads one file after
    the other with one thread, at about 1.4 GB/s on this machine.
  • Virtual machine: zpaqfranz -m7 -turbo is the fastest way to make an
    archive of it: 5.7 s against 8.1 s of tar | zstd (1.4 times faster), 3.8
    times faster than 7-Zip, 7 times faster than WinRAR. Its archive is smaller
    than zstd's (5.97 GB against 6.22 GB), using more CPU than tar | zstd (53
    s against 22 s) but a tenth of 7-Zip's. -m1 writes the smallest archive
    (5.64 GB): about as fast as 7-Zip (5% slower) with a quarter of its CPU, and with -turbo 1.6
    times faster than 7-Zip.
  • Many small files: here tar | zstd wins clearly, 8.6 s against 13.8 s
    of zpaqfranz -m7 -turbo, with less CPU (18 s against 30 s). Most likely zpaqfranz
    spends that time in its own work on each file (path, attributes, hashes of
    every fragment for the deduplication), which tar does not do.
    -turbo works only on files of 4 MB or more, so it helps little here. In
    exchange zpaqfranz -m1 writes the smallest archive of the test, 1.41 GB
    against 1.88 GB of tar | zstd -1 (25% smaller), with about the time of
    7-Zip and half its CPU.
  • WinRAR is the slowest in both cases (about 7 times slower than the
    fastest), with a big archive, and it uses as much CPU as 7-Zip on the
    virtual machine and twice as much on the small files.
  • The second backup is another story, and it is not in the tables:
    tar | zstd, 7-Zip and WinRAR rewrite everything, every time. zpaqfranz
    keeps every version in the same archive and stores only what changed: on
    a virtual machine, typically a small fraction of the disk. That is where
    its time goes, and where it is paid back.
  • On real storage (SSD, NAS, spinning disks) the times get closer to each
    other, because reading becomes the limit: 2-4 GB/s is more than most disks
    deliver. The CPU column stays the same.

Note

65.3 is a big step, and a delicate one on a. Nothing changes unless you
ask for it
: the default a writes the same bytes as 65.2, only faster.
-turbo, -m6, -m7 are opt-in and marked experimental. Please open an Issue on GitHub
for anything strange, with the banner line (zpaqfranz v65.3-JIT67+M,..., see
1.8): it tells which features were built in.

Download zpaqfranz

Don't miss a new zpaqfranz release

NewReleases is sending notifications on new releases.