github KoalaBear84/OpenDirectoryDownloader v4.1.0.0
OpenDirectoryDownloader v4.1.0.0

4 hours ago

Performance and memory

Big directory listings now scan several times faster and use much less memory. Scans of real servers are still limited by the
network, so the largest gains are on very large directories, long scans and --resume.

Benchmarks

Measured against a local test server that generates Apache- and nginx-style listings, on one Windows machine. They are single runs, so a few percent is noise. "Before" is the code at the start of this work.

Scenario Before After
One directory with 100,000 files, scan time 23.6 s 1.9 s
One directory with 100,000 files, peak memory 437 MB 251 MB
4,681 directories / 187,000 files, scan time 19.3 s 7.7 s
4,681 directories / 187,000 files, memory left after the scan 129 MB 90 MB
Parsing one Apache listing with 150,000 files 82 s under 8 s
Memory allocated per file while scanning (nginx style) about 27 KB 2.6 KB
Memory allocated per file while scanning (Apache style) not measured 3.3 KB
Small scan (21 directories), fixed overhead 209 ms 33 ms
Sorting 100,000 URLs for the URL list 3.1 s 0.46 s
Remembering 1 million scanned directories 209 MB 48 MB
An empty WebDirectory object 584 bytes 264 bytes
--resume of a database with 200,000 directories, peak memory 198 MB 119 MB

What changed

Speed

  • Fixed slow parsing of large Apache and nginx style listings, which grew faster than the listing size.
  • Listing lines are read without building a DOM for each one. For large pages (200 KB or more) with a plain listing, the lines
    are read straight from the HTML. Anything unusual falls back to the old path.
  • Regular nginx and Apache rows are read directly instead of through a regex, and a file's name no longer needs a second URL
    parse.
  • The finished URL list sorts much faster.
  • A response body is decoded once instead of being copied through several buffers.
  • Workers pick up new directories faster, so short scans finish sooner.
  • Release builds are now optimized.
  • SQLite is tuned for the scan database (cache size, temp storage, batch size) and for the viewer (memory-mapped reads).

Memory

  • Scanned directories are remembered as 128-bit hashes instead of full URLs.
  • The directory content check uses a hash instead of a copy of every file name.
  • Lighter per-directory lists and locks.
  • --resume rebuilds the tree from one compact entry per directory instead of several copies.
  • The URL list is written line by line, without extra copies.
  • About 10x fewer allocations per file during a scan.

Also in this release

  • Fixed a crash on Copyparty listings with no files or directories, and an error when resuming a scan whose database was still being written.
  • --resume with nothing left to do no longer adds a new run, repeats the speedtest or uploads the URL list again.
  • The link to an uploaded URL list is now stored in the database.

Don't miss a new OpenDirectoryDownloader release

NewReleases is sending notifications on new releases.