pypi figwig 0.1.0
figwig 0.1.0

2 days ago

figwig 0.1.0 is the first release of figwig, a fast, multithreaded reader and writer of bigWig files, into and out of numpy. It is for anyone who reads the signal under many genomic windows, such as the training data of a sequence-to-function model, writes bigWigs from numpy arrays or model predictions, or converts reads to per-base count tracks with figwig bam2bw.

Pull requests

  • #1 Prepare figwig 0.1.0: read_windows, missing windows, Windows and macOS reading, fixes and CI
  • #2 Add BigWigWriter and write_bigwig: multithreaded bigWig writing with zoom levels
  • #3 Make writing mirror reading: read_bigwig, write_bigwig and BigWigWriter.write
  • #4 FIX measure each writer benchmark run's own peak memory, not its parent's
  • #5 UPDATE rename BigWig to BigWigReader
  • #6 ADD figwig bam2bw, bam2bw with the speed search's readers and figwig's writer
  • #7 UPDATE figwig bam2bw: read each file on every core, SAM-like files in a pool, no joblib
  • #9 UPDATE figwig bam2bw: drop counted keys before the keys array grows
  • #10 FIX figwig bam2bw's BGZF inflate kernel using its scratch buffer after freeing it
  • #11 DOC split the README into a shorter README, user-guide docs and a Claude Code skill
  • #12 DOC fix the figwig skill and docs where agent tests found them wrong or short
  • #13 FIX np.str_ names in messages; DOC figwig bam2bw on Windows and a comment's timings
  • #14 RELEASE figwig 0.1.0

The full history is git log v0.1.0.

Reading

New: BigWigReader and read_bigwig.

  • BigWigReader(path).read(chroms, starts, width, out=None, n_jobs=8, missing=0.0) reads the per-base values of many windows of one width into a float32 array of shape (n, width).
  • read_bigwig given a list of bigWigs reads them all into one array of shape (n, len(bigwigs), width), which is the (batch, channels, length) layout that sequence models take.
  • The work runs on up to n_jobs threads. These run in parallel because zlib's uncompress() and figwig's numba kernels release the GIL.
  • On 167,750 windows of 1,000 bp from an ENCODE ATAC-seq bigWig, figwig took 0.135 s on 8 threads, against pybigtools' 1.25 s, on a 2x AMD EPYC 9575F.

The values are those pybigtools' values() gives, bit for bit:

  • a base that no interval covers is missing, 0 unless given;
  • a base past the end of its chromosome is NaN;
  • a window on a chromosome the file does not have is missing throughout, with a warning.

What it reads:

  • bedGraph, varStep and fixedStep sections, compressed or not.
  • A file it cannot read exactly raises a ValueError that says what it found, such as a bigBed, a big-endian file, a corrupt index or overlapping intervals.
  • A BigWigReader can be read from several threads at once, and pickled into DataLoader workers.

The reader began inside tangermeme's extract_loci (tangermeme #107).

Writing

New: BigWigWriter and write_bigwig.

  • write_bigwig writes what read_bigwig reads: windows, intervals or single bases, from (n, width) or (n, len(paths), width) arrays.
  • BigWigWriter writes the same a batch at a time, so a track larger than memory can be written in pieces.
  • Data blocks are compressed on several threads, with zlib, or with libdeflate through the new fast extra.
  • Each window gets the smallest section type: fixedStep, varStep or bedGraph.
  • Writing 15.9 million per-base counts took 0.18 s on 8 threads, against pyBigWig's 4.20 s.

Files match pyBigWig's. Files are laid out as libBigWig lays them out. A file of intervals written with missing=numpy.nan, zlib at level 6 and no zoom levels is byte for byte the file pyBigWig writes. Where libBigWig writes a wrong value, figwig writes the right one:

  • the header's maximum;
  • the end of the last fixedStep block of each call;
  • the sums of the last zoom record of each zoom block.

figwig bam2bw

New: the figwig bam2bw command. It is bam2bw 0.5.1 with the same arguments, outputs and messages.

  • A BAM file is read with libdeflate and a numba kernel on -p threads, and a BED/tsv file with a numba kernel.
  • The counts are written by BigWigWriter.
  • On a 2.4 GB ATAC-seq BAM it took 4.65 s and 910 MB at -p 2, against bam2bw's 58.48 s and 3,039 MB.
  • It gave the same exit code, messages and decoded entries as bam2bw in all 3,781 runs that bam2bw completed, over 591 public test files.
  • It needs the new bam2bw extra, and runs on Linux and macOS.

Claude Code skill

New: figwig install-skill. It installs a Claude Code skill that teaches a coding agent how to use figwig:

  • reading windows, and data loaders;
  • writing predictions, intervals and counts;
  • converting reads with figwig bam2bw;
  • the rules a write follows, and what each error message means;
  • the measured gain from threads.

Run figwig install-skill --force after upgrading figwig, to replace the installed copy.

Platforms

  • Tested on Linux with Python 3.10 to 3.14, and on macOS and Windows with Python 3.10 and 3.13.
  • It depends only on numpy (1.23 or later) and numba (0.58 or later).

This is the first release. Reading the per-base signal under hundreds of thousands of windows, from one or many bigWigs, used to mean a Python call per window that more threads did not speed up. With figwig it is one call whose threads run in parallel. Writing bigWigs from model predictions, and converting reads to count tracks, are each faster than the tools they replace. There is no earlier version to be compatible with.

🤖 Generated with Claude Code

Don't miss a new figwig release

NewReleases is sending notifications on new releases.