figwig 0.1.0 is the first release of figwig, a fast, multithreaded reader and writer of bigWig files, into and out of numpy. It is for anyone who reads the signal under many genomic windows, such as the training data of a sequence-to-function model, writes bigWigs from numpy arrays or model predictions, or converts reads to per-base count tracks with figwig bam2bw.
Pull requests
- #1 Prepare figwig 0.1.0: read_windows, missing windows, Windows and macOS reading, fixes and CI
- #2 Add BigWigWriter and write_bigwig: multithreaded bigWig writing with zoom levels
- #3 Make writing mirror reading: read_bigwig, write_bigwig and BigWigWriter.write
- #4 FIX measure each writer benchmark run's own peak memory, not its parent's
- #5 UPDATE rename BigWig to BigWigReader
- #6 ADD figwig bam2bw, bam2bw with the speed search's readers and figwig's writer
- #7 UPDATE figwig bam2bw: read each file on every core, SAM-like files in a pool, no joblib
- #9 UPDATE figwig bam2bw: drop counted keys before the keys array grows
- #10 FIX figwig bam2bw's BGZF inflate kernel using its scratch buffer after freeing it
- #11 DOC split the README into a shorter README, user-guide docs and a Claude Code skill
- #12 DOC fix the figwig skill and docs where agent tests found them wrong or short
- #13 FIX np.str_ names in messages; DOC figwig bam2bw on Windows and a comment's timings
- #14 RELEASE figwig 0.1.0
The full history is git log v0.1.0.
Reading
New: BigWigReader and read_bigwig.
BigWigReader(path).read(chroms, starts, width, out=None, n_jobs=8, missing=0.0)reads the per-base values of many windows of one width into a float32 array of shape(n, width).read_bigwiggiven a list of bigWigs reads them all into one array of shape(n, len(bigwigs), width), which is the(batch, channels, length)layout that sequence models take.- The work runs on up to
n_jobsthreads. These run in parallel because zlib'suncompress()and figwig's numba kernels release the GIL. - On 167,750 windows of 1,000 bp from an ENCODE ATAC-seq bigWig, figwig took 0.135 s on 8 threads, against pybigtools' 1.25 s, on a 2x AMD EPYC 9575F.
The values are those pybigtools' values() gives, bit for bit:
- a base that no interval covers is
missing, 0 unless given; - a base past the end of its chromosome is NaN;
- a window on a chromosome the file does not have is
missingthroughout, with a warning.
What it reads:
- bedGraph, varStep and fixedStep sections, compressed or not.
- A file it cannot read exactly raises a
ValueErrorthat says what it found, such as a bigBed, a big-endian file, a corrupt index or overlapping intervals. - A
BigWigReadercan be read from several threads at once, and pickled into DataLoader workers.
The reader began inside tangermeme's extract_loci (tangermeme #107).
Writing
New: BigWigWriter and write_bigwig.
write_bigwigwrites whatread_bigwigreads: windows, intervals or single bases, from(n, width)or(n, len(paths), width)arrays.BigWigWriterwrites the same a batch at a time, so a track larger than memory can be written in pieces.- Data blocks are compressed on several threads, with zlib, or with libdeflate through the new
fastextra. - Each window gets the smallest section type: fixedStep, varStep or bedGraph.
- Writing 15.9 million per-base counts took 0.18 s on 8 threads, against pyBigWig's 4.20 s.
Files match pyBigWig's. Files are laid out as libBigWig lays them out. A file of intervals written with missing=numpy.nan, zlib at level 6 and no zoom levels is byte for byte the file pyBigWig writes. Where libBigWig writes a wrong value, figwig writes the right one:
- the header's maximum;
- the end of the last fixedStep block of each call;
- the sums of the last zoom record of each zoom block.
figwig bam2bw
New: the figwig bam2bw command. It is bam2bw 0.5.1 with the same arguments, outputs and messages.
- A BAM file is read with libdeflate and a numba kernel on
-pthreads, and a BED/tsv file with a numba kernel. - The counts are written by
BigWigWriter. - On a 2.4 GB ATAC-seq BAM it took 4.65 s and 910 MB at
-p 2, against bam2bw's 58.48 s and 3,039 MB. - It gave the same exit code, messages and decoded entries as bam2bw in all 3,781 runs that bam2bw completed, over 591 public test files.
- It needs the new
bam2bwextra, and runs on Linux and macOS.
Claude Code skill
New: figwig install-skill. It installs a Claude Code skill that teaches a coding agent how to use figwig:
- reading windows, and data loaders;
- writing predictions, intervals and counts;
- converting reads with
figwig bam2bw; - the rules a write follows, and what each error message means;
- the measured gain from threads.
Run figwig install-skill --force after upgrading figwig, to replace the installed copy.
Platforms
- Tested on Linux with Python 3.10 to 3.14, and on macOS and Windows with Python 3.10 and 3.13.
- It depends only on numpy (1.23 or later) and numba (0.58 or later).
This is the first release. Reading the per-base signal under hundreds of thousands of windows, from one or many bigWigs, used to mean a Python call per window that more threads did not speed up. With figwig it is one call whose threads run in parallel. Writing bigWigs from model predictions, and converting reads to count tracks, are each faster than the tools they replace. There is no earlier version to be compatible with.
🤖 Generated with Claude Code