github teaguesterling/plinking_duck v0.5.0
v0.5.0: read_plink_vcf() — fast VCF genotype reader

latest releases: v0.9.2, v0.9.0, v0.8.2...
4 months ago

What's New

read_plink_vcf(path) — a new table function for fast biallelic genotype extraction from VCF files using plink-ng's optimized GT parsing. ~2.8x faster than htslib-based VCF reading for genotype-only queries.

SELECT CHROM, POS, ID, genotypes
FROM read_plink_vcf('cohort.vcf.gz', region := 'chr1', min_gq := 30);

Features

  • Compressed VCF (.vcf.gz) support via DuckDB VFS
  • Genotype output modes: 'array', 'list', 'columns'
  • Phased haplotype output (phased := true)
  • Per-sample quality filtering (min_gq, min_dp, max_dp)
  • Region filtering (region := 'chr:start-end')
  • Projection pushdown — skips GT parsing entirely for metadata-only queries
  • Half-call handling — 'missing', 'reference', 'haploid', or 'error'

Performance (30K variants × 10K samples, 18MB .vcf.gz)

Query DuckHTS read_bcf() PlinkingDuck read_plink_vcf() Speedup
COUNT(*) 4.47s 1.63s 2.7x
Metadata only 4.63s 1.62s 2.9x
Full scan (genotypes) 4.67s 1.63s 2.9x

Limitations

  • Biallelic variants only (multiallelic skipped with warning)
  • No INFO/QUAL/FILTER field access (use DuckHTS for that)
  • Single-threaded sequential scan
  • BCF not supported

See the full documentation for details.

Don't miss a new plinking_duck release

NewReleases is sending notifications on new releases.