What's New
read_plink_vcf(path) — a new table function for fast biallelic genotype extraction from VCF files using plink-ng's optimized GT parsing. ~2.8x faster than htslib-based VCF reading for genotype-only queries.
SELECT CHROM, POS, ID, genotypes
FROM read_plink_vcf('cohort.vcf.gz', region := 'chr1', min_gq := 30);Features
- Compressed VCF (
.vcf.gz) support via DuckDB VFS - Genotype output modes:
'array','list','columns' - Phased haplotype output (
phased := true) - Per-sample quality filtering (
min_gq,min_dp,max_dp) - Region filtering (
region := 'chr:start-end') - Projection pushdown — skips GT parsing entirely for metadata-only queries
- Half-call handling —
'missing','reference','haploid', or'error'
Performance (30K variants × 10K samples, 18MB .vcf.gz)
| Query | DuckHTS read_bcf()
| PlinkingDuck read_plink_vcf()
| Speedup |
|---|---|---|---|
COUNT(*)
| 4.47s | 1.63s | 2.7x |
| Metadata only | 4.63s | 1.62s | 2.9x |
| Full scan (genotypes) | 4.67s | 1.63s | 2.9x |
Limitations
- Biallelic variants only (multiallelic skipped with warning)
- No INFO/QUAL/FILTER field access (use DuckHTS for that)
- Single-threaded sequential scan
- BCF not supported
See the full documentation for details.