What's New
Genotype Output Modes
genotypes='struct'— STRUCT with named fields per sample (variant orient) or per variant (sample orient). Like'columns'but as a single logical column with named field access (genotypes.SAMPLE1).genotypes='counts'—STRUCT(hom_ref, het, hom_alt, missing)using PgrGetCounts for zero-decompression fast counting. Orders of magnitude faster than reading individual genotypes for large cohorts.genotypes='stats'— Extends counts with derived statistics:n,af,maf,missing_rate,carrier_count,het_rate. Cross-validates againstplink_freqandplink_missing.
Unified Variants Parameter
Flexible variant selection for read_pfile and read_pgen:
- Single index:
variants := 0 - Single rsid:
variants := 'rs123' - CPRA string:
variants := '1:10000:A:G' - CPRA struct:
variants := {chrom: '1', pos: 10000, ref: 'A', alt: 'G'} - Index range:
variants := {start: 0, stop: 100} - Identifier range:
variants := {start: 'rs1', stop: 'rs3'} - Lists of any of the above
Flexible Companion Sources
- Parquet auto-discovery:
.pvar.parquetand.psam.parquetare preferred over text formats when present (configurable viaSET plinking_use_parquet_companions) - Arbitrary sources:
pvar :=andpsam :=accept CSV files, DuckDB tables, and views read_pvar/read_psamaccept any DuckDB-readable source as positional argument
Configuration & Filtering
plinking_max_threadsconfig option: cap parallel scan threads across all functionsplink_glmp_threshold: scan-time p-value filtering for GWAS output reduction