The one deliberate output change — read this first
Streaming (cytome-path) gene-set scores move once in this release, by
summation-order noise only (max ~1e-1 on scores of order 1, before L2
normalisation; rankings unaffected). Pass 1 of the scorer used to accumulate
its per-feature statistics at whatever block size the pass-2 scoring chunk
resolved to, which made two tuning knobs (score_chunk_size,
max_score_chunk_bytes) quietly output-affecting. Pass 1 now accumulates at
a fixed internal block, so from this release on the knobs — and any future
retuning of their defaults — are genuinely output-neutral, pinned by
tests that require bit-identical scores across chunk sizes, byte budgets,
and cache on/off. One discontinuity, taken once. The AnnData fast path is
unaffected. Every other performance change below is bit-identical.
Performance — streaming scoring and GDR
- The Rust scoring kernel's parallel work unit was mis-sized on both
backends, leaving most threads idle. Fixed, plus: control-gene neighbours
are computed only for the features gene sets actually use; a batch's
chunks are cached between the scorer's two passes instead of being
decompressed twice; the control matrix is built in one pass; values cross
the Rust boundary as float32 when exactly representable (accumulation
stays float64). - Measured on a 200k-cell, 35-library atlas: the GDR scoring stage is
3.3× faster, the full pipeline (INFOG → SVD → GDR → neighbours →
UMAP) runs in 17 minutes at 5.7 GB peak RSS — see the new
GDR at scale tutorial. AnnData multi-set scoring: 2.4× (20k cells ×
200 sets: 0.60 s → 0.25 s). - GDR's stages use measured thread policies: stage 1 parallelises across
batches, scoring splits a shared core budget;max_workersis one budget,
not three different meanings. - On the cytome backend, GDR now recomputes INFOG per batch (matching
what the AnnData path always did), so each batch's HVGs come from its own
variance. - Masked streaming reads skip chunks with no selected cells, and scoring
allocates at the masked height — batch-wise operations on large files no
longer pay full-dataset costs.
Added
- Ligand–receptor database loaders (
piaso.data.fetch_lr_database/
load_lr_database, human and mouse) for SCALAR/LARIS workflows. - ChEMBL drug-target gene sets (
piaso.data.fetch_chembl,
load_chembl_targets,filter_chembl_activities) with a KEGG/drug-target
tutorial. - Motif database loaders (
load_jaspar_meme,load_cisbp_meme,
resolve_jaspar_path, …) and genome sequence access (2bit fetch/open,
revcomp) for the motif-analysis workflow. - h5ad → cytome converter with counts-source verification, and
Cell Ranger readers (read_10x_h5,read_10x_mtx) that match
scanpy's output byte for byte without requiring scanpy. - Five ready-to-stream cytome datasets on Zenodo (record 22012620),
registered forfetch_dataset/load_dataset— including a 200k-nuclei
developing visual cortex atlas and a 1.5M-nuclei human PFC atlas
(25.7 GB).load_dataset(name, return_type="cytome")opens the download
directly. - Configurable data root:
data_dir=per call,
piaso.settings.data_dirper session,PIASO_DATA_DIRper machine;
datasets, genomes and the registry cache share the resolved root. - Cytome-backend ergonomics:
neighbors/umap/leidenaccept a
cytome path and amodality=; embedding names resolve through one
helper;runSVDcan normalise on the fly from stored INFOG parameters
(no materialised layer needed); computed-layer dtypes have one knob. - Reproducibility provenance:
runGDRrecords its parameters and the
package version inadata.uns['gdr']/ cytome metadata, and the marker
sets each batch contributed. - Tutorials: GDR at scale (200k cells end to end, with the memory
knobs explained against measured curves), category order and colours on
both backends, human PBMC end-to-end pair, cells-vs-nuclei, projectGDR,
motif analysis, PIASOscore, KEGG/drug targets — on the relaunched
piaso.org with executed notebooks and figures. piaso[tbb]optional extra for the numba tbb threading layer (not
required; the default layer is thread-safe out of the box).
Fixed
- Boolean columns that read back all-True.
highly_variablewritten by
older versions could be stored as text; reading it selected every
feature as highly variable. Both backends now coerce boolean columns
safely and warn when a mask selects everything. expressed_pctwas silently ignored by marker detection on the
AnnData path; it now reaches COSG on both backends.- INFOG refuses non-count input. Normalised/log-transformed matrices
produced plausible-looking, meaningless results; now a clear error names
the fixes (layer='counts',adata.raw,infog_layer=), with
allow_non_integer=Truefor Smart-seq2 TPM and similar. runGDR(batch_key=...)crashed on dense matrices.umap()ignoredneighbors_keyon the AnnData path.- One group order for every plot: dotplot, stacked bar, Sankey, heatmap and
the embedding legends all resolve category order (and persisted cytome
category order) through one helper; three plotting layout defects fixed
and figure sizing defaults tightened. - Threaded runs are reproducible: the graph-clustering RNG is
serialised and control-gene sampling uses a local RNG — concurrent
batches no longer interleave draws. import piasono longer fails on a matplotlib withoutcm.get_cmap.runGDRParallelis a deprecated alias that forwards torunGDR
(previously a drifted copy).- The website's visitor map is served live;
/visitor-map/redirects
instead of snapshotting. - Docstring audit: public docstrings state current behaviour; internal
jargon removed, measured numbers kept.
Changed
- Pass-1↔2 chunk cache default 1 GB (a cap, not an allocation; small
datasets are unaffected). - Requires
cosg>=1.1.2(earlier 1.1.x raise ImportError on plain AnnData
when the optional cytome extra is absent) andcytome>=0.2.4.
Full suite: 887 passed, 170 skipped.