pypi piaso-tools 1.2.1
PIASO v1.2.1 — streaming at scale: faster GDR, new data loaders, reproducible tuning

latest releases: 1.2.6, 1.2.5, 1.2.4...
one month ago

The one deliberate output change — read this first

Streaming (cytome-path) gene-set scores move once in this release, by
summation-order noise only (max ~1e-1 on scores of order 1, before L2
normalisation; rankings unaffected). Pass 1 of the scorer used to accumulate
its per-feature statistics at whatever block size the pass-2 scoring chunk
resolved to, which made two tuning knobs (score_chunk_size,
max_score_chunk_bytes) quietly output-affecting. Pass 1 now accumulates at
a fixed internal block, so from this release on the knobs — and any future
retuning of their defaults — are genuinely output-neutral, pinned by
tests that require bit-identical scores across chunk sizes, byte budgets,
and cache on/off. One discontinuity, taken once. The AnnData fast path is
unaffected. Every other performance change below is bit-identical.

Performance — streaming scoring and GDR

  • The Rust scoring kernel's parallel work unit was mis-sized on both
    backends, leaving most threads idle. Fixed, plus: control-gene neighbours
    are computed only for the features gene sets actually use; a batch's
    chunks are cached between the scorer's two passes instead of being
    decompressed twice; the control matrix is built in one pass; values cross
    the Rust boundary as float32 when exactly representable (accumulation
    stays float64).
  • Measured on a 200k-cell, 35-library atlas: the GDR scoring stage is
    3.3× faster, the full pipeline (INFOG → SVD → GDR → neighbours →
    UMAP) runs in 17 minutes at 5.7 GB peak RSS — see the new
    GDR at scale tutorial. AnnData multi-set scoring: 2.4× (20k cells ×
    200 sets: 0.60 s → 0.25 s).
  • GDR's stages use measured thread policies: stage 1 parallelises across
    batches, scoring splits a shared core budget; max_workers is one budget,
    not three different meanings.
  • On the cytome backend, GDR now recomputes INFOG per batch (matching
    what the AnnData path always did), so each batch's HVGs come from its own
    variance.
  • Masked streaming reads skip chunks with no selected cells, and scoring
    allocates at the masked height — batch-wise operations on large files no
    longer pay full-dataset costs.

Added

  • Ligand–receptor database loaders (piaso.data.fetch_lr_database /
    load_lr_database, human and mouse) for SCALAR/LARIS workflows.
  • ChEMBL drug-target gene sets (piaso.data.fetch_chembl,
    load_chembl_targets, filter_chembl_activities) with a KEGG/drug-target
    tutorial.
  • Motif database loaders (load_jaspar_meme, load_cisbp_meme,
    resolve_jaspar_path, …) and genome sequence access (2bit fetch/open,
    revcomp) for the motif-analysis workflow.
  • h5ad → cytome converter with counts-source verification, and
    Cell Ranger readers (read_10x_h5, read_10x_mtx) that match
    scanpy's output byte for byte without requiring scanpy.
  • Five ready-to-stream cytome datasets on Zenodo (record 22012620),
    registered for fetch_dataset/load_dataset — including a 200k-nuclei
    developing visual cortex atlas and a 1.5M-nuclei human PFC atlas
    (25.7 GB). load_dataset(name, return_type="cytome") opens the download
    directly.
  • Configurable data root: data_dir= per call,
    piaso.settings.data_dir per session, PIASO_DATA_DIR per machine;
    datasets, genomes and the registry cache share the resolved root.
  • Cytome-backend ergonomics: neighbors/umap/leiden accept a
    cytome path and a modality=; embedding names resolve through one
    helper; runSVD can normalise on the fly from stored INFOG parameters
    (no materialised layer needed); computed-layer dtypes have one knob.
  • Reproducibility provenance: runGDR records its parameters and the
    package version in adata.uns['gdr'] / cytome metadata, and the marker
    sets each batch contributed.
  • Tutorials: GDR at scale (200k cells end to end, with the memory
    knobs explained against measured curves), category order and colours on
    both backends, human PBMC end-to-end pair, cells-vs-nuclei, projectGDR,
    motif analysis, PIASOscore, KEGG/drug targets — on the relaunched
    piaso.org with executed notebooks and figures.
  • piaso[tbb] optional extra for the numba tbb threading layer (not
    required; the default layer is thread-safe out of the box).

Fixed

  • Boolean columns that read back all-True. highly_variable written by
    older versions could be stored as text; reading it selected every
    feature as highly variable. Both backends now coerce boolean columns
    safely and warn when a mask selects everything.
  • expressed_pct was silently ignored by marker detection on the
    AnnData path; it now reaches COSG on both backends.
  • INFOG refuses non-count input. Normalised/log-transformed matrices
    produced plausible-looking, meaningless results; now a clear error names
    the fixes (layer='counts', adata.raw, infog_layer=), with
    allow_non_integer=True for Smart-seq2 TPM and similar.
  • runGDR(batch_key=...) crashed on dense matrices.
  • umap() ignored neighbors_key on the AnnData path.
  • One group order for every plot: dotplot, stacked bar, Sankey, heatmap and
    the embedding legends all resolve category order (and persisted cytome
    category order) through one helper; three plotting layout defects fixed
    and figure sizing defaults tightened.
  • Threaded runs are reproducible: the graph-clustering RNG is
    serialised and control-gene sampling uses a local RNG — concurrent
    batches no longer interleave draws.
  • import piaso no longer fails on a matplotlib without cm.get_cmap.
  • runGDRParallel is a deprecated alias that forwards to runGDR
    (previously a drifted copy).
  • The website's visitor map is served live; /visitor-map/ redirects
    instead of snapshotting.
  • Docstring audit: public docstrings state current behaviour; internal
    jargon removed, measured numbers kept.

Changed

  • Pass-1↔2 chunk cache default 1 GB (a cap, not an allocation; small
    datasets are unaffected).
  • Requires cosg>=1.1.2 (earlier 1.1.x raise ImportError on plain AnnData
    when the optional cytome extra is absent) and cytome>=0.2.4.

Full suite: 887 passed, 170 skipped.

Don't miss a new piaso-tools release

NewReleases is sending notifications on new releases.