pypi piaso-tools 1.2.0
PIASO v1.2.0 — Cytome datasets, atlas-scale streaming, and a self-contained workflow

latest releases: 1.2.6, 1.2.5, 1.2.4...
one month ago
pip install -U piaso-tools

The distribution is piaso-tools; the import is import piaso.

Cytome datasets and streaming

PIASO now works directly on cytome
files. Anywhere a function took an AnnData, it also takes a path to a
.cytome file or an open dataset:

import piaso, cytome

ds = cytome.open("atlas.cytome")
piaso.tl.runGDR(ds, groupby="cell_type", layer="infog")

The matrix is read in chunks rather than loaded, so peak memory is set by the
batch size instead of by the number of cells, and normalized layers (log1p,
infog, tfidf) are computed on the fly when they are not stored on the file.
Results are written back onto the dataset. cytome is installed automatically —
there is no extra to remember.

The scoring kernels are Rust-accelerated, with an automatic pure-Python
fallback that warns once and produces identical results, so a machine without a
compiled extension still works.

A self-contained analysis workflow

From raw UMI counts to clusters, an embedding and marker genes, without
leaving PIASO:

import piaso, cosg

# 1. INFOG normalization + feature selection.
#    Input must be RAW UMI counts. By default this reads adata.X; if your
#    raw counts live in a layer, pass it explicitly:
#        piaso.tl.infog(adata, layer="counts", n_top_genes=3000)
#    Writes adata.layers['infog'] and adata.var['highly_variable'].
piaso.tl.infog(adata, n_top_genes=3000)

# 2. Dimensionality reduction. Pass layer="infog" — runSVD defaults to
#    adata.X, which would silently ignore the normalization above.
piaso.tl.runSVD(adata, layer="infog", n_components=50, key_added="X_svd")

# 3. Neighbour graph, clustering, embedding
piaso.tl.neighbors(adata, use_rep="X_svd", n_neighbors=15)
piaso.tl.leiden(adata, resolution=1.0, key_added="leiden")
piaso.tl.umap(adata, use_rep="X_svd")

# 4. Marker genes
cosg.cosg(adata, groupby="leiden", key_added="cosg")

# 5. Plot
piaso.pl.embedding(adata, basis="X_umap", color="leiden")
piaso.pl.dotplot(adata, markers, groupby="leiden")

Every step above runs on a plain pip install piaso-tools — no scanpy
required
. piaso.tl.runGDR offers marker-gene-guided dimensionality
reduction as an alternative to step 2.

scanpy remains an optional extra (pip install 'piaso-tools[scanpy]') for
interoperability with scanpy-based workflows — reading 10x .h5 files,
scanpy's highly_variable_genes as an alternative feature-selection method,
and its score_genes.

Plotting

piaso.pl.embedding(adata, color="cell_type")
piaso.pl.dotplot(adata, markers, groupby="cell_type")

Embeddings and UMAPs, dot plots, heatmaps, violins, scatter plots, stacked bar
plots, Sankey diagrams, dendrograms, group-metric panels, and
split-by-category embeddings. Short aliases (embedding, umap, dotplot,
violin, scatter, heatmap, sankey) sit alongside the long names.
piaso.settings centralises figure defaults, including a set_figure_params()
one-liner and a publication preset.

Preprocessing and QC

calculateCellMetrics, calculateFeatureMetrics, calculateGroupMetrics,
filter_cells, filter_features, subset_cells, normalize_log1p,
scrublet for doublet detection, read_10x and importCellRanger.

Reference data

piaso.data is new: it fetches and caches what analyses need — genome sequence
and annotation (fetch_genome, resolve_genome_files), example datasets
(fetch_dataset, load_dataset, list_datasets), motif databases (JASPAR,
CIS-BP, cisTarget) with a PWM class, and the SCREEN cCRE registry.

Motif scanning

hits = piaso.pp.scan_motifs(sequences, pwms, threshold=...)

scan_motifs is a pure-numpy PWM scanner; scan_motifs_rust is a
rayon-parallel implementation with the same contract. pvalue_to_threshold and
estimate_background calibrate a score cutoff against a background model.

Changed

Requires cosg>=1.1.0, which provides the marker-detection entry point used by
runGDR on cytome datasets.

Security

pyo3 upgraded to 0.29.2, closing GHSA-36hh-v3qg-5jq4 (out-of-bounds read in
nth/nth_back for list and tuple iterators) and GHSA-chgr-c6px-7xpp (missing
Sync bound on PyCFunction::new_closure). PIASO called neither API, so
exposure was nil, but the crate linked the affected code.

BSD-3-Clause.

Don't miss a new piaso-tools release

NewReleases is sending notifications on new releases.