pypi piaso-tools 1.2.4
PIASO v1.2.4

latest releases: 1.2.6, 1.2.5
15 days ago

piaso-tools 1.2.4

pip install -U piaso-tools

SCALAR: a detection floor, applied in the test

piaso.tl.runSCALAR gains expressed_pct (default 0.1) and groupby. An interaction is tested only if the ligand is detected in at least that fraction of the sender's cells and the receptor in the receiver's, and each gene's matched control genes are drawn from genes that clear the same floor in the same cell type. The null therefore never contains a control that is not expressed where it is used, and the k × k support stays full. Rows that fail the floor are returned with testable = False, p_value = 1 and are excluded from FDR; two new columns, ligand_expressed_pct and receptor_expressed_pct, carry the fractions the floor was applied to. groupby names the cell-type column and is required when the floor is on; expressed_pct=None reproduces the previous behaviour exactly.

Why in the test and not in the matrix: zeroing lowly-detected genes in the specificity matrix leaves the observed pair conditioned on passing the floor while its controls are zeroed wherever they fall below it, so most of the null becomes zeros and the p-value drifts toward "detected in more cells than its peers". Measured on the SCALAR tutorial data, a quarter of the calls that approach produces are significant only because of the zero padding.

runSCALAR now validates its input. NaN specificity scores raise (they are what cosg.indexByGene leaves for genes outside a top-N list unless set_nan_to_zero=True, and a NaN observation would otherwise score at the p-value floor). Negative scores raise (COSG marks genes failing its own expressed_pct filter with −1, and two such sentinels multiply to +1, the largest product in the table). Existing callers that pass a clean matrix and no groupby will now see an error asking for one; add groupby= or expressed_pct=None.

The SCALAR tutorial's recipe changes with it: COSG at mu=10 over all genes, raw scores, no COSG-side filter, and the floor applied by runSCALAR. Any per-cell-type linear rescaling of the specificity matrix (for instance dividing by an IQR) leaves every p-value unchanged, because the observed product and its null scale together; cosg.iqrLogNormalize applies a log1p on top and is not recommended as SCALAR input.

SCALAR: a readable score, and a magnitude that compares across pairs

The interaction score is a product of two specificities, each below 1, so a pair of two 0.02 specificities scored 0.0004 — a number no one can hold, and one that is not comparable across sender–receiver pairs because specificity is calibrated within a cell type. Four columns fix both without touching a single p-value:

  • ligand_specificity and receptor_specificity, the two inputs.
  • specificity_geomean, their geometric mean: a monotone function of the product, so it ranks exactly as the score does, on the specificity scale. The tutorial's plots are sized by it.
  • null_mean and log2_enrichment: the mean of the interaction's own matched null and the log2 ratio of the observation to it (a 10⁻¹² guard against division by zero, nothing more), so every pair is referenced to what expressed, expression-matched gene pairs achieve between the same two cell types. This is the number to compare across pairs. On the tutorial data the largest raw score in the table (PTPRC→CD22) is ten times its null; a somatostatin pair scoring 0.0008 is six thousand times its null.

statistic="min" is a second test statistic: the smaller of the two specificities, so both sides must beat the controls' weaker sides — the "are both specific" question. Its exact null factorises into two one-sided counts. The default remains the product.

piaso.tl.umap and piaso.tl.leiden no longer return a value

Breaking. Both wrote their result and returned it on AnnData, while returning None on a cytome — one function with two contracts, and a 161,027-row array echoed into every notebook cell that called it. They now return None on both backends; read the result from adata.obsm[key_added] or adata.obs[key_added], as with every other piaso.tl function.

The two calls that have nowhere to write still return their result: data=None with a knn_result (in-memory), and leiden(..., cell_mask=...), whose labels cover the masked cells only and would mis-align if written full-length.

If you have emb = piaso.tl.umap(adata) in a script, it now binds None.

TF lists: one source that works, and a file that is checked

piaso.data.fetch_tf_list(species, source=...) replaces fetch_animaltfdb_tf_list as the entry point, defaulting to the cisTarget/SCENIC+ TF universe (~1,890 human, ~1,860 mouse).

AnimalTFDB was removed as a source. Its only host (guolab.wchscu.cn) sits behind a WAF that answers every non-browser client — urllib, requests and curl alike, under any User-Agent — with a captcha page or 405, and the old bioinfo.life.hust.edu.cn mirror no longer resolves. source="animaltfdb" now raises and says so; fetch_animaltfdb_tf_list, which shipped in 1.2.3, raises without making a request and will be deleted in 1.3. An AnimalTFDB table saved from a browser still works: pass it to load_tf_list(path=...), whose TSV parser reads its Symbol column.

Any TF-list file is now validated. The captcha page arrived as HTTP 200 and used to parse into a dozen plausible-looking "symbols" that silently shrank the TF universe. A file that starts with markup, or that yields tokens no gene symbol could be, is refused; a downloaded catalogue must also yield at least 50 symbols, and one that does not is deleted rather than cached. No minimum applies to a file you pass yourself — a ten-TF list is a legitimate thing to hand in.

Fixed: piaso.tl.umap(data=None, knn_result=...) crashed

The in-memory path — documented, and advertised in this release as one of the two calls that still return their result — asked None for .obsm and died with AttributeError. It had been invisible because the test covering it skips wherever the installed umap-learn and scikit-learn disagree, which is most machines. Calling it with no knn_result now explains itself instead of raising from three frames down.

Fixed: importCellRanger reported a broken install

It imported the ATAC fragment leg unconditionally, so in the released package every call failed with ModuleNotFoundError, including modality='rna', which needs nothing from it. RNA-only now works; asking for fragments says which method is unavailable and why.

Fixed: embedding dots were drawn as rings, at 3x the size requested

plotEmbedding, plot_embeddings_split and scatter never passed linewidths to ax.scatter, so matplotlib stroked every marker at rcParams['patch.linewidth'] — 1.0 pt, in edgecolors='face'.

On a cell-sized point that stroke is most of the dot. point_size=0.2 is a disc 0.45 pt across, so it drew 1.45 pt wide: 3.2x the size asked for. With alpha < 1 the stroke and the fill composite twice around the rim, so each point rendered as an annulus with a washed-out core whose colour was neither the palette colour nor the background. Measured on a 600 dpi render, ink peaked two pixels off-centre.

All three scatter calls now pass linewidths=0, edgecolors='none', and both are exposed as parameters for anyone who wants a deliberate outline. The same fix is applied in piaso.pl.scatter and the plotByCluster jitter points. No visible change on large markers; the workaround plt.rc_context({"patch.linewidth": 0}) is no longer needed.

Fixed: plot_embeddings_split silently ignored extra keywords

It accepted **kwargs and never forwarded them, so styling a split panel had no effect and no error. Unrecognised keywords now raise TypeError, as plotEmbedding does. show is accepted as an alias for show_figure — it was among the dropped keywords, which meant show=False displayed the figure anyway.

Codebook motifs: about twice the TF universe, opt-in

piaso.data.fetch_codebook() and load_codebook() add the representative PWM set of Jolma, Laverty, Fathi et al., Nature 657, 275-283 (2026) — 1,421 human TFs against JASPAR 2024 CORE vertebrates' 754, covering all but eleven of them. On SEA-AD middle temporal gyrus that lifts the count of TFs with a motif and expression from 483 to 1,023. The archive is ~1 MB, downloaded from the publisher under the article's CC-BY-4.0 licence on first use and cached; nothing is redistributed with PIASO.

cytorete.tl.inferRegulon(..., motif_db="codebook") uses it. JASPAR remains the default, and that is now measured rather than assumed. Both were run end-to-end on the same 20,000 SEA-AD nuclei:

JASPAR 2024 Codebook
TFs with motifs in the data 740 1,408
regulons 415 982 (412 of JASPAR's 415 among them)
zinc-finger regulons 65 (16%) 385 (39%)
agreement with the TF's own specificity 26.7% 12.7%
runtime 13 min 56 min

The result is mixed rather than bad. Codebook finds more agreeing regulons in absolute terms (125 against 111) and fixes cases JASPAR gets wrong — OLIG2 now tops OPC rather than oligodendrocytes, and SOX10, absent from JASPAR, tops oligodendrocytes. But it dilutes: agreement halves, the zinc-finger share more than doubles, and neuronal factors such as RORB and LHX6 fall in rank. Reach for it when the question is glial, zinc-finger, or coverage-limited; keep JASPAR when it is a survey.

The loader repairs what the archive carries: construct names (CASZ1.FL, ZNF729.DBD3) are filed under the gene symbol, three clone identifiers are dropped, and five CIS-BP conversions with all-zero positions are cleaned and renormalised.

Sankey: node heights, and one size key instead of two

piaso.pl.sankey gains min_node_size (height of the shortest category block, in points) and node_scale ('linear', 'sqrt', 'log'), the node-level parallel of min_flow_width and flow_scale. Both are solved the same affine way, so blocks stay monotone in group size and nothing is tied.

This fixes a distortion the default layout has always had: a node is exactly as tall as the ribbons landing on it, so under min_flow_width its height is m * size + b * fan_out, and two equally large categories are drawn at different heights when one splits more ways than the other. Measured on a two-category example at a 20 pt baseline, the four-way split draws twice as tall as the equal one-way split.

What it costs is stated in the docstring and enforced in the key: each node's ribbons are renormalised to fill its height, so the count-to-thickness map becomes per-node and min_flow_width stops being exact.

show_width_legend / width_legend_loc are now show_size_legend / size_legend_loc, and draw one key rather than two. Only one map can be global at a time, so only one key can be true: the key describes flows when the flow map is non-proportional, and switches to describing nodes — suppressing the flow key — as soon as the nodes are sized.

Confusion matrices: hierarchical ordering

order_query / order_reference accept 'hclust' (average linkage on that axis's own profiles) and 'hclust_symmetric' (one clustered order on both axes, for a square table over one label set). 'auto' now prefers the symmetric order when the two axes carry the same labels — two independent orders moved a real diagonal off the diagonal — and is unchanged otherwise.

Fixed

  • The GDR SettingWithCopyWarning. Each batch was scored by slicing an anndata view, whose constructor rewrites categorical obs columns through _remove_unused_categories — an anndata-internal pandas idiom, and pure overhead here since the scorer reads only the matrix and var_names. The batch object is now built directly. This also removes two warnings.catch_warnings blocks, one of them inside a ThreadPoolExecutor worker, where mutating the process-global filter list races between threads.
  • plotLigandReceptorLollipop tick labels were formatted to one decimal, so an axis spanning a few 10⁻⁴ — what neuronal ligand–receptor scores look like — read 0.0 at every tick. The precision now follows the axis range.

Don't miss a new piaso-tools release

NewReleases is sending notifications on new releases.