pypi pertpy 1.4.0
1.4.0 🌈

9 hours ago

🚀 Features

⚡ Performance

Hot loops across all tools were vectorized or moved into numba kernels, with results verified against the previous implementation or the reference library on real data.

  • Distances: DistanceTest 35x, ks_test 27x, pairwise with t_test/sym_kldiv 9-13x, MMD without n x n kernel matrices, numba KDE for mean_var_distribution, exact NB2 fits for nb_ll (#1118, #1134, #1135) @Zethson
  • Differential expression: Statsmodels contrasts 11-20x and batched OLS fits, sparse Wilcoxon and t-test kernels, vectorized PermutationTest statistic (#1124, #1136) @Zethson
  • Milo: batched neighbourhood refinement and pydeseq2 likelihood-ratio fits (3-7x), numba GLMM fits (~34x) (#1125, #1137) @Zethson
  • Dialogue: numba random-intercept fits (test_celltype_pairs 119 s to 0.9 s at 30k cells) (#1124, #1133) @Zethson
  • Mixscape and Mixscale: sparse signature assembly, numba EM, one DE call per reference (mixscale 11x, mixscape 6x) (#1140, #1143) @Zethson
  • Perturbation spaces, enrichment and metadata: Enrichment.score 88x, label_transfer 19x, CellLine.correlate 6x, chunked MLP training, parallel PubChem lookups (#1139, #1142) @Zethson
  • scCODA: make_arviz no longer simulates a discarded N x N counts tensor (minutes to under a second) (#1138) @Zethson
  • Guide assignment 120x on sparse input and faster CINEMA-OT without N x N intermediates (#1140) @Zethson

⚠️ Changed results

  • Mixscape holds the control variance fixed during EM as Seurat does; about 0.4% of KO/NP calls change on Papalexi 2021 (#1144) @Zethson
  • Euclidean e-distance and mean pairwise distance average over all n^2 cell pairs, so Distance.__call__, pairwise and DistanceTest agree (#1151) @Zethson
  • Distance.bootstrap resamples the second group with its own size (#1152) @Zethson
  • scGen batch_removal applies its latent correction, which was previously discarded (#1119) @Zethson
  • Dialogue p-values move slightly because the mixed-model fit converges where statsmodels' BFGS stopped early (#1133) @Zethson
  • nb_ll fits the overdispersed genes that statsmodels silently dropped (#1135) @Zethson
  • Wilcoxon, t-test, e-distance and MMD statistics on float32 input are accumulated in float64 (#1134, #1136) @Zethson
  • CellLine.correlate computes every pair instead of mirroring the asymmetric matrix (#1121) @Zethson
  • evaluate_clustering computes ASW on the cells instead of on rows of the distance matrix (#1120) @Zethson
  • Adjusted p-values of the simple DE tests are kept when some variables have NaN p-values (#1122) @Zethson

🐛 Bug Fixes

  • Keep scCODA and tascCODA working under numpyro 0.22 (#1104) @Zethson
  • Use regression-appropriate scorers for Augur regressor estimators (#1105) @Sizerta
  • Dialogue: use additional_covariates, handle degenerate genes and unused categorical levels (#1132, #1156) @Zethson
  • Mixscape and Mixscale: scale=False on sparse layers, epsilon for every split, control labels with spaces, duplicate obs_names (#1129, #1130, #1131, #1145, #1149, #1155) @Zethson
  • Distances: sparse MMD in pairwise, bounded memory, symmetric-matrix assumption, one-cell t-test groups (#1150, #1154) @Zethson
  • Perturbation spaces: sparse .X, pseudobulks keep obs columns that are NaN for some groups (#1146, #1158) @Zethson
  • CellLine.annotate_from_prism matches drug names case-insensitively (#1127) @Zethson
  • Scgen.plot_reg_var_plot on sparse X, scCODA tree plots after save=, deterministic from_scanpy column order (#1160, #1161, #1162) @Zethson
  • Make the PubChem compound annotation test robust to outages (#1111) @Zethson

🧰 Maintenance

Don't miss a new pertpy release

NewReleases is sending notifications on new releases.