🚀 Features
- Add signature reversal scoring to enrichment (#1082) @daveringelberg
- Add native Hill dose-response fitting (#1097) @daveringelberg
- Expose datasets as
pt.dsin addition topt.dt(#1103) @Zethson - Make Milo differential abundance testing match R Milo, including the pydeseq2 solver (#1109, #1110) @Zethson
- Add
n_jobstofit_dose_responseandbalance_classestoMLPClassifierSpace.compute(#1142, #1153) @Zethson - Seed the control gene sampling of
Enrichment.scoreviarandom_state(#1128) @Zethson - Accept
hdi_probin scCODA summaries with arviz 1.x (#1159) @Zethson
⚡ Performance
Hot loops across all tools were vectorized or moved into numba kernels, with results verified against the previous implementation or the reference library on real data.
- Distances:
DistanceTest35x,ks_test27x,pairwisewitht_test/sym_kldiv9-13x, MMD without n x n kernel matrices, numba KDE formean_var_distribution, exact NB2 fits fornb_ll(#1118, #1134, #1135) @Zethson - Differential expression: Statsmodels contrasts 11-20x and batched OLS fits, sparse Wilcoxon and t-test kernels, vectorized
PermutationTeststatistic (#1124, #1136) @Zethson - Milo: batched neighbourhood refinement and pydeseq2 likelihood-ratio fits (3-7x), numba GLMM fits (~34x) (#1125, #1137) @Zethson
- Dialogue: numba random-intercept fits (
test_celltype_pairs119 s to 0.9 s at 30k cells) (#1124, #1133) @Zethson - Mixscape and Mixscale: sparse signature assembly, numba EM, one DE call per reference (
mixscale11x,mixscape6x) (#1140, #1143) @Zethson - Perturbation spaces, enrichment and metadata:
Enrichment.score88x,label_transfer19x,CellLine.correlate6x, chunked MLP training, parallel PubChem lookups (#1139, #1142) @Zethson - scCODA:
make_arvizno longer simulates a discarded N x N counts tensor (minutes to under a second) (#1138) @Zethson - Guide assignment 120x on sparse input and faster CINEMA-OT without N x N intermediates (#1140) @Zethson
⚠️ Changed results
- Mixscape holds the control variance fixed during EM as Seurat does; about 0.4% of KO/NP calls change on Papalexi 2021 (#1144) @Zethson
- Euclidean e-distance and mean pairwise distance average over all n^2 cell pairs, so
Distance.__call__,pairwiseandDistanceTestagree (#1151) @Zethson Distance.bootstrapresamples the second group with its own size (#1152) @Zethson- scGen
batch_removalapplies its latent correction, which was previously discarded (#1119) @Zethson - Dialogue p-values move slightly because the mixed-model fit converges where statsmodels' BFGS stopped early (#1133) @Zethson
nb_llfits the overdispersed genes that statsmodels silently dropped (#1135) @Zethson- Wilcoxon, t-test, e-distance and MMD statistics on float32 input are accumulated in float64 (#1134, #1136) @Zethson
CellLine.correlatecomputes every pair instead of mirroring the asymmetric matrix (#1121) @Zethsonevaluate_clusteringcomputes ASW on the cells instead of on rows of the distance matrix (#1120) @Zethson- Adjusted p-values of the simple DE tests are kept when some variables have NaN p-values (#1122) @Zethson
🐛 Bug Fixes
- Keep scCODA and tascCODA working under numpyro 0.22 (#1104) @Zethson
- Use regression-appropriate scorers for Augur regressor estimators (#1105) @Sizerta
- Dialogue: use
additional_covariates, handle degenerate genes and unused categorical levels (#1132, #1156) @Zethson - Mixscape and Mixscale:
scale=Falseon sparse layers,epsilonfor every split, control labels with spaces, duplicateobs_names(#1129, #1130, #1131, #1145, #1149, #1155) @Zethson - Distances: sparse MMD in
pairwise, bounded memory, symmetric-matrix assumption, one-cell t-test groups (#1150, #1154) @Zethson - Perturbation spaces: sparse
.X, pseudobulks keep obs columns that are NaN for some groups (#1146, #1158) @Zethson CellLine.annotate_from_prismmatches drug names case-insensitively (#1127) @ZethsonScgen.plot_reg_var_ploton sparseX, scCODA tree plots aftersave=, deterministicfrom_scanpycolumn order (#1160, #1161, #1162) @Zethson- Make the PubChem compound annotation test robust to outages (#1111) @Zethson