Changelog

Changelog#

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased]#

Added#

  • mantispy.io: read_profiles for profile files, CellProfiler ExportToSpreadsheet directories and CytoTable parquet parts; read_plate for a Cell Painting Gallery source or an ExportForSpatialData plate folder as SpatialData; read_jump, read, write and validate

  • mantispy.ds: the generated synthetic_plate and blobs; bbbc021, rohban, pki and jump_target2 with the annotations the analyses need; five further Cell Painting Gallery accessions

  • mantispy.ds: jump_cells, jump_export and jump_plate, the single cells, one CellProfiler export directory and the images of BR00121438, the plate jump_target2 reads well profiles for; jump_cells(selected=True) returns only the features var['selected'] marks

  • mantispy.pp: quality control at cell, image and well level, normalization, feature selection, outlier detection, sphering, plate-position correction and Harmony

  • mantispy.tl: aggregation, consensus profiles, mAP and replicate retrieval, hit calling, effect sizes, dose response, mechanism-of-action retrieval and enrichment, differential features, transport across sites and single-cell heterogeneity

  • mantispy.metrics, mantispy.get and mantispy.pl, for judging a correction, reading results out and plotting them

  • mantispy.settings, holding the verbosity and the cache directory the datasets download into

  • mantispy.io: stamp, which puts an AnnData built elsewhere — a published h5ad, another pipeline’s output, a matrix of learned embeddings — on the mantispy API surface

  • mantispy.metrics: known_relationships, the share of annotated perturbation pairs whose similarity falls in either tail of the distribution over all pairs, and evaluate_correction(covariates=...), which reports what a representation spends its variance on besides the batch and the label

  • mantispy.pp: tvn, typical variation normalization with per-batch CORAL, which aligns each batch’s controls onto the pooled controls

  • mantispy.ds: jump_lite, the same 1,536 JUMP Target-2 wells embedded by five models and measured by cp_measure, and jump_lite_targets, the gene each compound is annotated to act on

  • mantispy.metrics: known_relationships(n_permutations=...) measures chance by shuffling which perturbation each annotation row names, keeping every set’s size and every perturbation’s number of sets, and reports it as null with a p_value. Chance is 2 × percentile only when every perturbation belongs to the same number of sets

  • mantispy.pp: feature_select’s drop_degenerate, run first by default, drops the features normalize flagged in var["degenerate_scale"] so that they no longer decide which other features are kept

  • mantispy.ds: jump_crispr joins JUMP’s CRISPR annotation by default, naming the gene each well targets, its control type and the chromosome arm the gene sits on, and corum returns CORUM’s protein complexes in the shape known_relationships and pathway_coherence read

  • mantispy.pp: annotate_jump(kind="crispr"), which also reads profiles that already carry Metadata_JCP2022, as JUMP’s assembled profiles do

  • mantispy.pl: hits and feature_volcano write how many points sit above and below the significance line, next to it

  • mantispy.pp: regress_out(reference=...) fits the covariate on the reference rows, re-expresses each group at their mean and clips it to their range, so a cell count is regressed out where density varies for technical reasons only; without a reference it warns when a group never reaches the value it is re-expressed at

  • mantispy.ds: scallops_arv471, the drug arm of the SCALLOPS optical pooled CRISPR screen at single-cell resolution, treated with the estrogen-receptor degrader ARV-471 (vepdegestrant). It carries gene symbols, guides and the non-targeting controls over nine phenotype features, so hit_calling scores each guide against the non-targeting cells; CRBN, DDB1, CUL4A, CUL4B and ESR1 are the known-mechanism rescuers

  • mantispy.ds: cp_posh, the insitro cp-POSH 124-gene proof-of-concept screen at single-cell resolution, a broad-morphology Cell Painting pooled CRISPR knockout in A549 cells. It carries about 1,278 well-normalized CellStats features (so mt.pp.normalize is not needed), gene symbols, guides and the non-targeting and intergenic controls, so hit_calling scores each gene against them; KIF18A, the proteasome, the mitochondrial ribosome, ARP2/3 and COPI are the known-mechanism genes. The broad-morphology complement to scallops_arv471

  • mantispy.pp: feature_select’s corr_window and corr_stride, an opt-in fast correlation_threshold for large screens. It sorts features by name and prunes redundancy within sliding windows of corr_window features (a two-pass approximation: the windowed pre-filter whittles the list down, then the exact all-pairs pass runs on the survivors to catch cross-family redundancy). Exact stays the default; the fast path keeps a different set of features but preserves the information (reconstruction R2 near 1) and the downstream replicate signal, several times faster on screens with many thousands of features.

  • mantispy.tl: enrich now exposes decoupler’s full set of set-scoring methods (ulm, mlm, ora, aucell, gsea, gsva, zscore, waggr, viper) and a method="consensus" that combines a panel of them (CONSENSUS_PANEL, overridable via methods=) into a robust score_consensus; n_permutations replaces a single method’s padj with calibrated permutation p-values, and check_collinearity warns when the feature sets are near-collinear

Changed#

  • mantispy.tl: consensus now defaults to method="median" (was modz) — it matches pycytominer and is more robust; breaking change

  • mantispy.tl: aggregate and every other grouped reduction of a backed object take each group’s rows from one stable ordering, rather than scanning the group codes once per group. The scan was two full-length passes per group, so its cost was set by the group count: reducing 1,000,000 cells to 50,000 wells spent 21 s on index arithmetic before a single row was read, against 0.03 s now when the rows arrive grouped, as they do from a plate-ordered file, or 0.14 s when they are shuffled

  • mantispy: a backed read whose rows are already in increasing order, which is what every grouped path asks for, goes to h5py as it stands instead of being sorted and then gathered back into the order it was already in. That gather was a full-size copy of the block just read, 720 MB at JUMP well scale. Rows that form a run with no gaps, which is every group of a file stored in the grouping’s own order and every row under by=None, are read as one block rather than selected point by point: 0.022 ms against 0.168 ms for a 200-row well on an uncompressed 20,000 x 200 file, and the same 7x at 10,000 rows

  • mantispy: reduce_grouped rejects a mask that does not hold one entry per row with one message on both the backed and the in-memory path, where each previously raised a different error from inside numpy

  • mantispy: a grouped reduction of a matrix stored column-major (CSC) on disk says that it cannot be read row by row, and names the two ways out. anndata falls back to reading the whole matrix for each group, so a loop over 50,000 wells reads the screen 50,000 times with nothing said

  • mantispy.pp: variance_threshold and drop_outliers reduce one column block at a time, so feature_select no longer builds a full-matrix copy and fits large screens on a commodity node.

Fixed#

  • mantispy.tl: cluster_composition p-values are now calibrated — the overdispersion is estimated leave-one-out and the scaled statistic is referred to an F distribution, so the pure-null false positive rate no longer exceeds the nominal level (it ran to 0.14 at 8 control wells) (#89)

  • mantispy: a sparse matrix on disk can be read whole. anndata’s CSRDataset and CSCDataset have no __array__, so every read that did not name rows raised setting an array element with a sequence, pp.calculate_qc_metrics on a backed sparse object among them

  • mantispy.io: an ExportToSpreadsheet directory takes its channels from the features it measured, so a run whose images are named OrigDNA or IllumDNA no longer leaves every feature without a channel

  • mantispy.io: CellProfiler 4’s AreaShape_Center_X/Y and bounding-box corners are read as where an object sits rather than as features, and the centroid goes to Metadata_Center_X/Y; nine of the packaged datasets carried them in their profiles

  • mantispy.ds: jump_cells keeps the image quality of every field of view, each under its own image number, instead of the first field’s alone

  • mantispy.pl: feature_groups counts features that have no channel, such as AreaShape, under “none”; under pandas 3 it left them out of the bars

  • mantispy.io: cp_measure column names are read as <object>_<channel>/<aggregation>/<group><Feature> rather than through the CellProfiler grammar, which left var['channel'] empty and split one feature group into as many as the channels and aggregations it was written with

  • mantispy.pp: well_qc says it expects cell resolution, instead of counting one row per well and failing every well on a well-level object

  • mantispy.pp: feature_select warns when it selects nothing, rather than leaving an empty matrix for whatever runs next; noise_removal’s stdev_cutoff is documented as an absolute threshold on the scale normalize left the values on

  • mantispy.tl: map(mode="activity") retrieves against the controls on the query’s own plate. It pooled every plate’s controls, so a perturbation with no effect of its own looked more active the more controls the other plates carried

  • mantispy.tl: map leaves out, with a warning, a query whose replicates have no negative pair to be ranked against, such as one on a plate without controls under mode="activity". Scored, it came out at an average precision of 1 and the smallest p-value

  • mantispy.tl: map warns when null_size is too small for the multiple-testing correction to call a group on its own

  • mantispy.ds: jump_lite returns every feature set with the wells in one order, sorted by source, plate and well, where each file lists them in its own

  • mantispy.tl: map draws its permutation nulls afresh on every call instead of through copairs’ cache in the home directory. The cache keys a null without the seed it was drawn with, so a p-value depended on whichever earlier call had written that null

  • mantispy.tl: feature_signature and dose_trajectory build var from the annotation columns every non-CellProfiler object supplies, so io.validate accepts their results and io.write writes them. Both stamped an object carrying three and one of the ten columns the schema requires

  • mantispy.tl: cluster_composition takes its annotation from the same helper rather than supplying all ten by hand, so its empty channel, radial_bin and params are categoricals rather than floats

  • mantispy.io: stamp records the schema version by assigning into uns rather than through uns.setdefault. On a view, setdefault is dict’s own, so it neither materialised the view nor wrote where the caller could see, and it handed back the parent’s own store: stamping a subset left the subset unstamped and rewrote the resolution of the object it came from

  • mantispy.pp: calculate_qc_metrics refuses an object no column of which names an area, rather than asking whether the feature column holds anything. qc_pass is an and over its checks, and a parsed Intensity-only export passed the old guard and still got an all-false area flag, which made qc_pass a weaker statement than it claims

  • mantispy.tl: feature_signature’s by columns keep the dtype and the missing values var held them in, instead of the strings that name the family; grouping on scale no longer writes text into a float column, on is_feature no longer writes "True" into a boolean one, and a genuinely absent channel is missing rather than the string "none". Naming the same column twice now raises

  • mantispy.tl: dose_trajectory names a position with four decimals rather than two, and refuses a grid it cannot name. Past 101 positions two decimals gave several positions the same name, so the object carried duplicate var_names, which neither validate nor the writer objects to and which makes a per-column lookup return more than one column

  • mantispy.tl: cluster_composition counts every cell of a well in Metadata_CellCount, not only the clustered ones. It is tl.cytotoxicity’s default count_key, so a well the clustering merely left cells out of read as cell loss

  • mantispy.tl: cluster_composition leaves out a well none of whose cells the clustering assigned, rather than reporting it as zero in every cluster. Its Metadata_ClusteredCellCount records the denominator the fractions are over, which differs from Metadata_CellCount when cells were left unassigned

  • mantispy.tl: cluster_composition refuses a clustering that assigned no cell at all, instead of returning a zero-feature object io.validate rejects

  • mantispy.tl: feature_signature refuses a by whose values name two families identically, and refuses one whose columns are all empty rather than averaging every feature in the object into a single column called none | none | none

  • mantispy.tl: feature_signature marks a missing component none on every supported pandas. Below pandas 3 astype(str) wrote the string nan first, so the sentinel never applied

  • mantispy.pp: standardize_feature_names leaves a feature whose annotation names nothing under its own name. empty_annotation marks every column is_feature while leaving the descriptive ones empty, and the canonical name is built from exactly those, so an embedding renamed every feature to the empty string and then refused the collision

  • mantispy.pp: calculate_qc_metrics checks for an area feature before it reads the matrix, so a call it will refuse no longer densifies X or copies the object first

  • mantispy.tl: feature_sets returns an empty network when no feature carries every component of by, instead of failing with AttributeError: 'DataFrame' object has no attribute 'str'. It also no longer drops a feature whose family name merely contains the letters nan, such as a Nanog channel

  • mantispy.tl: cluster_composition leaves a cell the clustering did not assign out of the fractions, reported through report_drop. The label it contributed became a cluster literally called nan under pandas 2, and raised under pandas 3, where sorted cannot order a missing value against the cluster names. subpopulation_hits drops it too