Changelog#
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]#
Added#
mantispy.io:read_profilesfor profile files, CellProfilerExportToSpreadsheetdirectories and CytoTable parquet parts;read_platefor a Cell Painting Gallery source or anExportForSpatialDataplate folder asSpatialData;read_jump,read,writeandvalidatemantispy.ds: the generatedsynthetic_plateandblobs;bbbc021,rohban,pkiandjump_target2with the annotations the analyses need; five further Cell Painting Gallery accessionsmantispy.ds:jump_cells,jump_exportandjump_plate, the single cells, one CellProfiler export directory and the images ofBR00121438, the platejump_target2reads well profiles for;jump_cells(selected=True)returns only the featuresvar['selected']marksmantispy.pp: quality control at cell, image and well level, normalization, feature selection, outlier detection, sphering, plate-position correction and Harmonymantispy.tl: aggregation, consensus profiles, mAP and replicate retrieval, hit calling, effect sizes, dose response, mechanism-of-action retrieval and enrichment, differential features, transport across sites and single-cell heterogeneitymantispy.metrics,mantispy.getandmantispy.pl, for judging a correction, reading results out and plotting themmantispy.settings, holding the verbosity and the cache directory the datasets download intomantispy.io:stamp, which puts anAnnDatabuilt elsewhere — a published h5ad, another pipeline’s output, a matrix of learned embeddings — on the mantispy API surfacemantispy.metrics:known_relationships, the share of annotated perturbation pairs whose similarity falls in either tail of the distribution over all pairs, andevaluate_correction(covariates=...), which reports what a representation spends its variance on besides the batch and the labelmantispy.pp:tvn, typical variation normalization with per-batch CORAL, which aligns each batch’s controls onto the pooled controlsmantispy.ds:jump_lite, the same 1,536 JUMP Target-2 wells embedded by five models and measured bycp_measure, andjump_lite_targets, the gene each compound is annotated to act onmantispy.metrics:known_relationships(n_permutations=...)measures chance by shuffling which perturbation each annotation row names, keeping every set’s size and every perturbation’s number of sets, and reports it asnullwith ap_value. Chance is 2 ×percentileonly when every perturbation belongs to the same number of setsmantispy.pp:feature_select’sdrop_degenerate, run first by default, drops the featuresnormalizeflagged invar["degenerate_scale"]so that they no longer decide which other features are keptmantispy.ds:jump_crisprjoins JUMP’s CRISPR annotation by default, naming the gene each well targets, its control type and the chromosome arm the gene sits on, andcorumreturns CORUM’s protein complexes in the shapeknown_relationshipsandpathway_coherencereadmantispy.pp:annotate_jump(kind="crispr"), which also reads profiles that already carryMetadata_JCP2022, as JUMP’s assembled profiles domantispy.pl:hitsandfeature_volcanowrite how many points sit above and below the significance line, next to itmantispy.pp:regress_out(reference=...)fits the covariate on the reference rows, re-expresses each group at their mean and clips it to their range, so a cell count is regressed out where density varies for technical reasons only; without a reference it warns when a group never reaches the value it is re-expressed atmantispy.ds:scallops_arv471, the drug arm of the SCALLOPS optical pooled CRISPR screen at single-cell resolution, treated with the estrogen-receptor degrader ARV-471 (vepdegestrant). It carries gene symbols, guides and the non-targeting controls over nine phenotype features, sohit_callingscores each guide against the non-targeting cells;CRBN,DDB1,CUL4A,CUL4BandESR1are the known-mechanism rescuersmantispy.ds:cp_posh, the insitro cp-POSH 124-gene proof-of-concept screen at single-cell resolution, a broad-morphology Cell Painting pooled CRISPR knockout in A549 cells. It carries about 1,278 well-normalized CellStats features (somt.pp.normalizeis not needed), gene symbols, guides and the non-targeting and intergenic controls, sohit_callingscores each gene against them;KIF18A, the proteasome, the mitochondrial ribosome, ARP2/3 and COPI are the known-mechanism genes. The broad-morphology complement toscallops_arv471mantispy.pp:feature_select’scorr_windowandcorr_stride, an opt-in fastcorrelation_thresholdfor large screens. It sorts features by name and prunes redundancy within sliding windows ofcorr_windowfeatures (a two-pass approximation: the windowed pre-filter whittles the list down, then the exact all-pairs pass runs on the survivors to catch cross-family redundancy). Exact stays the default; the fast path keeps a different set of features but preserves the information (reconstruction R2 near 1) and the downstream replicate signal, several times faster on screens with many thousands of features.mantispy.tl:enrichnow exposes decoupler’s full set of set-scoring methods (ulm,mlm,ora,aucell,gsea,gsva,zscore,waggr,viper) and amethod="consensus"that combines a panel of them (CONSENSUS_PANEL, overridable viamethods=) into a robustscore_consensus;n_permutationsreplaces a single method’spadjwith calibrated permutation p-values, andcheck_collinearitywarns when the feature sets are near-collinear
Changed#
mantispy.tl:consensusnow defaults tomethod="median"(wasmodz) — it matches pycytominer and is more robust; breaking changemantispy.tl:aggregateand every other grouped reduction of a backed object take each group’s rows from one stable ordering, rather than scanning the group codes once per group. The scan was two full-length passes per group, so its cost was set by the group count: reducing 1,000,000 cells to 50,000 wells spent 21 s on index arithmetic before a single row was read, against 0.03 s now when the rows arrive grouped, as they do from a plate-ordered file, or 0.14 s when they are shuffledmantispy: a backed read whose rows are already in increasing order, which is what every grouped path asks for, goes to h5py as it stands instead of being sorted and then gathered back into the order it was already in. That gather was a full-size copy of the block just read, 720 MB at JUMP well scale. Rows that form a run with no gaps, which is every group of a file stored in the grouping’s own order and every row underby=None, are read as one block rather than selected point by point: 0.022 ms against 0.168 ms for a 200-row well on an uncompressed 20,000 x 200 file, and the same 7x at 10,000 rowsmantispy:reduce_groupedrejects amaskthat does not hold one entry per row with one message on both the backed and the in-memory path, where each previously raised a different error from inside numpymantispy: a grouped reduction of a matrix stored column-major (CSC) on disk says that it cannot be read row by row, and names the two ways out. anndata falls back to reading the whole matrix for each group, so a loop over 50,000 wells reads the screen 50,000 times with nothing saidmantispy.pp:variance_thresholdanddrop_outliersreduce one column block at a time, sofeature_selectno longer builds a full-matrix copy and fits large screens on a commodity node.
Fixed#
mantispy.tl:cluster_compositionp-values are now calibrated — the overdispersion is estimated leave-one-out and the scaled statistic is referred to an F distribution, so the pure-null false positive rate no longer exceeds the nominal level (it ran to 0.14 at 8 control wells) (#89)mantispy: a sparse matrix on disk can be read whole. anndata’sCSRDatasetandCSCDatasethave no__array__, so every read that did not name rows raisedsetting an array element with a sequence,pp.calculate_qc_metricson a backed sparse object among themmantispy.io: anExportToSpreadsheetdirectory takes its channels from the features it measured, so a run whose images are namedOrigDNAorIllumDNAno longer leaves every feature without a channelmantispy.io: CellProfiler 4’sAreaShape_Center_X/Yand bounding-box corners are read as where an object sits rather than as features, and the centroid goes toMetadata_Center_X/Y; nine of the packaged datasets carried them in their profilesmantispy.ds:jump_cellskeeps the image quality of every field of view, each under its own image number, instead of the first field’s alonemantispy.pl:feature_groupscounts features that have no channel, such asAreaShape, under “none”; under pandas 3 it left them out of the barsmantispy.io:cp_measurecolumn names are read as<object>_<channel>/<aggregation>/<group><Feature>rather than through the CellProfiler grammar, which leftvar['channel']empty and split one feature group into as many as the channels and aggregations it was written withmantispy.pp:well_qcsays it expects cell resolution, instead of counting one row per well and failing every well on a well-level objectmantispy.pp:feature_selectwarns when it selects nothing, rather than leaving an empty matrix for whatever runs next;noise_removal’sstdev_cutoffis documented as an absolute threshold on the scalenormalizeleft the values onmantispy.tl:map(mode="activity")retrieves against the controls on the query’s own plate. It pooled every plate’s controls, so a perturbation with no effect of its own looked more active the more controls the other plates carriedmantispy.tl:mapleaves out, with a warning, a query whose replicates have no negative pair to be ranked against, such as one on a plate without controls undermode="activity". Scored, it came out at an average precision of 1 and the smallest p-valuemantispy.tl:mapwarns whennull_sizeis too small for the multiple-testing correction to call a group on its ownmantispy.ds:jump_litereturns every feature set with the wells in one order, sorted by source, plate and well, where each file lists them in its ownmantispy.tl:mapdraws its permutation nulls afresh on every call instead of through copairs’ cache in the home directory. The cache keys a null without the seed it was drawn with, so a p-value depended on whichever earlier call had written that nullmantispy.tl:feature_signatureanddose_trajectorybuildvarfrom the annotation columns every non-CellProfiler object supplies, soio.validateaccepts their results andio.writewrites them. Both stamped an object carrying three and one of the ten columns the schema requiresmantispy.tl:cluster_compositiontakes its annotation from the same helper rather than supplying all ten by hand, so its emptychannel,radial_binandparamsare categoricals rather than floatsmantispy.io:stamprecords the schema version by assigning intounsrather than throughuns.setdefault. On a view,setdefaultisdict’s own, so it neither materialised the view nor wrote where the caller could see, and it handed back the parent’s own store: stamping a subset left the subset unstamped and rewrote the resolution of the object it came frommantispy.pp:calculate_qc_metricsrefuses an object no column of which names an area, rather than asking whether thefeaturecolumn holds anything.qc_passis anandover its checks, and a parsed Intensity-only export passed the old guard and still got an all-false area flag, which madeqc_passa weaker statement than it claimsmantispy.tl:feature_signature’sbycolumns keep the dtype and the missing valuesvarheld them in, instead of the strings that name the family; grouping onscaleno longer writes text into a float column, onis_featureno longer writes"True"into a boolean one, and a genuinely absent channel is missing rather than the string"none". Naming the same column twice now raisesmantispy.tl:dose_trajectorynames a position with four decimals rather than two, and refuses a grid it cannot name. Past 101 positions two decimals gave several positions the same name, so the object carried duplicatevar_names, which neithervalidatenor the writer objects to and which makes a per-column lookup return more than one columnmantispy.tl:cluster_compositioncounts every cell of a well inMetadata_CellCount, not only the clustered ones. It istl.cytotoxicity’s defaultcount_key, so a well the clustering merely left cells out of read as cell lossmantispy.tl:cluster_compositionleaves out a well none of whose cells the clustering assigned, rather than reporting it as zero in every cluster. ItsMetadata_ClusteredCellCountrecords the denominator the fractions are over, which differs fromMetadata_CellCountwhen cells were left unassignedmantispy.tl:cluster_compositionrefuses a clustering that assigned no cell at all, instead of returning a zero-feature objectio.validaterejectsmantispy.tl:feature_signaturerefuses abywhose values name two families identically, and refuses one whose columns are all empty rather than averaging every feature in the object into a single column callednone | none | nonemantispy.tl:feature_signaturemarks a missing componentnoneon every supported pandas. Below pandas 3astype(str)wrote the stringnanfirst, so the sentinel never appliedmantispy.pp:standardize_feature_namesleaves a feature whose annotation names nothing under its own name.empty_annotationmarks every columnis_featurewhile leaving the descriptive ones empty, and the canonical name is built from exactly those, so an embedding renamed every feature to the empty string and then refused the collisionmantispy.pp:calculate_qc_metricschecks for an area feature before it reads the matrix, so a call it will refuse no longer densifiesXor copies the object firstmantispy.tl:feature_setsreturns an empty network when no feature carries every component ofby, instead of failing withAttributeError: 'DataFrame' object has no attribute 'str'. It also no longer drops a feature whose family name merely contains the lettersnan, such as aNanogchannelmantispy.tl:cluster_compositionleaves a cell the clustering did not assign out of the fractions, reported throughreport_drop. The label it contributed became a cluster literally callednanunder pandas 2, and raised under pandas 3, wheresortedcannot order a missing value against the cluster names.subpopulation_hitsdrops it too