Pooled rare variants#
pooled_rare() serves cpg0032-pooled-rare, already aggregated to one profile per barcode. There is no plate, well or cell count in obs: this is the example of a pre-aggregated input, the shape a screen takes once wells have been collapsed to perturbations upstream. The 290 barcodes cover 132 genes, each barcode either a wild-type gene or one specific coding variant of it, so the genes come as allele series.
Use it if your own data is already aggregated to one profile per perturbation, with no well or cell counts, or is an allele series (a wild type plus coding variants of the same gene). No tutorial is built on it; the profile-space tools apply directly, since the aggregation is already done.
import numpy as np
import pandas as pd
import plotly.express as px
import scanpy as sc
import mantispy as mt
adata = mt.ds.pooled_rare()
adata
AnnData object with n_obs × n_vars = 290 × 4417
obs: 'Metadata_Foci_Barcode_MatchedTo_GeneCode', 'Metadata_Perturbation', 'Metadata_Perturbation_Type', 'Metadata_Gene', 'Metadata_Allele'
var: 'object', 'feature_group', 'feature', 'channel', 'scale', 'angle', 'gray_levels', 'radial_bin', 'params', 'is_feature'
uns: 'mantispy'
layers: None (.X)
An allele series#
Metadata_Gene is the gene and Metadata_Allele the full variant label, so the rows of one gene are its wild type and its variants.
obs = adata.obs
obs.loc[obs["Metadata_Gene"] == "ACTB", ["Metadata_Gene", "Metadata_Allele", "Metadata_Perturbation"]]
| Metadata_Gene | Metadata_Allele | Metadata_Perturbation | |
|---|---|---|---|
| 0 | ACTB | ACTB | ACTB |
| 1 | ACTB | ACTB E364K | ACTB E364K |
| 2 | ACTB | ACTB R183W | ACTB R183W |
Barcodes in profile space#
On the first two principal components each point is a barcode. Colouring the wild-type genes apart from the variant alleles shows how far a coding change moved the profile from its reference.
scaled = adata.copy()
sc.pp.scale(scaled)
scaled.X = np.nan_to_num(scaled.X)
sc.pp.pca(scaled, n_comps=10)
coords = pd.DataFrame(scaled.obsm["X_pca"][:, :2], columns=["PC1", "PC2"])
coords["allele"] = adata.obs["Metadata_Allele"].to_numpy()
is_variant = adata.obs["Metadata_Allele"].astype(str) != adata.obs["Metadata_Gene"].astype(str)
coords["kind"] = np.where(is_variant.to_numpy(), "variant", "wild type")
fig = px.scatter(
coords,
x="PC1",
y="PC2",
color="kind",
hover_name="allele",
title="Pooled rare variants in profile space",
)
fig.show()