Packages maintained by core team

These packages are considered foundational in that many other packages build upon them. Joint maintenance by the core team guarantees long-term stability.

Data structures

Data structures are the foundational building block for all scverse packages. Building upon common data structures ensures interoperability.
Logo for anndata
anndata AnnData is a Python package for handling annotated data matrices in memory and on disk, positioned between pandas and xarray. anndata offers a broad range of computationally efficient features including, among others, sparse data support, lazy operations, and a PyTorch interface.
Logo for mudata
mudata MuData is a format for annotated multimodal datasets where each modality is represented by an AnnData object. MuData’s reference implementation is in Python, and the cross-language functionality is achieved via HDF5-based .h5mu files with libraries in R and Julia.
Logo for spatialdata
spatialdata SpatialData is a data framework that comprises a FAIR storage format and a collection of python libraries for performant access, alignment, and processing of uni- and multi-modal spatial omics datasets. This repository contains the core spatialdata library. See the links below to learn more about other packages in the SpatialData ecosystem.

Analysis task-specific extensions

In addition to these packages, we define standards on how to represent certain data types in these data structures. For now, such a specification is available for Adaptive Immune Receptor Repertoire (AIRR) data.

Frameworks

Frameworks provide essential algorithms and plotting functions for specific analysis steps, building on our data structures.
Logo for scanpy
scanpy Scanpy is a scalable toolkit for analyzing single-cell gene expression data built jointly with anndata. It includes preprocessing, visualization, clustering, trajectory inference and differential expression testing. The Python-based implementation efficiently deals with datasets of more than one million cells.
Logo for muon
muon muon is a Python framework for multimodal omics analysis. While there are many features that muon brings to the table, there are three key areas that its functionality is focused on.
Logo for squidpy
squidpy Squidpy is a tool for the analysis and visualization of spatial molecular data. It builds on top of scanpy and anndata, from which it inherits modularity and scalability. It provides analysis tools that leverages the spatial coordinates of the data, as well as tissue images if available.
Logo for scvi-tools
scvi-tools scvi-tools is a library for developing and deploying machine learning models based on PyTorch and AnnData. With an emphasis on probabilistic models, scvi-tools streamlines the development process via training, data management, and user interface abstractions. scvi-tools also contains easy-to-use implementations of more than 14 state-of-the-art probabilistic models in the field.
Logo for scirpy
scirpy Scirpy is a scalable toolkit to analyse T-cell receptor or B-cell receptor repertoires from single-cell RNA sequencing data. It seamlessly integrates with scanpy and provides various modules for data import, analysis and visualization.
Logo for SnapATAC2
SnapATAC2 SnapATAC2 is a scalable and modular pipeline for analyzing single-cell ATAC-seq data, enabling efficient preprocessing, dimensionality reduction, clustering, and integration with single-cell RNA-seq.
Logo for rapids-singlecell
rapids-singlecell rapids-singlecell is a GPU-accelerated single-cell analysis library that serves as a drop-in replacement for scanpy, squidpy, and decoupler.
Logo for pertpy
pertpy Pertpy is a framework for analyzing large-scale single-cell perturbation experiments. It harmonizes datasets, automates metadata annotation, calculates perturbation distances, and analyzes cellular responses to genetic modifications, drugs, and environmental changes.
Logo for decoupler
decoupler decoupler is a framework containing different enrichment statistical methods to extract biologically driven scores from omics data within a unified framework.

Ecosystem packages maintained by scverse community

Many popular packages rely on scverse functionality. For instance, they take advantage of established data format standards such as AnnData and MuData, or are designed to be integrated into the workflow of analysis frameworks. Here, we list ecosystem packages following development best practices (continuous testing, documented, available through standard distribution tools).

This listing is a work in progress. See scverse/ecosystem-packages for inclusion criteria, and to submit more packages.

118 of 118 packages

AESTETIK

v0.3.1 MIT

Convolutional autoencoder that learns spot representations from spatial transcriptomics by jointly integrating transcriptomics, morphology (H&E), and spatial-neighborhood topology. The learned embeddings support downstream tasks such as spatial-domain clustering and multi-modal analysis.

AlphaPeptTools

0.2.0 Apache-2.0

AnnData-based downstream proteomics analysis package with comprehensive support of search engine report formats. AlphaPeptTools features an io module for loading and standardizing MS-proteomics data, as well as pp, pl, tl and metrics modules to facilitate common proteomics-specific data processing, visualization and analysis tasks with a focus on metrics for evaluating the effect of different processing steps. Additionally, it features several example notebooks to demonstrate different proteomics use cases.

AnnData

0.12.4 BSD-3-Clause

anndata is a Python package for handling annotated data matrices in memory and on disk, positioned between pandas and xarray. anndata offers a broad range of computationally efficient features including, among others, sparse data support, lazy operations, and GPU support.

anndata for R

0.7.5.5 MIT

⚠️ DEPRECATED: This package is deprecated in favor of anndataR. Please use anndataR for new projects. A ‘reticulate’ wrapper for the Python package ‘anndata’. Provides a scalable way of keeping track of data and learned annotations. Used to read from and write to the h5ad file format.

anndataR

1.0.0 MIT

An R implementation of the AnnData object including native reading and writing of H5AD files and conversion to common R data structures.

annsel

v0.0.8 MIT

Annsel is a user-friendly library that brings familiar dataframe-style operations to AnnData objects such as selection, filtering and group by’s.

benGRN

v1.2.1 MIT

Benchmarking tool for gene network inference from single cell RNAseq methods. It uses the grnndata/anndata modality and only contains biological ground truth networks

biolord

v0.0.1 BSD-3-Clause

biolord (biological representation disentanglement) is a deep generative framework for disentangling known and unknown attributes in single-cell data.

cell2location

v0.1 Apache-2.0

Cell2location is a Bayesian model that can resolve fine-grained cell types in spatial transcriptomic data and create comprehensive cellular maps of diverse tissues. Cell2location accounts for technical sources of variation and borrows statistical strength across locations, thereby enabling the integration of single-cell and spatial transcriptomics with higher sensitivity and resolution than existing tools.

CellAnnotator

v0.1.3 MIT

CellAnnotator is a lightweight tool to query large language models for cell type labels in scRNA-seq data. It can incorporate prior knowledge, and it creates consistent labels across samples in your study.

CellCharter

v0.3.1 BSD-3-Clause

CellCharter is a framework to identify, characterize and compare spatial domains from spatial omics and multi-omics data.

CellMapper

v0.1.2 MIT

CellMapper is a leightweight tool to transfer labels, expression values and embeddings from reference to query datasets using k-NN mapping. It’s fast and versatile, applicable to mapping scenarios in space, across modalities, or from an atlas to a new query dataset.

CellOracle

v0.10.12 Apache-2.0

A computational tool that integrates single-cell transcriptome and epigenome profiles to infer gene regulatory networks (GRNs), critical regulators of cell identity.

CellphoneDB

v5.0.0 MIT

CellphoneDB is a publicly available repository of HUMAN curated receptors, ligands and their interactions paired with a tool to interrogate your own single-cell transcriptomics data (or even bulk transcriptomics data if your samples represent pure populations!). A distinctive feature of CellphoneDB is that the subunit architecture of either ligands and receptors is taken into account, representing heteromeric complexes accurately. This is crucial, as cell communication relies on multi-subunit protein complexes that go beyond the binary representation used in most databases and studies. CellphoneDB also incorporates biosynthetic pathways in which we use the last representative enzyme as a proxy of ligand abundance, by doing so, we include interactions involving non-peptidic molecules. CellphoneDB includes only manually curated and reviewed molecular interactions with evidenced role in cellular communication.

CellRank

v1.5.1 BSD-3-Clause

CellRank is a toolkit to uncover cellular dynamics based on Markov state modeling of single-cell data. It contains two main modules - kernels compute cell-cell transition probabilities and estimators generate hypothesis based on these.

cellxgene

1.1.1 MIT

CZ CELLxGENE Annotate (pronounced “cell-by-gene”) is an interactive data explorer for single-cell datasets, such as those coming from the Human Cell Atlas.

dandelion

v0.3.0 AGPL-3.0-or-later

dandelion - A single cell BCR/TCR V(D)J-seq analysis package for 10X Chromium 5’ data. It streamlines the pre-processing, leveraging some tools from immcantation suite, and integrates with scanpy/anndata for single-cell BCR/TCR analysis. It also includes a couple of functions for visualization.

decoupler

2.1.1 BSD-3-Clause

decoupler is a framework containing different enrichment statistical methods to extract biologically driven scores from omics data within a unified framework.

delnx

v0.2.3 MIT

delnx is a python package for differential expression analysis of (single-cell) genomics data. It enables scalable analyses of atlas-level datasets through GPU-accelerated regression models and statistical tests implemented in JAX and provides a consistent interface to perform DE analysis with other methods, such as statsmodels and PyDESeq2.

DOTools_py

v0.0.2 MIT

Convenient and user-friendly package to streamline common workflows in single-cell RNA sequencing data analysis with improved visualisation

DRVI

0.2.0 BSD-3-Clause

DRVI is a tool for the unsupervised disentanglement and integration of single-cell omics. By providing interpretable latent dimensions, it allows users to identify cellular heterogeneity and biological processes beyond traditional cell types, identify rare cell types, and highlight developmental stages. DRVI is implemented using scvi-tools and includes a set of utility functions for interacting with latent dimensions.

dynamo-release

v1.1.0 BSD-3-Clause

Inclusive model of expression dynamics with metabolic labeling based scRNA-seq / multiomics, vector field reconstruction, potential landscape mapping, differential geometry analyses, and most probably paths / in silico perturbation predictions.

ecosystem-packages

BSD-3-Clause

Registry for scverse ecosystem packages (https://scverse.org/packages/#ecosystem)

epiScanpy

v0.3.2 BSD-3-Clause

EpiScanpy is a toolkit to analyse single-cell open chromatin (scATAC-seq) and single-cell DNA methylation (for example scBS-seq) data.

eschr

v1.0.1 MIT

ESCHR is an ensemble clustering method that provides hard clustering along with uncertainty scores and soft clustering outputs for enhanced interpretability.

fava

v0.3.9.4 MIT

FAVA uses Variational Autoencoders to infer functional associations from large-scale scRNA-seq (and proteomics) data.

flashdeconv

v0.1 BSD-3-Clause

FlashDeconv is a high-performance spatial transcriptomics deconvolution tool that enables cell type mapping at atlas scale. Using structure-preserving sketching via randomized numerical linear algebra, FlashDeconv achieves linear time and memory complexity, processing 1 million spots in approximately 3 minutes on a standard laptop without GPU. It provides accurate cell type proportion estimation with Pearson r = 0.944 on the Spotless benchmark, and preserves rare cell type detection through leverage-score weighted sampling.

flowsom

v0.0.1 GPL-3.0-only

The complete FlowSOM package known from R, now available in Python! Analyze high-dimensional cytometry data using FlowSOM, a clustering and visualization algorithm based on a self-organizing map (SOM). FlowSOM is used to distinguish cell populations from cytometry data in an unsupervised way and can help to gain deeper insights in fields such as immunology and oncology.

GPTBioInsightor

v0.3.0 BSD-3-Clause

GPTBioInsightor is a tool designed for single-cell data analysis, particularly beneficial for newcomers to a biological field or those in interdisciplinary areas who may lack sufficient biological background knowledge. GPTBioInsightor utilizes the powerful capabilities of large language models to help people quickly gain knowledge and insight, enhancing their work efficiency.

grassp

v0.1.0 BSD-3-Clause

grassp (GRaph-based Analysis of Subcellular/Spatial Proteomics) is a Python module for fast, flexible, and scalable analysis of mass-spectrometry-based subcellular proteomics datasets. It uses the anndata format to store mass-spec data and results. scanpy is used for dimensionality reduction and visualization functions.

GRnnData

v1.1.4 MIT

An overload of anndata to more easily work with gene networks. Allows easy conversion between anndata and grnndata and provide loads of useful utilities functions.

illico

0.1.1 Apache-2.0

illico runs fast, CPU-based, wilcoxon rank-sum tests to identify differentially expressed genes for single-cell RNA-seq data.

A repo for integration testing core packages against upstream core packages

kompot

v0.6.1 GPL-3.0-or-later

Statistical framework for holistic comparison of multi-condition single-cell datasets over continuous phenotypic manifolds. Kompot performs differential abundance and expression analysis agnostically to cell-type annotation or clustering, detecting differences that span the continuous manifold even when effects go in opposite directions across cell states. It models cell distributions and gene expression as continuous functions over cell state representations of arbitrary dimension (typically ~50 dimensions), enabling single-cell resolution inference with calibrated uncertainty estimates and capturing both global and cell-state-specific effects of perturbation.

LazySlide

v0.3.0 MIT

LazySlide is a Python library for processing whole slide images (WSI) analysis. It provides a simple interface to perform robust preprocessing and advanced analysis for WSI.

liana

v1.0.0a1 GPL-3.0-only

Python package to infer cell-cell communication events from omics data using a collection of methods.

moscot

v0.4.0 BSD-3-Clause

moscot is a scalable toolbox for multiomics single-cell optimal transport applications.

Mowgli

v0.2.0 GPL-3.0-only

Paired single-cell multi-omics data integration with Optimal Transport-flavored Nonnegative Matrix Factorization

MuData

0.3.2 BSD-3-Clause

MuData is a format for annotated multimodal datasets where each modality is represented by an AnnData object. MuData’s reference implementation is in Python, and the cross-language functionality is achieved via HDF5-based .h5mu files with libraries in R and Julia.

Muon

0.1.7 BSD-3-Clause

muon is a Python framework for multimodal omics analysis. It provides functionality for working with multimodal data, including preprocessing, integration, and visualization.

nichepca

v0.0.3 MIT

A Python package for PCA-based spatial domain identification in single-cell spatial transcriptomics data.

novae

v0.2.1 BSD-3-Clause

Graph-based foundation model for spatial transcriptomics data. Zero-shot spatial domain inference, batch-effect correction, and many other features.

omicverse

v1.4.12 GPL-3.0-only

OmicVerse is the fundamental package for multi omics included bulk and single cell analysis with Python. The original name of the omicverse was Pyomic, but we wanted to address a whole universe of transcriptomics, so we changed the name to OmicVerse, it aimed to solve all task in RNA-seq.

Palantir

v1.3.3 GPL-2.0-or-later

Palantir is an algorithm to align cells along differentiation trajectories. Palantir models differentiation as a stochastic process where stem cells differentiate to terminally differentiated cells by a series of steps through a low dimensional phenotypic manifold. Palantir effectively captures the continuity in cell states and the stochasticity in cell fate determination. Palantir has been designed to work with multidimensional single cell data from diverse technologies such as Mass cytometry and single cell RNA-seq.

ParTIpy

v0.0.04 MIT

Implements Pareto task inference and archetypal analysis for analyzing functional trade-offs in single-cell and spatial omics data.

pcdl

v4.0.4 BSD-3-Clause

physicell data output loader for downstream data analysis in python3.

Pertpy

1.0.3 MIT

Pertpy is a framework for analyzing large-scale single-cell perturbation experiments. It harmonizes datasets, automates metadata annotation, calculates perturbation distances, and analyzes cellular responses to genetic modifications, drugs, and environmental changes.

PILOT

v2.0.6 MIT

PILOT is a Python library for Detection of PatIent-Level distances from single cell genomics and pathomics data with Optimal Transport.

popV

v0.5.2 MIT

p(opular)V(oting) is a consensus tool for transfering labels from an annotated reference dataset to an unannotated query dataset. Consensus calling allows interpretable scores that quantify certainty.

pycea

v0.1.0 BSD-3-Clause

Pycea is a Python toolkit for single-cell lineage tracing analysis and visualization.

pyCrossTalkeR

v2.1.0 MIT

pyCrossTalkeR is a framework for network analysis and visualisation of LR networks. pyCrossTalkeR identifies relevant ligands, receptors and cell types contributing to changes in cell communication when contrasting two biological states: disease vs. homeostasis.

PyDESeq2

v0.3.0 MIT

PyDESeq2 is a python package for bulk RNA-seq differential expression analysis. It is a re-implementation from scratch of the main features of the R package DESeq2 (Love et al. 2014).

pySCENIC

v0.12.0 GPL-3.0-only

pySCENIC is a lightning-fast python implementation of the SCENIC pipeline (Single-Cell rEgulatory Network Inference and Clustering) which enables biologists to infer transcription factors, gene regulatory networks and cell types from single-cell RNA-seq data.

pytximport

v0.2.0 GPL-3.0-only

A Python port of the tximport R package for importing transcript-level quantification data from various RNA-seq quantification tools such as salmon and kallisto and summarizing it to the gene level.

pyUCell

v0.3.0 MIT

pyUCell is a package for evaluating gene signatures in single-cell datasets. pyUCell signature scores, based on the Mann-Whitney U statistic, are robust to dataset size and heterogeneity, and their calculation demands less computing time and memory than other available methods, enabling the processing of large datasets in a few minutes even on machines with limited computing power.

Rectangle

v0.1.6 MIT

Rectangle is a python package for computational deconvolution. Rectangle presents a novel approach to second-generation deconvolution, characterized by hierarchical processing, an estimation of unknown cellular content and a significant reduction in data volume during signature matrix computation.

SC2Spa

v1.2 BSD-3-Clause

SC2Spa is a deep learning-based tool for predicting the spatial coordinates of single cells based on transcriptome. Two paired single cell and spatial transcriptomic datasets are required to run SC2Spa. SC2Spa is trained on a ST reference dataset to learn the relationship of gene expression and spatial coordinates. The trained fully-connected neural network can be used to predict the locations of a single cell with only the transcriptomic profile as input. The predicted locations of single cells can be further used to study the communication of the single cells.

SCALEX

v1.0.3 MIT

SCALEX is an integration and projection tool for atlas-level single-cell RNA-seq and ATAC-seq data.

Scanpy

1.11.5 BSD-3-Clause

Scanpy is a scalable toolkit for analyzing single-cell gene expression data built jointly with anndata. It includes preprocessing, visualization, clustering, trajectory inference and differential expression testing. The Python-based implementation efficiently deals with datasets of more than one million cells.

scCellFie

v0.4.5 MIT

scCellFie infers metabolic activities from single-cell and spatial transcriptomics and offers a variety of downstream analyses.

scDataLoader

v1.2.2 MIT

A dataloader for large single cell databases like cellxgene. Does weighted random sampling, downloading and preprocessing. works with anndata, zarr, and h5ad files.

scib-rapids

0.1.0 BSD-3-Clause

GPU-accelerated single-cell integration benchmarking metrics using RAPIDS (cuML, CuPy) as a drop-in replacement for the JAX-based metrics in scib-metrics.

Scirpy

0.22.3 BSD-3-Clause

Scirpy is a scalable toolkit to analyse T-cell receptor or B-cell receptor repertoires from single-cell RNA sequencing data. It seamlessly integrates with scanpy and provides various modules for data import, analysis and visualization.

scPRINT

v1.6.2 MIT

A single cell foundation model for Gene network inference and more…

scPRINT-2

v1.0.0 GPL-3.0-or-later

A next generation single cell foundation model

scVelo

v0.2.5 BSD-3-Clause

scVelo is a scalable toolkit for RNA velocity analysis in single cells, based on Bergen et al., Nature Biotech, 2020.

scvi-tools

1.4.0.post1 BSD-3-Clause

scvi-tools is a library for developing and deploying machine learning models based on PyTorch and AnnData. With an emphasis on probabilistic models, scvi-tools streamlines the development process via training, data management, and user interface abstractions. scvi-tools also contains easy-to-use implementations of more than 14 state-of-the-art probabilistic models in the field.

scxmatch

v0.1.0 MIT

Single-cell Cross Match (scxmatch) is a is a Python package that implements Rosenbaum’s cross-match test using distance-based matching to assess distribution shifts between two groups of high-dimensional data. This is particularly useful in analyzing multivariate distributions in structured data, such as single-cell RNA-seq or ATAC-seq.

scXpand

v0.4.3 MIT

scXpand is a machine learning framework for pan-cancer detection of T-cell clonal expansion directly from single-cell RNA sequencing (scRNA-seq), without paired T-cell receptor (TCR) sequencing.

scyan

v1.5.0 BSD-3-Clause

Biology-driven deep generative model for cell-type annotation in cytometry. Scyan is an interpretable model that also corrects batch-effect and can be used for debarcoding or population discovery.

sift-sc

v0.1.0 BSD-3-Clause

SiFT is a computational framework which aims to uncover the underlying structure by filtering out previously exposed biological signals. SiFT can be applied to a wide range of tasks, from (i) the removal of unwanted variation as a pre-processing step, through (ii) revealing hidden biological structure by utilizing prior knowledge with respect to existing signal, to (iii) uncovering trajectories of interest using reference data to remove unwanted variation.

sincei

v0.5.1 MIT

sincei provides a flexible, easy-to-use command-line interface and python API to work with single-cell epigenomics data directly from BAM files. It can: Aggregate signal in bins, genes or any feature of interest from single-cells. Perform read-level and count-level quality control. Perform dimensionality reduction and clustering of various types of single-cell data (open chromatin, histone marks, methylation etc..). Create coverage files (bigwigs) for visualization.

SnapATAC2

2.5.0 MIT

SnapATAC2 is a scalable and modular pipeline for analyzing single-cell ATAC-seq data, enabling efficient preprocessing, dimensionality reduction, clustering, and integration with single-cell RNA-seq.

Sobolev alignment of deep probabilistic models for comparing single cell profiles from pre-clinical models and patients

sopa

v1.0.0 BSD-3-Clause

Technology-invariant pipeline for spatial-omics analysis that scales to millions of cells. It includes segmentation, annotation, spatial statistics, and efficient visualization.

Python package designed to transfer information from multiple spatial-transcriptomics data sets to a single reference representing a Common Coordinate Framework (CCF).

SpatialData

0.5.0 BSD-3-Clause

SpatialData is a data framework that comprises a FAIR storage format and a collection of python libraries for performant access, alignment, and processing of uni- and multi-modal spatial omics datasets.

Spatialproteomics is an interoperable toolbox for analyzing highly multiplexed fluorescence image data. This analysis involves a sequence of steps, including segmentation, image processing, marker quantification, cell type classification, and neighborhood analysis.

spatiomic

v0.5.0 GPL-3.0-only

spatiomic is a computational library for the analysis of spatial proteomics, mainly via pixel-based clustering, differential cluster abundance analysis and spatial statistics.

Squidpy

1.6.5 BSD-3-Clause

Squidpy is a tool for the analysis and visualization of spatial molecular data. It builds on top of scanpy and anndata, from which it inherits modularity and scalability. It provides analysis tools that leverages the spatial coordinates of the data, as well as tissue images if available.

STMiner

v1.1.0 GPL-3.0-or-later

Gene-centric spatial transcriptomics for deciphering complex spatial omics data

Symphonypy

v0.2.1 GPL-3.0-only

Symphonypy is a pure Python port of Symphony label transfer algorithm for reference-based cell type annotation.

Evolutionary graph community detection combining genetic search with Leiden refinement, with native support for clustering Scanpy neighbor graphs stored in AnnData.

TreeData

v0.2.2 BSD-3-Clause

TreeData is a lightweight wrapper around AnnData which adds two additional attributes, obst and vart, to store nx.DiGraph trees for observations and variables.

vitessce

v3.5.7 MIT

Vitessce is an integrative visualization framework for multimodal and 2D/3D spatially resolved single-cell data. It consists of reusable, interactive linked views including scatterplot embedding views, 2D/3D spatial image views, genome browser tracks, statistical plots, and control views, built on web technologies such as WebGL and WebXR.

wsidata

v0.3.0 MIT

wsidata is a data-structure for efficient IO of Whole Slide Images based on spatialdata.

zellkonverter

1.20.0 MIT

An Bioconductor R package that enables reading and writing of H5AD files and conversion between AnnData and SingleCellExperiment objects by wrapping the Python anndata package.