Skip to content
scDrugPerturb-Bench / Dataset

The largest open atlas for mechanism-aware single-cell drug perturbation prediction.

scDrugPerturb-Bench is a virtual cell dataset for drug and gene perturbation prediction models. Each retained case connects drug, dose, time, cellular context and key-gene response evidence, making the dataset useful for building and evaluating intelligent virtual-cell models.

Versioned resourceMechanism-aware annotationsSingle-cell RNA-seq
Comparison of scDrugPerturb-Bench with representative perturbation prediction datasets across data coverage and mechanism annotation.
Figure 1c. Data coverage and mechanism annotation compared with representative perturbation-prediction datasets.
At a glance

Large scale, but built around evidence rather than size alone.

The benchmark starts from more than 50,000 PubMed records and retains manually verified studies where measured single-cell responses are aligned with experimentally supported drug-response mechanisms.

2.5Msingle cells after unified processing
181drug perturbation datasets
93source studies retained
101curated drugs
137cellular contexts
717unique annotated key genes
Construction

From literature to benchmark-ready matrices.

The curation workflow extracts biological claims from papers, connects them to single-cell matrices, verifies the measured response and removes cases where literature evidence and expression data disagree.

Data collection workflow for scDrugPerturb-Bench.
Data-collection workflow. Candidate studies are retrieved, screened, processed, manually verified, cleaned and quality controlled before entering the benchmark.

Literature-grounded cases

Each case records the cellular context, drug identity, dose, time and post-perturbation gene changes supported by the source paper.

Matched expression matrices

The final resource includes matched control matrices, drug-perturbed matrices, drug metadata and curated response annotations.

Unified processing

All retained data pass through a consistent workflow for cell filtering, gene mapping, gene filtering and normalization.

Coverage

Five experimental sources, broad biological context.

The dataset preserves the heterogeneity found in real perturbation studies, from controlled cell-line experiments to organoids, primary cultures and patient-derived samples.

Dataset sources

Cell lines
66
Primary cultures
47
Organoids
47
Patient samples
18
PDX
3

Patient-derived xenograft data are retained as a training source because the category is too small for a standalone test split.

Dataset-level composition of scDrugPerturb-Bench across sources, tissue groups, platforms, dose-time designs and response cases.
Dataset composition across experimental source, tissue group, profiling platform, dose-time design and response direction.
Annotations

Mechanism-relevant gene changes are the core signal.

Instead of treating expression reconstruction as the only objective, scDrugPerturb-Bench records directionally annotated response genes so predictions can be judged by whether they recover the biology reported in the literature.

Example of extracting drug, dose, cell type, gene-response direction and supporting evidence from a source paper into a structured benchmark record.
Data-extraction case. Evidence from article text, figures and expression results is cross-validated and organized into a structured perturbation record.

Curated response cases

The current benchmark contains 423 directionally annotated response cases, enabling model evaluation at the level of key genes, gene sets and pathways.

195upregulated response cases
126downregulated response cases
102non-significant response cases
10x Chromium, 135 datasets Combinatorial indexing Drop-seq TempO-LINC Microwell Plate-based assays Multi-platform studies

Experimental designs range from single-dose, single-time measurements to multidose and multitime studies.

Why it matters

Designed for mechanism-aware virtual-cell modeling.

Existing perturbation resources are valuable expression-profile collections, but many do not provide case-specific directional response annotations. scDrugPerturb-Bench couples multi-source perturbation matrices with manually curated mechanism-level evidence.

Beyond expression similarity

Models can be evaluated on whether they recover response direction, effect magnitude, mechanism specificity and pathway-level response polarity.

Generalization-ready splits

The benchmark supports cell-line and source-aware scenarios, including OOD drug, cell, tissue, drug-cell pair and unseen-source settings.

Versioned and extensible

The dataset is designed to grow as new single-cell drug perturbation studies and mechanism annotations become available.

Access

Public sample available now.

Twenty representative cell-line cases are available on Hugging Face. For access to the full dataset, contact the Simucella team.

Research page

Read the benchmark story, mechanism metrics and evaluation results.

Open research page