Skip to contents

Hierarchical and high-resolution cell-type identification for single-cell RNA-seq data inspired by ScType.

Features

Installation

You can install the development version of hitype like so:

if (!requireNamespace("devtools", quietly = TRUE)) {
    install.packages("devtools")
}
devtools::install_github("pwwang/hitype")

Optional packages for the individual weight-learning methods: glmnet (method glmnet), ranger (method rf), xgboost (method xgb), keras + innsight (method lrp).

Quick start

Prepare the dataset

See also https://satijalab.org/seurat/articles/pbmc3k_tutorial.html#setup-the-seurat-object

Click to expand
pbmc <- pbmc3k.SeuratData::pbmc3k
pbmc <- Seurat::UpdateSeuratObject(pbmc)
pbmc[["percent.mt"]] <- Seurat::PercentageFeatureSet(pbmc, pattern = "^MT-")
pbmc <- subset(pbmc, subset = nFeature_RNA > 200 & nFeature_RNA < 2500 & percent.mt < 5)
pbmc <- Seurat::NormalizeData(pbmc)
pbmc <- Seurat::FindVariableFeatures(pbmc, selection.method = "vst", nfeatures = 2000)
pbmc <- Seurat::ScaleData(pbmc, features = rownames(pbmc))
pbmc <- Seurat::RunPCA(pbmc, features = Seurat::VariableFeatures(object = pbmc))
pbmc <- Seurat::FindNeighbors(pbmc, dims = 1:10)
pbmc <- Seurat::FindClusters(pbmc, resolution = 0.5)
pbmc <- Seurat::RunUMAP(pbmc, dims = 1:10)

Use as a Seurat extension

library(hitype)

markers <- data.frame(
    cellName = c(
        "Naive CD4+ T", "CD14+ Mono", "Memory CD4+", "B",
        "CD8+ T", "FCFR3A+ Mono", "NK", "DC", "Platelet"
    ),
    geneSymbolmore1 = c(
        "IL7R,CCR7",  "CD14,LYZ", "IL7R,S100A4", "MS4A1",
        "CD8A", "FCGR3A,MS4A7", "GNLY,NKG7", "FCER1A,CST3", "PPBP"
    ),
    geneSymbolmore2 = rep("", 9)
)

# Load gene sets
gs <- gs_prepare(markers)

# Assign cell types
obj <- RunHitype(pbmc, gs)

Seurat::DimPlot(obj, group.by = "hitype", label = TRUE, label.box = TRUE) +
  Seurat::NoLegend()

Compared to the manual marked cell types: Seurat manual marked cell types

See also https://satijalab.org/seurat/articles/pbmc3k_tutorial.html#assigning-cell-type-identity-to-clusters

Use as standalone functions

scores <- hitype_score(Seurat::GetAssayData(pbmc, layer = "data"), gs)
cell_types <- hitype_assign(pbmc$seurat_clusters, scores, gs)
summary(cell_types)

You may see that we have exactly the same assignment in the Seurat tutorial:

Cluster ID Markers Cell Type
0 IL7R, CCR7 Naive CD4+ T
1 CD14, LYZ CD14+ Mono
2 IL7R, S100A4 Memory CD4+
3 MS4A1 B
4 CD8A CD8+ T
5 FCGR3A, MS4A7 FCGR3A+ Mono
6 GNLY, NKG7 NK
7 FCER1A, CST3 DC
8 PPBP Platelet

hitype_score() accepts log-normalized data by default (scaled = FALSE; it performs its own z-scoring) or pre-scaled data with scaled = TRUE. When scoring with learned weights (see below), use norm = "weight", use_sensitivity = FALSE — this normalizes scores by the total marker weight and avoids double-penalizing shared markers.

Train marker weights

train_weights() learns per-marker weights from a labeled dataset (clusters or cell types) with 7 backends:

method Description Requires
uniform Equal weights (baseline) —
correlation Correlation of each marker with the cluster label —
lr Logistic regression (one-vs-rest) coefficients —
glmnet Sparse elastic-net logistic regression (default) glmnet
rf Random forest permutation importance ranger
xgb XGBoost gain-based feature importance xgboost
lrp Neural network + Layer-wise Relevance Propagation keras, innsight
weights <- train_weights(
    path_to_gs = markers,
    exprs = pbmc,
    method = "glmnet",
    cv_folds = 5
)
gs <- gs_prepare(weights)

cv_folds > 1 performs stratified cross-validation inside the training data and averages the weights across folds for stability. Do not evaluate on data used for training — train the weights on a train split and score on a held-out split (or another dataset) to avoid optimistic, circular results. See vignette("train-marker-weights") for a full walkthrough including cross-dataset transfer.

Find markers from your data

find_markers() discovers marker genes from a labeled dataset and returns them directly in the database format consumed by gs_prepare(). The default method = "fc" is dependency-light; "seurat" and "presto" backends are also available:

markers <- find_markers(
    exprs = pbmc, clusters = pbmc$seurat_clusters, method = "fc"
)
weights <- train_weights(
    path_to_gs = markers, exprs = pbmc, method = "glmnet"
)
gs <- gs_prepare(weights)

Fair-use warning: markers and weights derived from the same dataset are for exploratory use only. For publications, keep the marker database fixed and follow the benchmark protocol (see paper_plan.md) to avoid circular results.