Published November 2021 | Version v1
Journal article

High-throughput single cell data analysis – A tutorial

  • 1. Radboud University, Institute for Molecules and Materials, Analytical Chemistry, P.O. Box 9010, 6500, GL, Nijmegen (Netherlands)
  • 2. Department of Internal Medicine, Laboratory of Metabolism and Vascular Medicine, P.O. Box 616 (UNS50/14), 6200, MD, Maastricht (Netherlands)
  • 3. Corbion, Arkelsedijk 46, 4206, AC, Gorinchem (Netherlands)

Description

Highlights: • Single cell analysis may measure multiple parameters on millions of single cells. • How to: analyze the study design and define a research question. • Prepare high quality data, including compensation and removal of unwanted cells. • Convert the millions of single cells per sample into a cellular distribution. • Top model to highlight the most deviating cell (sub)types. White blood cells protect the body against disease but may also cause chronic inflammation, auto-immune diseases or leukemia. There are many different white blood cell types whose identity and function can be studied by measuring their protein expression. Therefore, high-throughput analytical instruments were developed to measure multiple proteins on millions of single cells. The information-rich biochemistry information may only be fully extracted using multivariate statistics. Here we show an overview of the most essential steps for multivariate data analysis of single cell data. We used white blood cells (immunology) as a case study, but a similar approach may be used in environment or biotech research. The first step is analyzing the study design and subsequently formulating a research question. The three main designs are immunophenotyping (finding different cell types), cell activation and rare cell discovery. When preparing the data it is essential to consider the design and focus on the cell type of interest by removing all unwanted events. After pre-processing, the ten-thousands to millions of single cells per sample need to be converted into a cellular distribution. For immunophenotyping a clustering method such as Self-Organizing Maps is useful and for cell activation a model that describes the covariance such as Principal Component Analysis is useful. In rare cell discovery it is useful to first model all common cells and remove them to find the rare cells. Finally discriminant analysis based on the cellular distribution may highlight which cell (sub)types are different between groups.

Availability note (English)

Available from http://dx.doi.org/10.1016/j.aca.2021.338872

Additional details

Identifiers

DOI
10.1016/j.aca.2021.338872;
PII
S000326702100698X;

Publishing Information

Journal Title
Analytica Chimica Acta
Journal Volume
1185
Journal Page Range
vp.
ISSN
0003-2670
CODEN
ACACAM

Optional Information

Copyright
Copyright (c) 2021 The Authors. Published by Elsevier B.V.