• Home
  • Search
  • Robust principal component analysis for accurate outlier sample detection in RNA-Seq data
  • Cite Icon111
  • https://doi.org/10.1186/s12859-020-03608-0Copy DOI Icon

Robust principal component analysis for accurate outlier sample detection in RNA-Seq data

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

BackgroundHigh throughput RNA sequencing is a powerful approach to study gene expression. Due to the complex multiple-steps protocols in data acquisition, extreme deviation of a sample from samples of the same treatment group may occur due to technical variation or true biological differences. The high-dimensionality of the data with few biological replicates make it challenging to accurately detect those samples, and this issue is not well studied in the literature currently. Robust statistics is a family of theories and techniques aim to detect the outliers by first fitting the majority of the data and then flagging data points that deviate from it. Robust statistics have been widely used in multivariate data analysis for outlier detection in chemometrics and engineering. Here we apply robust statistics on RNA-seq data analysis.ResultsWe report the use of two robust principal component analysis (rPCA) methods, PcaHubert and PcaGrid, to detect outlier samples in multiple simulated and real biological RNA-seq data sets with positive control outlier samples. PcaGrid achieved 100% sensitivity and 100% specificity in all the tests using positive control outliers with varying degrees of divergence. We applied rPCA methods and classical principal component analysis (cPCA) on an RNA-Seq data set profiling gene expression of the external granule layer in the cerebellum of control and conditional SnoN knockout mice. Both rPCA methods detected the same two outlier samples but cPCA failed to detect any. We performed differentially expressed gene detection before and after outlier removal as well as with and without batch effect modeling. We validated gene expression changes using quantitative reverse transcription PCR and used the result as reference to compare the performance of eight different data analysis strategies. Removing outliers without batch effect modeling performed the best in term of detecting biologically relevant differentially expressed genes.ConclusionsrPCA implemented in the PcaGrid function is an accurate and objective method to detect outlier samples. It is well suited for high-dimensional data with small sample sizes like RNA-seq data. Outlier removal can significantly improve the performance of differential gene detection and downstream functional analysis.

Loading PDF

Similar Papers
  • Research Article

Abstract 1817: Differential expression of long non-coding RNA in colon adenocarcinoma RNA-sequence data set

  • Jul 01, 2019
  • Cancer Research
  • Stephen J O'Brien +4
  • Research Article

PanGraphRNA: An efficient and flexible bioinformatics platform for graph pangenome-based RNA-seq data analysis.

  • Mar 19, 2026
  • Journal of integrative plant biology
  • Yifan Bu +10
  • Research Article
  • Citations20

Classifying next-generation sequencing data using a zero-inflated Poisson model.

  • Nov 27, 2017
  • Bioinformatics
  • Yan Zhou +3
  • PDF
  • Research Article
  • Citations17

Differential expression analysis of RNA sequencing data by incorporating non-exonic mapped reads

  • Jun 11, 2015
  • BMC Genomics
  • Hung-I Harry Chen +6
  • Research Article

Development and validation of the PipeSeq program for RNA-seq data analysis in the Chlamydomonas reinhardtii as a model.

  • Apr 01, 2026
  • Vavilovskii zhurnal genetiki i selektsii
  • A M Nerezenko +3
  • Research Article

Abstract 2489: The new approach for measuring nonuniformity of read coverages reveals the quality of RNA-seq data

  • Apr 21, 2025
  • Cancer Research
  • Wonyoung Choi +4
  • Research Article
  • Citations178

Protein Identification Using Customized Protein Sequence Databases Derived from RNA-Seq Data

  • Dec 14, 2011
  • Journal of Proteome Research
  • Xiaojing Wang +6
  • Book Chapter
  • Citations1

Chapter 19 - Dimensionality Reduction and Latent Variable Modeling

  • Jan 01, 2020
  • Machine Learning
  • Sergios Theodoridis
  • Conference Article

Research of liquid precursor chemicals recognition based on principal component analysis

  • Mar 01, 2013
  • Daoyang Yu +4
  • Research Article
  • Citations2

RNA-seq reproducibility of Pseudomonas aeruginosa in laboratory models of cystic fibrosis

  • Dec 03, 2024
  • Microbiology Spectrum
  • Rebecca P Duncan +9
  • PDF
  • Research Article
  • Citations28

Getting the most out of RNA-seq data analysis

  • Oct 29, 2015
  • PeerJ
  • Tsung Fei Khang +1
  • Research Article
  • Citations19

High-Throughput Sequence Analysis of Peripheral T-Cell Lymphomas Indicates Subtype-Specific Viral Gene Expression Patterns and Immune Cell Microenvironments.

  • Jul 10, 2019
  • mSphere
  • Hani Nakhoul +5
  • Research Article
  • Citations5

Combining a robust PCA of logratio transformed data and geostatistical sequential Gaussian simulation approach for geochemical characterization of orogenic gold deposits: a case study from the Alut area, NW of Iran

  • May 09, 2019
  • Geochemistry: Exploration, Environment, Analysis
  • Fereydoun Sharifi +3
  • PDF
  • Research Article
  • Citations8

Identification of hub genes and transcription factor regulatory network for heart failure using RNA-seq data and robust rank aggregation analysis

  • Oct 28, 2022
  • Frontiers in Cardiovascular Medicine
  • Dingyuan Tu +6
  • Conference Article
  • Citations5

Robust Principal Component Analysis Using Alpha Divergence

  • Oct 01, 2020
  • Aref Miri Rekavandi +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.