• Home
  • Search
  • Classifying next-generation sequencing data using a zero-inflated Poisson model.
  • Cite Icon20
  • https://doi.org/10.1093/bioinformatics/btx768Copy DOI Icon

Classifying next-generation sequencing data using a zero-inflated Poisson model.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

With the development of high-throughput techniques, RNA-sequencing (RNA-seq) is becoming increasingly popular as an alternative for gene expression analysis, such as RNAs profiling and classification. Identifying which type of diseases a new patient belongs to with RNA-seq data has been recognized as a vital problem in medical research. As RNA-seq data are discrete, statistical methods developed for classifying microarray data cannot be readily applied for RNA-seq data classification. Witten proposed a Poisson linear discriminant analysis (PLDA) to classify the RNA-seq data in 2011. Note, however, that the count datasets are frequently characterized by excess zeros in real RNA-seq or microRNA sequence data (i.e. when the sequence depth is not enough or small RNAs with the length of 18-30 nucleotides). Therefore, it is desired to develop a new model to analyze RNA-seq data with an excess of zeros. In this paper, we propose a Zero-Inflated Poisson Logistic Discriminant Analysis (ZIPLDA) for RNA-seq data with an excess of zeros. The new method assumes that the data are from a mixture of two distributions: one is a point mass at zero, and the other follows a Poisson distribution. We then consider a logistic relation between the probability of observing zeros and the mean of the genes and the sequencing depth in the model. Simulation studies show that the proposed method performs better than, or at least as well as, the existing methods in a wide range of settings. Two real datasets including a breast cancer RNA-seq dataset and a microRNA-seq dataset are also analyzed, and they coincide with the simulation results that our proposed method outperforms the existing competitors. The software is available at http://www.math.hkbu.edu.hk/∼tongt. xwan@comp.hkbu.edu.hk or tongt@hkbu.edu.hk. Supplementary data are available at Bioinformatics online.

Similar Papers
  • Research Article

Abstract 2489: The new approach for measuring nonuniformity of read coverages reveals the quality of RNA-seq data

  • Apr 21, 2025
  • Cancer Research
  • Wonyoung Choi +4
  • Research Article

PanGraphRNA: An efficient and flexible bioinformatics platform for graph pangenome-based RNA-seq data analysis.

  • Mar 19, 2026
  • Journal of integrative plant biology
  • Yifan Bu +10
  • PDF
  • Research Article
  • Citations3

Functional regression method for whole genome eQTL epistasis analysis with sequencing data

  • May 18, 2017
  • BMC Genomics
  • Kelin Xu +2
  • Research Article
  • Citations1

Abstract 5478: CICERO: An accurate method for detecting complex and diverse driver fusions using cancer transcriptome sequencing (RNA-seq) data

  • Aug 13, 2020
  • Cancer Research
  • Liqing Tian +18
  • Research Article
  • Citations2

RNA-seq reproducibility of Pseudomonas aeruginosa in laboratory models of cystic fibrosis

  • Dec 03, 2024
  • Microbiology Spectrum
  • Rebecca P Duncan +9
  • Research Article

Abstract 3017: Advancing PDX research through model, data, and bioinformatics with the PDXNet Portal

  • Jul 01, 2021
  • Cancer Research
  • Soner Koc +35
  • PDF
  • Research Article
  • Citations17

Differential expression analysis of RNA sequencing data by incorporating non-exonic mapped reads

  • Jun 11, 2015
  • BMC Genomics
  • Hung-I Harry Chen +6
  • PDF
  • Research Article
  • Citations62

RNA-QC-chain: comprehensive and fast quality control for RNA-Seq data

  • Feb 14, 2018
  • BMC Genomics
  • Qian Zhou +4
  • Research Article
  • Citations93

A marginalized zero-inflated Poisson regression model with overall exposure effects.

  • Sep 14, 2014
  • Statistics in Medicine
  • D Leann Long +3
  • Abstract
  • Citations1

Spatial Scan Statistics for Models with Excess Zeros and Overdispersion

  • Apr 04, 2013
  • Online Journal of Public Health Informatics
  • Max Sousa De Lima +2
  • Abstract

1270 Pretreatment prediction of non-responders to PD-1 axis inhibitors in advanced urothelial carcinomas using a hybrid multimodal deep learning algorithm

  • Nov 01, 2022
  • Journal for ImmunoTherapy of Cancer
  • Fahad Ahmed
  • Research Article
  • Citations10

Accurate Quantification of Overlapping Herpesvirus Transcripts from RNA Sequencing Data.

  • Oct 27, 2021
  • Journal of virology
  • Alejandro Casco +5
  • PDF
  • Research Article
  • Citations41

ShrinkBayes: a versatile R-package for analysis of count-based sequencing data in complex study designs

  • Apr 26, 2014
  • BMC Bioinformatics
  • Mark A Van De Wiel +4
  • Research Article

The impact of misspecification of nuisance parameters on test for homogeneity in zero-inflated Poisson model: A simulation study

  • Jul 30, 2019
  • Communications in Statistics - Simulation and Computation
  • Siyu Gao +2
  • Conference Article

<span>Can Fusion Transcripts between a transposable element and an exon generate piRNAs in mouse (</span><span>Mus musculus)</span><span>?</span>

  • Jan 05, 2019
  • Bairon Hernández +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.