• Home
  • Search
  • Comparing Four Genome-Wide Association Study (GWAS) Programs with Varied Input Data Quantity
  • Cite Icon5
  • https://doi.org/10.1109/bibm.2018.8621425Copy DOI Icon

Comparing Four Genome-Wide Association Study (GWAS) Programs with Varied Input Data Quantity

  • Dec 1, 2018
  • Yan Yan +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Genome-wide association studies (GWAS) have served as primary methods for the past decade for identifying associations between genetic variants and traits or diseases. Many software packages have been developed for GWAS analysis based on different statistical models. One key factor influencing the statistical reliability of GWAS is the amount of input data used. Few studies have been conducted to investigate this effect by comparing the performance of GWAS programs using varied amounts of experimental data, especially in the context of plants and plant genomes. In this paper, we investigate how input data quantity influences output of four widely used GWAS programs, PLINK, TASSEL, GAPIT, and FaST-LMM. Both synthetic and real data are used. Standard GWAS output includes single nucleotide polymorphisms (SNPs) and their p-values. To evaluate the programs, p-values and q-values of SNPs, and Kendall rank correlation between output SNP lists, are used. Results show that with the same GWAS program, different Arabidopsis thaliana datasets demonstrate similar trends of rank correlation with varied input quantity, but differentiate on the numbers of SNPs passing a given p- or q-value threshold. In practice, experimental datasets may have samples containing varied numbers of biological replicates. We show that this variation in replicates influences the p-values of SNPs, but does not strongly affect the rank correlation. When comparing synthetic and real data, the output SNPs from synthetic data have similar rank correlation trends across all four GWAS programs, but the same measure from real data is diverse across the programs. In addition, the real data results in a linear-like increase in the numbers of significant SNPs with more input data, but the synthetic data does not follow this trend. This study provides guidance on selecting GWAS programs when varied experimental data is present and on selecting significant SNPs for subsequent study. It contributes to understanding how much input data is necessary to yield satisfying GWAS results.

Similar Papers
  • Research Article
  • Citations9

Effects of input data quantity on genome-wide association studies (GWAS)

  • Jan 01, 2019
  • International Journal of Data Mining and Bioinformatics
  • N.A Yan +4
  • Research Article
  • Citations114

IL28B single nucleotide polymorphisms in the treatment of hepatitis C

  • Mar 25, 2011
  • Journal of Hepatology
  • Christian M Lange +1
  • Research Article

Abstract 868: Replication of melanoma GWAS hits and exploration of pleiotropic effects of cancer GWAS hits with melanoma risk in the PAGE study

  • Apr 15, 2011
  • Cancer Research
  • Jonathan D Kocarnik +6
  • PDF
  • Book Chapter
  • Citations2

SNPpattern: A Genetic Tool to Derive Haplotype Blocks and Measure Genomic Diversity in Populations Using SNP Genotypes

  • Nov 02, 2011
  • Stephen J. +1
  • Research Article
  • Citations65

Detection of quantitative trait loci in Bos indicus and Bos taurus cattle using genome-wide association studies

  • Oct 29, 2013
  • Genetics, Selection, Evolution : GSE
  • Sunduimijid Bolormaa +8
  • Research Article
  • Citations3

Breast Cancer Risk Estimation Using the OncoVue® Model Compared to Combined GWAS Single Nucleotide Polymorphisms.

  • Dec 15, 2009
  • Cancer Research
  • E Jupe +3
  • Research Article

Abstract 2764: Cosegregating variants in chronic lymphocytic leukemia (CLL) families that are located in loci discovered by genome wide association studies (GWAS)

  • Aug 01, 2015
  • Cancer Research
  • Sara Beiggi +10
  • Research Article
  • Citations1

Abstract PL01-02: The heritable component of cancer: Insights from genome-wide association studies and beyond

  • Apr 15, 2011
  • Cancer Research
  • Stephen J Chanock
  • Research Article

Abstract 5224: Genome-wide interaction study of smoking and bladder cancer risk: results from the COBLAnCE cohort

  • Apr 04, 2023
  • Cancer Research
  • Maryam Karimi +12
  • Research Article
  • Citations77

Further improvements to linear mixed models for genome-wide association studies.

  • Nov 12, 2014
  • Scientific Reports
  • Christian Widmer +7
  • Research Article
  • Citations210

SNPHarvester: a filtering-based approach for detecting epistatic interactions in genome-wide association studies

  • Dec 19, 2008
  • Bioinformatics
  • Can Yang +5
  • PDF
  • Research Article
  • Citations4

Genome-wide Association Study (GWAS) and Its Application for Improving the Genomic Estimated Breeding Values (GEBV) of the Berkshire Pork Quality Traits

  • Sep 03, 2015
  • Asian-Australasian Journal of Animal Sciences
  • Young-Sup Lee +6
  • Research Article
  • Citations4

Gene-guided therapy for catheter-ablation of atrial fibrillation: are we there yet?

  • Dec 11, 2015
  • Journal of Interventional Cardiac Electrophysiology
  • Henry Huang +1
  • Research Article

Abstract 236: Identification of novel cancer target genes by combining data from the cancer genome-wide association studies (GWAS), regulatory DNA elements and The Cancer Genome Atlas (TCGA)

  • Jul 01, 2018
  • Cancer Research
  • Diptee A Kulkarni +7
  • Abstract
  • Citations1

Combined Donor and Recipient Non-HLA Genotypes Show Evidence of Genome Wide Association with Transplant Related Mortality (TRM) after HLA-Matched Unrelated Donor Blood and Marrow Transplantation (URD-BMT) (DISCOVeRY-BMT study)

  • Dec 03, 2015
  • Blood
  • Lara E Sucheston‐Campbell +19
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.