• Home
  • Search
  • Measuring similarities between gene expression profiles through new data transformations.
  • Cite Icon30
  • https://doi.org/10.1186/1471-2105-8-29Copy DOI Icon

Measuring similarities between gene expression profiles through new data transformations.

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

BackgroundClustering methods are widely used on gene expression data to categorize genes with similar expression profiles. Finding an appropriate (dis)similarity measure is critical to the analysis. In our study, we developed a new measure for clustering the genes when the key factor is the shape of the profile, and when the expression magnitude should also be accounted for in determining the gene relationship. This is achieved by modeling the shape and magnitude parameters separately in a gene expression profile, and then using the estimated shape and magnitude parameters to define a measure in a new feature space.ResultsWe explored several different transformation schemes to construct the feature spaces that include a space whose features are determined by the mutual differences of the original expression components, a space derived from a parametric covariance matrix, and the principal component space in traditional PCA analysis. The former two are the newly proposed and the latter is explored for comparison purposes. The new measures we defined in these feature spaces were employed in a K-means clustering procedure to perform analyses. Applying these algorithms to a simulation dataset, a developing mouse retina SAGE dataset, a small yeast sporulation cDNA dataset, and a maize root affymetrix microarray dataset, we found from the results that the algorithm associated with the first feature space, named TransChisq, showed clear advantages over other methods.ConclusionThe proposed TransChisq is very promising in capturing meaningful gene expression clusters. This study also demonstrates the importance of data transformations in defining an efficient distance measure. Our method should provide new insights in analyzing gene expression data. The clustering algorithms are available upon request.

Loading PDF

Similar Papers
  • Research Article

Abstract P1-04-02: Immune milieu associated with PD-L1 status in TNBC is dependent on time of biomarker assessment and treatment received: A secondary analysis of the NeoTRIPaPDL1 trial

  • Feb 15, 2022
  • Cancer Research
  • Maurizio Callari +19
  • PDF
  • Research Article
  • Citations1

A Seriation Approach for Visualization-Driven Discovery of Co-Expression Patterns in Serial Analysis of Gene Expression (SAGE) Data

  • Sep 12, 2008
  • PLoS ONE
  • Olena Morozova +4
  • Conference Article
  • Citations12

A new cluster validity measure for bioinformatics relational datasets

  • Jun 01, 2008
  • Mihail Popescu +4
  • Front Matter
  • Citations1

The genome is now accessible to the endoscopist

  • Jun 27, 2006
  • Gastrointestinal Endoscopy
  • Anson W Lowe
  • Abstract

Detailed Genome-Wide DNA-Mapping of CD34+ Cells Purified from Patients with MDS Using High-Resolution SNP Arrays Identifies Significant Regions of Genomic Alterations.

  • Nov 16, 2006
  • Blood
  • Daniel Nowak +7
  • Research Article
  • Citations30

Complex patterns of gene expression in human T cells during in vivo aging

  • Jan 01, 2010
  • Molecular BioSystems
  • Daniel Remondini +9
  • PDF
  • Research Article
  • Citations15

Robust Detection and Genotyping of Single Feature Polymorphisms from Gene Expression Data

  • Mar 13, 2009
  • PLoS Computational Biology
  • Minghui Wang +8
  • Research Article
  • Citations927

Model-based clustering and data transformations for gene expression data.

  • Oct 01, 2001
  • Bioinformatics
  • K Y Yeung +4
  • Abstract

Combined Clinical and Gene Expression Score Identifies Follicular Lymphoma Patients with High Risk of Transformation in the Rituximab Era

  • Jun 25, 2021
  • Blood
  • Chloe Beate Steen +13
  • PDF
  • Research Article
  • Citations9

Identification of candidate cancer drivers by integrative Epi-DNA and Gene Expression (iEDGE) data analysis

  • Nov 15, 2019
  • Scientific Reports
  • Amy Li +4
  • PDF
  • Research Article
  • Citations14

Random Subspace Aggregation for Cancer Prediction with Gene Expression Profiles

  • Jan 01, 2016
  • BioMed Research International
  • Liying Yang +4
  • PDF
  • Research Article
  • Citations83

In Situ Proteomic Analysis of Human Breast Cancer Epithelial Cells Using Laser Capture Microdissection: Annotation by Protein Set Enrichment Analysis and Gene Ontology

  • Nov 01, 2010
  • Molecular & Cellular Proteomics
  • Sangwon Cha +6
  • Conference Article
  • Citations1

Microarray Meta-Miner (MMM): An integrated method and a web tool to identify genes with similar expression profile

  • Nov 01, 2010
  • Vijayaraj Nagarajan +4
  • Abstract

DNA microarray analysis in down syndrome

  • Sep 01, 2002
  • Fertility and Sterility
  • I.H Jung +5
  • PDF
  • Research Article
  • Citations70

Importance of Correlation between Gene Expression Levels: Application to the Type I Interferon Signature in Rheumatoid Arthritis

  • Oct 17, 2011
  • PLoS ONE
  • Frédéric Reynier +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.