• Home
  • Search
  • GRID distribution supports clustering validation of large mixed microarray data sets
  • Cite Icon6
  • https://doi.org/10.14806/ej.17.1.205Copy DOI Icon

GRID distribution supports clustering validation of large mixed microarray data sets

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Microarray data are a rich source of information, containing the collected expression values of thousands of genes for well defined states of a cell or tissue. Vast amounts of data (thousands of arrays) are publicly available and ready for analysis, e.g. to scrutinise correlations between genes at the level of gene expression. The large variety of arrays available makes it possible to combine different independent experiments to extract new knowledge. Starting with a large set of data, relevant information can be isolated for further analysis. To extract the required information from data sets of such size and complexity requires an appropriate and powerful analysis method. In this study, we chose to use an unsupervised hierarchical clustering algorithm, Chaotic Map Clustering (CMC), in a coupled two-way approach to analyse such data. However, the clustering approach is intrinsically difficult, both in terms of the unknown structure of the data and interpretation of the clustering results. It is therefore critical to evaluate the quality of any unsupervised procedure for such a complex set of data and to validate the clustering results, separating those clusters that are due simply to noise or statistical fluctuations. We used a resampling method to perform this validation. The resampling procedure applies the clustering algorithm to a large number of random sub-samples of the original data matrix and, consequently, the whole process becomes computationally intensive and time consuming. Using Grid technology, we show that we can drastically speed up this process by distributing the clustering of each matrix to a separate worker node, and thus retrieve resampling results within a few hours instead of several days. Further, we offer an online service to cluster large microarray data sets and conduct the subsequent validation described in this paper.

Similar Papers
  • Research Article
  • Citations4

Gene microarray data analysis using parallel point-symmetry-based clustering.

  • Aug 25, 2010
  • International journal of data mining and bioinformatics
  • Anasua Sarkar +1
  • Research Article
  • Citations381

Control Genes and Variability: Absence of Ubiquitous Reference Transcripts in Diverse Mammalian Expression Studies

  • Feb 01, 2002
  • Genome Research
  • Peter D Lee +3
  • Research Article
  • Citations9

NORMALITY OF GENE EXPRESSION REVISITED

  • Mar 01, 2007
  • Journal of Biological Systems
  • Linlin Chen +2
  • PDF
  • Research Article
  • Citations14

A transcriptome-based classifier to determine molecular subtypes in medulloblastoma

  • Oct 29, 2020
  • PLoS Computational Biology
  • Komal S Rathi +8
  • Research Article
  • Citations235

Minerva and minepy: a C engine for the MINE suite and its R, Python and MATLAB wrappers

  • Dec 14, 2012
  • Bioinformatics
  • Davide Albanese +5
  • PDF
  • Research Article
  • Citations5

Association Analysis Techniques for Discovering Functional Modules from Microarray Data

  • Aug 13, 2008
  • Nature Precedings
  • Gaurav Pandey +3
  • Research Article
  • Citations61

Genome-wide identification and analysis of Catharanthus roseus RLK1-like kinases in rice.

  • Nov 16, 2014
  • Planta
  • Quynh-Nga Nguyen +5
  • Research Article
  • Citations128

Human Breast Tumor Cells Induce Self-Tolerance Mechanisms to Avoid NKG2D-Mediated and DNAM-Mediated NK Cell Recognition

  • Oct 30, 2011
  • Cancer Research
  • Emilie Mamessier +9
  • Conference Article

Effectiveness of Principal Component Analysis in Functional Mapping of Gene Expression Profiles

  • May 01, 2023
  • Rajashree Sahoo +1
  • PDF
  • Research Article
  • Citations211

Biclustering of gene expression data by Non-smooth Non-negative Matrix Factorization.

  • Feb 17, 2006
  • BMC Bioinformatics
  • Pedro Carmona-Saez +4
  • Research Article
  • Citations108

The derivation of diagnostic markers of chronic myeloid leukemia progression from microarray data

  • Oct 08, 2009
  • Blood
  • Vivian G Oehler +5
  • Research Article
  • Citations68

Interpretation of ANOVA models for microarray data using PCA

  • Nov 14, 2006
  • Bioinformatics
  • J R De Haan +5
  • Research Article
  • Citations33

Fuzzy Expert System based on a Novel Hybrid Stem Cell (HSC) Algorithm for Classification of Micro Array Data.

  • Feb 21, 2018
  • Journal of Medical Systems
  • S Arul Antran Vijay +1
  • Book Chapter
  • Citations1

Chapter 5 - The Purpose of Data Analysis Is to Enable Data Reanalysis

  • Jan 01, 2015
  • Repurposing Legacy Data
  • Jules J Berman
  • Preprint Article

Data from Human Breast Tumor Cells Induce Self-Tolerance Mechanisms to Avoid NKG2D-Mediated and DNAM-Mediated NK Cell Recognition

  • Mar 30, 2023
  • Emilie Mamessier +9
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.