• Home
  • Search
  • Resampling Techniques in Cluster Analysis: Is Subsampling Better Than Bootstrapping?
  • Cite Icon3
  • https://doi.org/10.1007/978-3-662-44983-7_10Copy DOI Icon

Resampling Techniques in Cluster Analysis: Is Subsampling Better Than Bootstrapping?

  • Jan 1, 2015
  • Hans-Joachim Mucha +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Abstract In the case of two small toy data sets, we found out that subsampling has a much weaker behavior in the finding of the true number of clusters K than bootstrapping (Mucha and Bartel, Soft bootstrapping in cluster analysis and its comparison with other resampling methods. In: M. Spiliopoulou, L. Schmidt-Thieme, R. Janning (eds.) Data analysis, machine learning and knowledge discovery. Springer, Cham, 2014). In contradiction, Möller and Dörte (Intell Data Anal 10:139–162, 2006) pointed out that “subsampling … clearly outperformed the bootstrapping technique in the detection of correct clustering consensus results.” Obviously, there is a need for further investigations. Therefore here we compare these two resampling techniques based on real and artificial data sets by means of different indices: ARI or Jaccard. We consider hierarchical cluster analysis methods because they find all partitions into K = 2, 3, … clusters in one run only, and, moreover, these results are (usually) unique (Spaeth, Cluster analysis algorithms for data reduction and classification of objects. Ellis Horwood, Chichester, 1982). The methods are tested on two synthetic data sets and two real data sets. Obviously, bootstrapping is better than subsampling in finding the true number of clusters.

Similar Papers
  • Research Article

Evaluating the Sampling Performance of Exploratory and Cross-Validated DETECT Procedure with Imperfect Models

  • Nov 02, 2015
  • Multivariate Behavioral Research
  • Cengiz Zopluoglu
  • Research Article
  • Citations142

Extensive Copy-Number Variation of the Human Olfactory Receptor Gene Family

  • Jul 31, 2008
  • American journal of human genetics
  • Janet M Young +5
  • Research Article
  • Citations53

Validation of Synthetic U.S. Electric Power Distribution System Data Sets

  • Sep 01, 2020
  • IEEE Transactions on Smart Grid
  • Venkat Krishnan +8
  • Book Chapter
  • Citations1

Content Extraction of Chinese Archive Images via Synthetic and Real Data

  • Jan 01, 2022
  • Xin Jin +4
  • Research Article
  • Citations58

Temporal Matrix Factorization for Tracking Concept Drift in Individual User Preferences

  • Mar 01, 2018
  • IEEE Transactions on Computational Social Systems
  • Yung-Yin Lo +3
  • Research Article
  • Citations927

Model-based clustering and data transformations for gene expression data.

  • Oct 01, 2001
  • Bioinformatics
  • K Y Yeung +4
  • Research Article
  • Citations8

Effects of errors and biases on the scaling of earthquake spatial pattern: application to the 2004 Sumatra–Andaman sequence

  • Dec 07, 2013
  • Natural Hazards
  • Simanchal Padhy +5
  • PDF
  • Research Article
  • Citations2

A New Method of Generating Index Label for Dynamic XML Data

  • Mar 01, 2011
  • Journal of Computer Science
  • Paramasivam
  • Research Article
  • Citations8

SPREd: a simulation-supervised neural network tool for gene regulatory network reconstruction.

  • Jan 05, 2024
  • Bioinformatics Advances
  • Zijun Wu +1
  • Research Article

FCM-FCS: Hybridization of Fractional Cuckoo Search with FCM for High Dimensional Data Clustering Process

  • Jan 01, 2013
  • International Review on Computers and Software
  • G Victo Sudha George +1
  • Research Article
  • Citations1

The Comparison Between Gumbel and Exponentiated Gumbel Distributions and Their Applications in Hydrological Process

  • Sep 10, 2021
  • Advances in Machine Learning & Artificial Intelligence
  • Anita Abdollahi Nanvapisheh
  • Conference Article
  • Citations1

A validity index based on cluster symmetry

  • Oct 01, 2007
  • Sriparna Saha +1
  • Research Article
  • Citations27

A novel clustering algorithm based on the natural reverse nearest neighbor structure

  • Apr 18, 2019
  • Information Systems
  • Qi-Zhu Dai +5
  • Research Article
  • Citations24

Simultaneous cancer classification and gene selection with Bayesian nearest neighbor method: An integrated approach

  • Oct 18, 2008
  • Computational Statistics & Data Analysis
  • Sounak Chakraborty
  • Research Article
  • Citations41

Evaluating data mining procedures: techniques for generating artificial data sets

  • Jun 01, 1999
  • Information and Software Technology
  • P.D Scott +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.