• https://doi.org/10.14264/uql.2020.783Copy DOI Icon

Clustering with mixed variables

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

A common and old problem in statistics is the separation of a heterogeneous population into more homogeneous subpopulations. A wide variety of approaches and techniques for tackling this problem now exist. The finite mixture model is one such technique used to analyse a given data set and create natural groups. The main advantage of mixture models is that they offer a very flexible and relatively easy method of fitting to a data set.The main focus of this thesis will be the clustering of data with the aid of mixed variable mixture models. There has been a lot of research and many papers written with regard to the analysis of continuous variables. This thesis, however, will be concerned with the techniques that will be able to analyse data from continuous to mixed variables, a combination of continuous and discrete variables.The expectation-maximization (EM) algorithm of Dempster et al. (1977) will be used in this thesis for the clustering of mixed variable data. As demonstrated effectively in numerous publications, this algorithm solves estimation of the model parameters iteratively under an advantageous property of monotonic convergence. An extension of the EM algorithm, the ECM algorithm introduced by Meng and Rubin (1993), will also be considered in this thesis for itsrole within more complicated mixed mixture models.On the application of mixed mixture models, this thesis will use a joint distribution specified by the conditional distribution of the continuous variables, given the values of the discrete variables times the marginal distribution of the latter. There are three models - naive, logistic and multinomial - considered for the discrete variables and two models - independent and location - considered for the continuous variables. These models are combined to make a total of sixmodels for the analysis of mixed variable data.The program EMM, its design and development will be discussed in this thesis. Many real mixed data sets and some simulations will also be analysed to investigate how these proposed methods will compare with existing methods. Some of these data sets will be transformed so as to create a suitable input for the program and hence provide a comparable output so that comparisons can be made.The program EMM is written in the FORTRAN computer language. This means that it is well organised with each segment arranged in separate sub-routines. This allows for easy modification to any part of the program or alternatively the addition of a new subroutine so as to increase the capabilities of the model.

Similar Papers
  • Research Article
  • Citations17

Online variational inference on finite multivariate Beta mixture models for medical applications

  • Mar 04, 2021
  • IET Image Processing
  • Narges Manouchehri +2
  • Single Book
  • Citations1654

Finite Mixture and Markov Switching Models

  • Jan 01, 2006
  • Sylvia Frühwirth‐Schnatter
  • Research Article
  • Citations1

A Mean Field Games model for finite mixtures of Bernoulli and categorical distributions

  • Dec 14, 2020
  • Journal of Dynamics and Games
  • Laura Aquilanti +3
  • Book Chapter
  • Citations11

Probabilistic Models for Clustering

  • Sep 03, 2018
  • Hongbo Deng +1
  • Research Article
  • Citations8

Using general regression with local tuning for learning mixture models from incomplete data sets

  • Nov 02, 2010
  • Egyptian Informatics Journal
  • Ahmed R Abas
  • Research Article

Variational learning for finite Beta-Liouville mixture models

  • Apr 01, 2014
  • The Journal of China Universities of Posts and Telecommunications
  • Yu-Ping Lai +4
  • Conference Article
  • Citations9

An algorithm for estimating number of components of Gaussian mixture model based on penalized distance

  • Jun 01, 2008
  • Daming Zhang +2
  • Research Article
  • Citations74

Examining the effect of initialization strategies on the performance of Gaussian mixture modeling.

  • Dec 31, 2015
  • Behavior Research Methods
  • Emilie Shireman +2
  • Research Article

How have high-impact scientific studies designing their experiments on mixed data clustering? A systematic map to guide better choices

  • Jun 08, 2021
  • Machine Learning with Applications
  • Nádia Junqueira Martarelli +1
  • Research Article
  • Citations3

Simrec: a similarity measure recommendation system for mixed data clustering algorithms

  • Feb 22, 2025
  • Journal of Big Data
  • Abdoulaye Diop +5
  • Research Article

Finite Mixture Model: A Comparison of Maximum Likelihood Estimation and Bayesian Analysis

  • Oct 08, 2021
  • Global Conference on Business and Social Sciences Proceeding
  • Seuk Yen Phoong +1
  • Research Article
  • Citations1

Pre-analysis of Multi-batch Bioprocesses Data with Finite Mixture Models in the Reduced Feature Subspace

  • Jan 01, 2010
  • IFAC Proceedings Volumes
  • Weilu Lin +2
  • Book Chapter
  • Citations6

Unsupervised Classification of Mixed Data Type of Attributes Using Genetic Algorithm (Numeric, Categorical, Ordinal, Binary, Ratio-Scaled)

  • Jan 01, 2014
  • Rohit Rastogi +4
  • Research Article
  • Citations2

Global estimation of finite mixture and misclassification models with an application to multiple equilibria

  • Apr 19, 2021
  • Econometric Reviews
  • Yingyao Hu +1
  • Research Article
  • Citations7

Using incremental general regression neural network for learning mixture models from incomplete data

  • Aug 25, 2011
  • Egyptian Informatics Journal
  • Ahmed R Abas
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.