• Home
  • Search
  • A Novel Density-based Technique for Outlier Detection of High Dimensional Data Utilizing Full Feature Space
  • Cite Icon6
  • https://doi.org/10.5755/j01.itc.50.1.25588Copy DOI Icon

A Novel Density-based Technique for Outlier Detection of High Dimensional Data Utilizing Full Feature Space

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Recently, anomaly detection has acquired a realistic response from data mining scientists as a graph of its reputation has increased smoothly in various practical domains like product marketing, fraud detection, medical diagnosis, fault detection and so many other fields. High dimensional data subjected to outlier detection poses exceptional challenges for data mining experts and it is because of natural problems of the curse of dimensionality and resemblance of distant and adjoining points. Traditional algorithms and techniques were experimented on full feature space regarding outlier detection. Customary methodologies concentrate largely on low dimensional data and hence show ineffectiveness while discovering anomalies in a data set comprised of a high number of dimensions. It becomes a very difficult and tiresome job to dig out anomalies present in high dimensional data set when all subsets of projections need to be explored. All data points in high dimensional data behave like similar observations because of its intrinsic feature i.e., the distance between observations approaches to zero as the number of dimensions extends towards infinity. This research work proposes a novel technique that explores deviation among all data points and embeds its findings inside well established density-based techniques. This is a state of art technique as it gives a new breadth of research towards resolving inherent problems of high dimensional data where outliers reside within clusters having different densities. A high dimensional dataset from UCI Machine Learning Repository is chosen to test the proposed technique and then its results are compared with that of density-based techniques to evaluate its efficiency.

Loading PDF

Similar Papers
  • Research Article
  • Citations2

Differential privacy protection algorithm for large data sources based on normalized information entropy Bayesian network

  • Aug 01, 2024
  • Journal of Physics: Conference Series
  • Guangyuan Ni +1
  • Conference Article
  • Citations11

Outlier detection via sampling ensemble

  • Dec 01, 2016
  • Hongfu Liu +3
  • Book Chapter
  • Citations2

Clustering of High Dimensional Handwritten Data by an Improved Hypergraph Partition Method

  • Jan 01, 2017
  • Tian Wang +2
  • Conference Article
  • Citations10

SSDP+: A Diverse and More Informative Subgroup Discovery Approach for High Dimensional Data

  • Jul 01, 2018
  • Tarcsio Lucas +2
  • Conference Article
  • Citations73

OutRank: ranking outliers in high dimensional data

  • Apr 01, 2008
  • Emmanuel Muller +3
  • Conference Article
  • Citations2

A validity index for outlier detection

  • Nov 01, 2010
  • Noha A Yousri
  • Research Article
  • Citations1

Design of feature selection algorithm for high-dimensional network data based on supervised discriminant projection.

  • Jun 26, 2023
  • PeerJ. Computer science
  • Zongfu Zhang +4
  • Conference Article
  • Citations2

Clustering High-Dimensional Stock Data using Data Mining Approach

  • Jul 01, 2019
  • Dhea Indriyanti +1
  • Research Article
  • Citations75

DBFS: An effective Density Based Feature Selection scheme for small sample size and high dimensional imbalanced data sets

  • Aug 17, 2012
  • Data & Knowledge Engineering
  • Mina Alibeigi +2
  • Research Article
  • Citations9

High-dimensional sparse vine copula regression with application to genomic prediction.

  • Jan 29, 2024
  • Biometrics
  • Özge Sahin +1
  • Research Article
  • Citations25

Anomaly Detection Using Proximity Graph and PageRank Algorithm

  • Aug 01, 2012
  • IEEE Transactions on Information Forensics and Security
  • Zhe Yao +2
  • Research Article

A Permutation Test for Multiple Correlation Coefficient in High Dimensional Normal Data

  • Sep 01, 2023
  • Journal of Statistical Sciences
  • Dariush Najarzadeh
  • Research Article
  • Citations5

A LoOP based outlier detection method for high dimensional fuzzy data set

  • Jan 13, 2017
  • Journal of Intelligent & Fuzzy Systems
  • Alireza Fakharzadeh Jahromi +1
  • Research Article

Intelligent Recommender System for High Dimensional Transaction Data Set with Complex Relationships among the Variables

  • May 30, 2016
  • Indian Journal of Science and Technology
  • Woo Kim Jun +1
  • Research Article
  • Citations149

The Role of Hubness in Clustering High-Dimensional Data

  • Mar 01, 2014
  • IEEE Transactions on Knowledge and Data Engineering
  • Nenad Tomasev +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.