• Home
  • Search
  • Feature Selection Techniques and Classification Accuracy of Supervised Machine Learning in Text Mining
  • Cite Icon1
  • https://doi.org/10.7176/jiea/9-3-06Copy DOI Icon

Feature Selection Techniques and Classification Accuracy of Supervised Machine Learning in Text Mining

Show More
  • Abstract
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Text mining is a special case of data mining which explore unstructured or semi-structured text documents, to establish valuable patterns and rules that indicate trends and significant features about specific topics. Text mining has been in pattern recognition, predictive studies, sentiment analysis and statistical theories in many areas of research, medicine, financial analysis, social life analysis, and business intelligence. Text mining uses concept of natural language processing and machine learning. Machine learning algorithms have been used and reported to give great results, but their performance of machine learning algorithms is affected by factors such as dataset domain, number of classes, length of the corpus, and feature selection techniques used. Redundant attribute affects the performance of the classification algorithm, but this can be reduced by using different feature selection techniques and dimensionality reduction techniques. Feature selection is a data preprocessing step that chooses a subset of input variable while eliminating features with little or no predictive information. Feature selection techniques are Information gain, Term Frequency, Term Frequency-Inverse document frequency, Mutual Information, and Chi-Square, which can use a filters, wrappers, or embedded approaches. To get the most value from machine learning, pairing the best algorithms with the right tools and processes is necessary. Little research has been done on the effect of feature selection techniques on classification accuracy for pairing of these algorithms with the best feature selection techniques for optimal results. In this research, a text classification experiment was conducted using incident management dataset, where incidents were classified into their resolver groups. Support vector machine (SVM), K-Nearest Neighbors (KNN), Naïve Bayes (NB) and Decision tree (DT) machine learning algorithms were examined. Filtering approach was used on the feature selection techniques, with different ranking indices applied for optimal feature set and classification accuracy results analyzed. The classification accuracy results obtained using TF were, 88% for SVM, 70% for NB, 79% for Decision tree, and KNN had 55%, while Boolean registered 90%, 83%, 82% and 75%, for SVM, NB, DT, and KNN respectively. TF-IDF, had 91%, 83%, 76%, and 56% for SVM, NB, DT, and KNN respectively. The results showed that algorithm performance is affected by feature selection technique applied. SVM performed best, followed by DT, KNN and finally NB. In conclusion, presence of noisy data leads to poor learning performance and increases the computational time. The classifiers performed differently depending on the feature selection technique applied. For optimal results, the classifier that performed best together with the feature selection technique with the best feature subset should be applied for all types of data for accurate classification performance. Keywords: Text Classification, Supervised Machine Learning, Feature Selection DOI : 10.7176/JIEA/9-3-06 Publication date :May 31 st 2019

Loading PDF

Similar Papers
  • Research Article
  • Citations14

Prediction and feature selection of low birth weight using machine learning algorithms

  • Oct 12, 2024
  • Journal of Health, Population and Nutrition
  • Tasneem Binte Reza +1
  • Conference Article
  • Citations1

Utilizing Various Machine-Learning Techniques in Breast Cancer Detection

  • Jan 01, 2024
  • Skala Hassan Hussen +1
  • Conference Article

Effects of Expression Recognition with Machine and Deep Learning Algorithms on Psychotherapy

  • Dec 06, 2025
  • Gulay Cicek +2
  • Research Article
  • Citations41

Performances of machine learning algorithms for individual thermal comfort prediction based on data from professional and practical settings

  • Sep 17, 2022
  • Journal of Building Engineering
  • Changyong Yu +6
  • Research Article

GlycoDetect a Diabetic Prediction Model using ML

  • Oct 05, 2024
  • INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Archana Nikose +4
  • Book Chapter
  • Citations3

Performance Stagnation of Meteorological Data of Kashmir

  • Sep 23, 2022
  • Sameer Kaul +3
  • Book Chapter
  • Citations5

Machine Learning-Based Depression Detection

  • Oct 14, 2022
  • Saikat Biswas +4
  • Conference Article
  • Citations10

Applying Feature Selection to Short Time Wavelet Transformed Vibration Data for Reliability Analysis of an Ocean Turbine

  • Dec 01, 2012
  • Janell Duhaney +2
  • Research Article
  • Citations2

The Use of Radiomics Data Obtained from ADC Map of Lumbar MRI and Machine Learning in Diagnosis of Osteoporosis

  • Jul 30, 2024
  • Iranian Journal of Radiology
  • Fatih Erdem +4
  • Conference Article
  • Citations60

A Comparative Study for Breast Cancer Prediction using Machine Learning and Feature Selection

  • May 01, 2019
  • R Dhanya +4
  • Conference Article
  • Citations3

Risk Analysis in Electronic Payments and Settlement System Using Dimensionality Reduction Techniques

  • Jan 01, 2018
  • B Emil Richard Singh +1
  • Research Article

Performance of Machine Learning Algorithms for Credit Risk Prediction with Feature Selection

  • Apr 05, 2025
  • Statistics, Optimization & Information Computing
  • Muhammad M Seliem +5
  • PDF
  • Research Article
  • Citations414

Benchmarking of Machine Learning for Anomaly Based Intrusion Detection Systems in the CICIDS2017 Dataset

  • Jan 01, 2021
  • IEEE Access
  • Ziadoon Kamil Maseer +4
  • Supplementary Content
  • Citations231

Review of Machine Learning Algorithms for Diagnosing Mental Illness

  • Apr 01, 2019
  • Psychiatry Investigation
  • Gyeongcheol Cho +4
  • Conference Article
  • Citations9

Hybrid Model of Correlation Based Filter Feature Selection and Machine Learning Classifiers Applied on Smart Meter Data Set

  • May 01, 2019
  • Janvier Omar Sinayobye +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.