Theme-Logo
  • Login
  • Home
  • Course
  • Publication
  • Theses
  • Reports
  • Published books
  • Workshops / Conferences
  • Supervised PhD
  • Supervised MSc
  • Supervised projects
  • Education
  • Language skills
  • Positions
  • Memberships and awards
  • Committees
  • Experience
  • Scientific activites
  • In links
  • Outgoinglinks
  • News
  • Gallery
publication name The Effect of Different Dimensionality Reduction Techniques on Machine Learning Overfitting Problem
Authors Mustafa Abdul Salam and Ahmad Taher Azar and Mustafa Samy Elgendy and Khaled Mohamed Fouad
year 2021
keywords Dimensionality reduction; feature subset selection; rough set; overfitting; underfitting; machine learning
journal International Journal of Advanced Computer Science and Applications
volume 12
issue 4
pages 17
publisher The Science and Information Organization
Local/International International
Paper Link http://dx.doi.org/10.14569/IJACSA.2021.0120480
Full paper download
Supplementary materials Not Available
Abstract

In most conditions, it is a problematic mission for a machine-learning model with a data record, which has various attributes, to be trained. There is always a proportional relationship between the increase of model features and the arrival to the overfitting of the susceptible model. That observation occurred since not all the characteristics are always important. For example, some features could only cause the data to be noisier. Dimensionality reduction techniques are used to overcome this matter. This paper presents a detailed comparative study of nine dimensionality reduction methods. These methods are missing-values ratio, low variance filter, high-correlation filter, random forest, principal component analysis, linear discriminant analysis, backward feature elimination, forward feature construction, and rough set theory. The effects of used methods on both training and testing performance were compared with two different datasets and applied to three different models. These models are, Artificial Neural Network (ANN), Support Vector Machine (SVM) and Random Forest classifier (RFC). The results proved that the RFC model was able to achieve the dimensionality reduction via limiting the overfitting crisis. The introduced RFC model showed a general progress in both accuracy and efficiency against compared approaches. The results revealed that dimensionality reduction could minimize the overfitting process while holding the performance so near to or better than the original one

Benha University © 2023 Designed and developed by portal team - Benha University