Feature selection for classification tasks: Expert knowledge or traditional methods?

Recently, available data has increased explosively in both number of samples and dimensionality. The huge number of high dimensional data generates the presence of noisy, redundant and irrelevant dimensions. Such dimensions can increase the time and computational cost in the learning process and eve...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:Journal of intelligent & fuzzy systems 2018-01, Vol.34 (5), p.2825-2835
Hauptverfasser: Corrales, David Camilo, Lasso, Emmanuel, Ledezma, Agapito, Corrales, Juan Carlos
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Beschreibung
Zusammenfassung:Recently, available data has increased explosively in both number of samples and dimensionality. The huge number of high dimensional data generates the presence of noisy, redundant and irrelevant dimensions. Such dimensions can increase the time and computational cost in the learning process and even degenerate the performance of learning tasks. One of the ways to reduce dimensionality is by Feature Selection (FS). The aim of this paper is study the feature selection based on expert knowledge and traditional methods (filter, wrapper and embedded) and analyze their performance in classification tasks. Three datasets related to cancer domain in humans were used for feature selection: Breast Cancer (BC), Primary Tumor (PT) and Central Nervous System (CNS). C4.5, K-Nearest Neighbors, Support Vector Machine and Multi Layer Perceptron were trained with the best subset of features for each cancer dataset. The subset of features selected by the wrapper method presents the best average accuracy in the datasets BC and PT, while the subset of features selected by the embedded method reaches the highest average accuracy in the CNS dataset.
ISSN:1064-1246
1875-8967
DOI:10.3233/JIFS-169470