Automatic missing value imputation for cleaning phase of diabetic’s readmission prediction model

Recently, the industry of healthcare started generating a large volume of datasets. If hospitals can employ the data, they could easily predict the outcomes and provide better treatments at early stages with low cost. Here, data analytics (DA) was used to make correct decisions through proper analys...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	International journal of electrical and computer engineering (Malacca, Malacca) Malacca), 2022-04, Vol.12 (2), p.2001
Hauptverfasser:	Mohd Zebaral Hoque, Jesmeen, Hossen, Jakir, Sayeed, Shohel, Mohammed Tawsif K., Chy, Ganesan, Jaya, Raja, J. Emerson
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Cost analysis Data sampling Datasets Decision analysis Decision trees Machine learning Prediction models Structured data Support vector machines
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Recently, the industry of healthcare started generating a large volume of datasets. If hospitals can employ the data, they could easily predict the outcomes and provide better treatments at early stages with low cost. Here, data analytics (DA) was used to make correct decisions through proper analysis and prediction. However, inappropriate data may lead to flawed analysis and thus yield unacceptable conclusions. Hence, transforming the improper data from the entire data set into useful data is essential. Machine learning (ML) technique was used to overcome the issues due to incomplete data. A new architecture, automatic missing value imputation (AMVI) was developed to predict missing values in the dataset, including data sampling and feature selection. Four prediction models (i.e., logistic regression, support vector machine (SVM), AdaBoost, and random forest algorithms) were selected from the well-known classification. The complete AMVI architecture performance was evaluated using a structured data set obtained from the UCI repository. Accuracy of around 90% was achieved. It was also confirmed from cross-validation that the trained ML model is suitable and not over-fitted. This trained model is developed based on the dataset, which is not dependent on a specific environment. It will train and obtain the outperformed model depending on the data available.
ISSN:	2088-8708 2722-2578 2088-8708
DOI:	10.11591/ijece.v12i2.pp2001-2013