A study on missing values imputation using K-Harmonic means algorithm: Mixed datasets

Data cleaning is one step in the preprocessing which in the process often found missing values in the dataset. Missing values is the condition of the absence of data items on a subject. A quick step that can be taken to handle missing values is to remove data containing missing values, but this can...

Full description

Saved in:
Bibliographic Details
Published inAIP conference proceedings Vol. 2202; no. 1
Main Authors Anwar, Taufik, Siswantining, Titin, Sarwinda, Devvi, Soemartojo, Saskya Mary, Bustamam, Alhadi
Format Journal Article Conference Proceeding
LanguageEnglish
Published Melville American Institute of Physics 27.12.2019
Subjects
Online AccessGet full text
ISSN0094-243X
1935-0465
1551-7616
1551-7616
DOI10.1063/1.5141651

Cover

More Information
Summary:Data cleaning is one step in the preprocessing which in the process often found missing values in the dataset. Missing values is the condition of the absence of data items on a subject. A quick step that can be taken to handle missing values is to remove data containing missing values, but this can reducing information in the data. Another way to handle missing values is by using imputation with mean, median, or mode, and several methods of imputation such as regression, likelihood, and the clustering approach. Imputation with the clustering approach is the focus of this study, where we used the K-Harmonic Means which has been adjusted to handle mixed data. K-Harmonic Means is an extension of K-Means by reducing random centroid initialization sensitivity problems. Imputation of the missing values is carried out by distributing missing values observation to the cluster and replacing the missing values with the information on the same centroid cluster. The results of the simulation were evaluated using the root mean square error and the accuracy values of each imputation value for numerical and categorical data respectively.
Bibliography:ObjectType-Conference Proceeding-1
SourceType-Conference Papers & Proceedings-1
content type line 21
ISSN:0094-243X
1935-0465
1551-7616
1551-7616
DOI:10.1063/1.5141651