Published 2019 | Version v1
Book

Evaluating Auto-encoder and Principal Component Analysis for Feature Engineering in Electronic Health Records

Description

Feature engineering is an important mechanism where we transform and represent the high-dimensional data into a lower-dimensional space. These representations can then be used to efficiently train machine-learning models. Auto-encoders are widely used in research for unsupervised feature learning. However, the application of auto-encoders for electronic health records (EHRs) containing features with binary values (binary-valued features) has been less studied. The primary objective of this research was to compare an auto-encoder with principal component analysis (PCA), a popular feature engineering technique, for feature selection in different (US and Indian) EHR datasets containing binary-valued features. The US dataset contained thousands of binary-valued features, and the Indian dataset contained nineteen binary-valued features. Results revealed that feature selection by the auto-encoder followed by different classification algorithms gave the highest accuracy on both the datasets compared to feature selection by PCA. We highlight the implications of using auto-encoders for learning features in EHR datasets.

Part of:
ITISE 2019. Proceedings of papers. Vol 2

Additional details

Publishing Information

Publisher
Universdad de Granada
Imprint Place
Granada (Spain)
Imprint Title
ITISE 2019. Proceedings of papers. Vol 2
Imprint Pagination
675 p.
Journal Page Range
12 p.

Conference

Title
International Conference on Time Series and Forecasting
Acronym
ITISE 2019
Dates
25-27 Sep 2019
Place
Granada (Spain)

INIS

Country of Publication
Spain
Country of Input or Organization
Spain
INIS RN
52049002
Subject category
S97: MATHEMATICAL METHODS AND COMPUTING;
Resource subtype / Literary indicator
Conference
Descriptors DEI
EVALUATION; FORECASTING; MATHEMATICS; NEURAL NETWORKS; STATISTICS
Descriptors DEC
MATHEMATICS

Optional Information