Evaluating Auto-encoder and Principal Component Analysis for Feature Engineering in Electronic Health Records
Description
Feature engineering is an important mechanism where we transform and represent the high-dimensional data into a lower-dimensional space. These representations can then be used to efficiently train machine-learning models. Auto-encoders are widely used in research for unsupervised feature learning. However, the application of auto-encoders for electronic health records (EHRs) containing features with binary values (binary-valued features) has been less studied. The primary objective of this research was to compare an auto-encoder with principal component analysis (PCA), a popular feature engineering technique, for feature selection in different (US and Indian) EHR datasets containing binary-valued features. The US dataset contained thousands of binary-valued features, and the Indian dataset contained nineteen binary-valued features. Results revealed that feature selection by the auto-encoder followed by different classification algorithms gave the highest accuracy on both the datasets compared to feature selection by PCA. We highlight the implications of using auto-encoders for learning features in EHR datasets.
Additional details
Identifiers
Publishing Information
- Publisher
- Universdad de Granada
- Imprint Place
- Granada (Spain)
- Imprint Title
- ITISE 2019. Proceedings of papers. Vol 2
- Imprint Pagination
- 675 p.
- Journal Page Range
- 12 p.
Conference
- Title
- International Conference on Time Series and Forecasting
- Acronym
- ITISE 2019
- Dates
- 25-27 Sep 2019
- Place
- Granada (Spain)
INIS
- Country of Publication
- Spain
- Country of Input or Organization
- Spain
- INIS RN
- 52049002
- Subject category
- S97: MATHEMATICAL METHODS AND COMPUTING;
- Resource subtype / Literary indicator
- Conference
- Descriptors DEI
- EVALUATION; FORECASTING; MATHEMATICS; NEURAL NETWORKS; STATISTICS
- Descriptors DEC
- MATHEMATICS