Data Mining Methods for Optimizing Feature Extraction and Model Selection
Abstract
How can we carry out on-the-fly data mining on massive amounts of data, to make relevant predictions, based on data for similar observations to the one currently under consideration? In this paper we show the benefit of using large numbers of computationally efficient analyses to tune the feature extraction, and prediction, steps in data mining, using cross-validated prediction accuracy as the evaluative criterion. Different feature extraction strategies are also compared in terms of their predictive effectiveness in this context. While the research reported here focused on clinical prediction of healthcare outcomes, the results should have broader implications for large scale data mining in general.
Authors: Masha Rouzbahman, Alexandra Jovicic, Lu Wang, Leon Zucherman, Zahid Abul-Basher, Nipon Charoenkitkarn, Mark Chignell