Predicting ICU Death with Summarized Data: The Emerging Health Data Search Engine
Abstract
Chignell et al. [1] previously described a methodology for converting a large set of confidential data records into a set of summaries of similar patients. They claimed that the resulting patient types could “capture important trends and patterns in the data set without disclosing the information in any of the individual data records.” In this paper we examine the predictive validity of an initial set of patient types developed by [1]. We ask the following question: To what extent can the summarized data derived from each cluster (patient type) be as informative as the original case level data (individuals) from which the clusters were inferred? We address this question by assessing how well predictions made with summarized data matched predictions made with original data. After reviewing relevant literature, and explaining how data is summarized in each cluster of similar patients, we compare the results of predicting death in the ICU1 using both summarized (regression analysis) and original case data (discriminant analysis and logistic regression analysis). When multiple clusters were used, prediction based on regression analysis of the summarized data was found to be better than prediction using either logistic regression or discriminant analysis on the raw data. We hypothesize that this result is due to segmentation of a heterogenous multivariate space into more homogeneous subregions. We see the present results as an important step towards the development of generalized health data search engines that can utilize non-confidential summarized data passed through health data repository firewalls.
Authors: Mahsa Rouzbahman, Mark Chignell
Published in: TSpace (University of Toronto) (2014)