Bayesian variable selection for linear regression in high dimensional microarray data

Autor: Veerabhadran Baladandayuthapani, David Sergio Matusevich, Carlos Ordonez, Wellington Cabrera
Rok vydání: 2013
Předmět:
Zdroj: DTMBIO
DOI: 10.1145/2512089.2512094
Popis: Variable selection is a fundamental problem in Bayesian statistics whose solution requires exploring a combinatorial search space. We study the solution of variable selection with a well-known MCMC method, which requires thousands of iterations. We present several algorithmic optimizations to accelerate the MCMC method to make it work efficiently inside a database system. Our optimizations include sufficient statistics, variable preselection, hash tables and calling a linear algebra library. We present experiments with very high dimensional microarray data sets to predict cancer survival time. We discuss encouraging findings, identifying specific genes likely to predict the survival time for brain cancer patients. We also show our DBMS-based algorithm is orders of magnitude faster than the R statistical package. Our work shows a DBMS is a promising platform to analyze microarray data.
Databáze: OpenAIRE