Bayesian variable selection for linear regression in high dimensional microarray data
Autor: | Veerabhadran Baladandayuthapani, David Sergio Matusevich, Carlos Ordonez, Wellington Cabrera |
---|---|
Rok vydání: | 2013 |
Předmět: |
Microarray
Computer science business.industry Microarray analysis techniques Markov chain Monte Carlo Feature selection computer.software_genre Machine learning Hash table Bayesian statistics symbols.namesake Variable (computer science) ComputingMethodologies_PATTERNRECOGNITION Linear regression Linear algebra symbols Combinatorial search Data mining Artificial intelligence business computer |
Zdroj: | DTMBIO |
DOI: | 10.1145/2512089.2512094 |
Popis: | Variable selection is a fundamental problem in Bayesian statistics whose solution requires exploring a combinatorial search space. We study the solution of variable selection with a well-known MCMC method, which requires thousands of iterations. We present several algorithmic optimizations to accelerate the MCMC method to make it work efficiently inside a database system. Our optimizations include sufficient statistics, variable preselection, hash tables and calling a linear algebra library. We present experiments with very high dimensional microarray data sets to predict cancer survival time. We discuss encouraging findings, identifying specific genes likely to predict the survival time for brain cancer patients. We also show our DBMS-based algorithm is orders of magnitude faster than the R statistical package. Our work shows a DBMS is a promising platform to analyze microarray data. |
Databáze: | OpenAIRE |
Externí odkaz: |