Building predictive models for MERS-CoV infections using data mining techniques

Autor: Tahani Mazyad Almutairi, Mona D. Al-Shahrani, Isra Al-Turaiki
Rok vydání: 2016
Předmět:
Male
J48
Saudi Arabia
Decision tree
Stability (learning theory)
02 engineering and technology
computer.software_genre
Article
lcsh:Infectious and parasitic diseases
MERS-CoV
Naive Bayes
03 medical and health sciences
Naive Bayes classifier
0302 clinical medicine
C4.5 algorithm
Health care
0202 electrical engineering
electronic engineering
information engineering

Data Mining
Humans
Medicine
lcsh:RC109-216
Computer Simulation
030212 general & internal medicine
Aged
Aged
80 and over

business.industry
lcsh:Public aspects of medicine
Decision tree learning
Public Health
Environmental and Occupational Health

lcsh:RA1-1270
General Medicine
Middle Aged
Prognosis
Classification
Survival Analysis
Treatment Outcome
Infectious Diseases
Female
020201 artificial intelligence & image processing
Christian ministry
Data mining
Coronavirus Infections
business
computer
Predictive modelling
Zdroj: Journal of Infection and Public Health
Journal of Infection and Public Health, Vol 9, Iss 6, Pp 744-748 (2016)
ISSN: 1876-0341
DOI: 10.1016/j.jiph.2016.09.007
Popis: Summary: Background: Recently, the outbreak of MERS-CoV infections caused worldwide attention to Saudi Arabia. The novel virus belongs to the coronaviruses family, which is responsible for causing mild to moderate colds. The control and command center of Saudi Ministry of Health issues a daily report on MERS-CoV infection cases. The infection with MERS-CoV can lead to fatal complications, however little information is known about this novel virus. In this paper, we apply two data mining techniques in order to better understand the stability and the possibility of recovery from MERS-CoV infections. Method: The Naive Bayes classifier and J48 decision tree algorithm were used to build our models. The dataset used consists of 1082 records of cases reported between 2013 and 2015. In order to build our prediction models, we split the dataset into two groups. The first group combined recovery and death records. A new attribute was created to indicate the record type, such that the dataset can be used to predict the recovery from MERS-CoV. The second group contained the new case records to be used to predict the stability of the infection based on the current status attribute. Results: The resulting recovery models indicate that healthcare workers are more likely to survive. This could be due to the vaccinations that healthcare workers are required to get on regular basis. As for the stability models using J48, two attributes were found to be important for predicting stability: symptomatic and age. Old patients are at high risk of developing MERS-CoV complications. Finally, the performance of all the models was evaluated using three measures: accuracy, precision, and recall. In general, the accuracy of the models is between 53.6% and 71.58%. Conclusion: We believe that the performance of the prediction models can be enhanced with the use of more patient data. As future work, we plan to directly contact hospitals in Riyadh in order to collect more information related to patients with MERS-CoV infections. Keywords: MERS-CoV, Data mining, Decision tree, J48, Naive Bayes, Classification
Databáze: OpenAIRE