Developing a dengue forecast model using machine learning: A case study in China

PLoS Negl Trop Dis. 2017 Oct 16;11(10):e0005973. doi: 10.1371/journal.pntd.0005973. eCollection 2017 Oct.

Abstract

Background: In China, dengue remains an important public health issue with expanded areas and increased incidence recently. Accurate and timely forecasts of dengue incidence in China are still lacking. We aimed to use the state-of-the-art machine learning algorithms to develop an accurate predictive model of dengue.

Methodology/principal findings: Weekly dengue cases, Baidu search queries and climate factors (mean temperature, relative humidity and rainfall) during 2011-2014 in Guangdong were gathered. A dengue search index was constructed for developing the predictive models in combination with climate factors. The observed year and week were also included in the models to control for the long-term trend and seasonality. Several machine learning algorithms, including the support vector regression (SVR) algorithm, step-down linear regression model, gradient boosted regression tree algorithm (GBM), negative binomial regression model (NBM), least absolute shrinkage and selection operator (LASSO) linear regression model and generalized additive model (GAM), were used as candidate models to predict dengue incidence. Performance and goodness of fit of the models were assessed using the root-mean-square error (RMSE) and R-squared measures. The residuals of the models were examined using the autocorrelation and partial autocorrelation function analyses to check the validity of the models. The models were further validated using dengue surveillance data from five other provinces. The epidemics during the last 12 weeks and the peak of the 2014 large outbreak were accurately forecasted by the SVR model selected by a cross-validation technique. Moreover, the SVR model had the consistently smallest prediction error rates for tracking the dynamics of dengue and forecasting the outbreaks in other areas in China.

Conclusion and significance: The proposed SVR model achieved a superior performance in comparison with other forecasting techniques assessed in this study. The findings can help the government and community respond early to dengue epidemics.

MeSH terms

  • Algorithms
  • China / epidemiology
  • Climate
  • Dengue / epidemiology*
  • Dengue / virology
  • Disease Outbreaks
  • Forecasting / methods*
  • Humans
  • Incidence
  • Linear Models
  • Machine Learning*
  • Predictive Value of Tests
  • Public Health / methods*
  • Temperature

Grants and funding

This work was supported by the National Natural Science Foundation of China (No. 81703323 and 81773497). TL received the Guangdong Provincial Science and Technology Project Funding (NO.2014A040401041). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.