Using country-level variables to classify countries according to the number of confirmed COVID-19 cases: An unsupervised machine learning approach

Background: The COVID-19 pandemic has attracted the attention of researchers and clinicians whom have provided evidence about risk factors and clinical outcomes. Research on the COVID-19 pandemic benefiting from open-access data and machine learning algorithms is still scarce yet can produce relevant and pragmatic information. With country-level pre-COVID-19-pandemic variables, we aimed to cluster countries in groups with shared profiles of the COVID-19 pandemic. Methods: Unsupervised machine learning algorithms (k-means) were used to define data-driven clusters of countries; the algorithm was informed by disease prevalence estimates, metrics of air pollution, socio-economic status and health system coverage. Using the one-way ANOVA test, we compared the clusters in terms of number of confirmed COVID-19 cases, number of deaths, case fatality rate and order in which the country reported the first case. Results: The model to define the clusters was developed with 155 countries. The model with three principal component analysis parameters and five or six clusters showed the best ability to group countries in relevant sets. There was strong evidence that the model with five or six clusters could stratify countries according to the number of confirmed COVID-19 cases (p<0.001). However, the model could not stratify countries in terms of number of deaths or case fatality rate. Conclusions: A simple data-driven approach using available global information before the COVID-19 pandemic, seemed able to classify countries in terms of the number of confirmed COVID-19 cases. The model was not able to stratify countries based on COVID-19 mortality data.

[1]  Md Zahidul Islam,et al.  Comparing sets of patterns with the Jaccard index , 2018, Australas. J. Inf. Syst..

[2]  Toshiya Murai,et al.  Distinct Patterns of Cerebral Cortical Thinning in Schizophrenia: A Neuroimaging Data-Driven Approach , 2016, Schizophrenia bulletin.

[3]  Zhaofeng Chen,et al.  Prevalence of comorbidities and its effects in patients infected with SARS-CoV-2: a systematic review and meta-analysis , 2020, International Journal of Infectious Diseases.

[4]  L. Groop,et al.  Novel subgroups of adult-onset diabetes and their association with outcomes: a data-driven cluster analysis of six variables. , 2018, The lancet. Diabetes & endocrinology.

[5]  S. Lo,et al.  A familial cluster of pneumonia associated with the 2019 novel coronavirus indicating person-to-person transmission: a study of a family cluster , 2020, The Lancet.

[6]  G. Castellanos-Dominguez,et al.  Weighted-PCA for unsupervised classification of cardiac arrhythmias , 2010, 2010 Annual International Conference of the IEEE Engineering in Medicine and Biology.

[7]  Roger Detels,et al.  Environmental Health: a Global Access Science Source Air Pollution and Case Fatality of Sars in the People's Republic of China: an Ecologic Study , 2022 .

[8]  Rui Ji,et al.  Prevalence of comorbidities and its effects in patients infected with SARS-CoV-2: a systematic review and meta-analysis , 2020, International Journal of Infectious Diseases.

[9]  P. Rousseeuw Silhouettes: a graphical aid to the interpretation and validation of cluster analysis , 1987 .

[10]  Y. Hu,et al.  Clinical features of patients infected with 2019 novel coronavirus in Wuhan, China , 2020, The Lancet.

[11]  Denny Meyer,et al.  Exploring Heterogeneity on the Wisconsin Card Sorting Test in Schizophrenia Spectrum Disorders: A Cluster Analytical Investigation , 2019, Journal of the International Neuropsychological Society.

[12]  Spiros Denaxas,et al.  Identifying clinically important COPD sub-types using data-driven approaches in primary care population based electronic health records , 2019, BMC Medical Informatics and Decision Making.

[13]  Ting Yu,et al.  Epidemiological and clinical characteristics of 99 cases of 2019 novel coronavirus pneumonia in Wuhan, China: a descriptive study , 2020, The Lancet.

[14]  Kuo-Lung Wu,et al.  Unsupervised possibilistic clustering , 2006, Pattern Recognit..

[15]  E. Dong,et al.  An interactive web-based dashboard to track COVID-19 in real time , 2020, The Lancet Infectious Diseases.