Predicting the characteristics of people living in the South USA using logistic regression and decision tree

Analysis of social data is at the core of social studies and an important application area of data mining and knowledge discovery. One aspect of such social data analysis is based on demographic and/or economic data. In this paper, we apply data mining techniques to find the characteristics of people living in the south of USA. The data used in our study is the WAGE2 data set with 935 observations that has been used in some previous social study research. The software tool SAS Enterprise Miner was used to analyze the data, particularly the regression and decision tree models. The results of our analysis show that the decision tree model produced a better variable selection than the logistic regression model did to predict if a person is likely to live in the south than the logistic regression model, at least from the given data set.