论文信息 - Classification of Imbalanced Data with Random sets and Mean-Variance Filtering

Classification of Imbalanced Data with Random sets and Mean-Variance Filtering

Imbalanced data represent a significant problem because the corresponding classifier has a tendency to ignore patterns which have smaller representation in the training set. We propose to consider a large number of balanced training subsets where representatives from the larger pattern are selected randomly. As an outcome, the system will produce a matrix of linear regression coefficients where rows represent random subsets and columns represent features. Based on the above matrix we make an assessment of the stability of the influence of the particular features. It is proposed to keep in the model only features with stable influence. The final model represents an average of the single models, which are not necessarily a linear regression. The above model had proven to be efficient and competitive during the PAKDD-2007 Data Mining Competition.

Vladimir Nikulin

[1] Ina Fourie. Social and Political Implications of Data Mining: Knowledge Management in E‐Government , 2010 .

[2] Ah Chung Tsoi,et al. Application of Text Mining Methodologies to Health Insurance Schedules , 2009 .

[3] Min Song,et al. Handbook of Research on Text and Web Mining Technologies , 2008 .

[4] Mingjun Wei,et al. A Solution to the Cross-Selling Problem of PAKDD-2007: Ensemble Model of TreeNet and Logistic Regression , 2008, Int. J. Data Warehous. Min..