论文信息 - Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Commonly used evaluation measures including Recall, Precision, F-Factor and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the statistic. Using these measures a system that performs worse in the objective sense of Informedness, can appear to perform better under any of these commonly used measures. We discuss several concepts and measures that reflect the probability that prediction is informed versus chance. Informedness and introduce Markedness as a dual measure for the probability that prediction is marked versus chance. Finally we demonstrate elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision, and outline the extension from the dichotomous case to the general multi-class case. .

David M. W. Powers | D. Powers

[1] Sandip Sinharay,et al. Editors Appointed for Journal of Educational and Behavioral Statistics , 2010 .

[2] Douglas G. Bonett,et al. Inferential Methods for the Tetrachoric Correlation Coefficient , 2005 .

[3] Ronald W. Manderscheid,et al. Approximating the Moments and Distribution of the Likelihood Ratio Statistic for Multinomial Goodness of Fit , 1981 .

[4] Johannes Fürnkranz,et al. ROC ‘n’ Rule Learning—Towards a Better Understanding of Covering Algorithms , 2005, Machine Learning.

[5] J. A. Adams,et al. Psychological bulletin. , 1962, Psychological bulletin.

[6] Pieter Reitsma,et al. Educational and Psychological Measurement , 2003 .

[7] D. Shanks. Is Human Learning Rational? , 1995, The Quarterly journal of experimental psychology. A, Human experimental psychology.

[8] Trent W. Lewis,et al. Audio-Visual Speech Recognition Using Red Exclusion and Neural Networks , 2002, ACSC.

[9] Raj Madhavan,et al. Performance Metrics for Intelligent Systems (PerMIS) 2006Workshop: Summary and Review , 2006, 35th IEEE Applied Imagery and Pattern Recognition Workshop (AIPR'06).

[10] F. James Rohlf,et al. Biometry: The Principles and Practice of Statistics in Biological Research , 1969 .

[11] Pierre Perruchet,et al. The exploitation of distributional information in syllable processing , 2004, Journal of Neurolinguistics.