论文信息 - More accurate tests for the statistical significance of result differences

More accurate tests for the statistical significance of result differences

Statistical significance testing of differences in values of metrics like recall, precision and balanced F-score is a necessary part of empirical natural language processing. Unfortunately, we find in a set of experiments that many commonly used tests often underestimate the significance and so are less likely to detect differences that exist between different techniques. This underestimation comes from an independence assumption that is often violated. We point out some useful tests that do not make this assumption, including computationally-intensive randomization tests.

Alexander S. Yeh | A. Yeh

[1] W. J. Langford. Statistical Methods , 1959, Nature.

[2] G. W. Snedecor. Statistical Methods , 1964 .

[3] Michael A. Malcolm,et al. Computer methods for mathematical computations , 1977 .

[4] S. Addelman. Statistics for experimenters , 1978 .

[5] Statistical methods , 1980 .

[6] R. Larsen. An introduction to mathematical statistics and its applications / Richard J. Larsen, Morris L. Marx , 1986 .

[7] K. J. Evans,et al. Computer Intensive Methods for Testing Hypotheses: An Introduction , 1990 .

[8] Kenneth Ward Church,et al. Introduction to the Special Issue on Computational Linguistics Using Large Corpora , 1993, Comput. Linguistics.

[9] Lynette Hirschman,et al. Evaluating Message Understanding Systems: An Analysis of the Third Message Understanding Conference (MUC-3) , 1993, CL.

[10] S.J.J. Smith,et al. Empirical Methods for Artificial Intelligence , 1995 .