Evidence for variable selection bias in classification tree algorithms based on the Gini Index is reviewed from the literature and embedded into a broader explanatory scheme: Variable selection bias in classification tree algorithms based on the Gini Index can be caused not only by the statistical effect of multiple comparisons, but also by an increasing estimation bias and variance of the splitting criterion when plug-in estimates of entropy measures like the Gini Index are employed. The relevance of these sources of variable selection bias in the different simulation study designs is examined. Variable selection bias due to the explored sources applies to all classification tree algorithms based on empirical entropy measures like the Gini Index, Deviance and Information Gain, and to both binary and multiway splitting algorithms.
[1]
Carolin Strobl,et al.
Variable Selection Bias in Classification Trees Based on Imprecise Probabilities
,
2005
.
[2]
Hyunjoong Kim,et al.
Classification Trees With Unbiased Multiway Splits
,
2001
.
[3]
Wei Zhong Liu,et al.
Bias in information-based measures in decision tree induction
,
1994,
Machine Learning.
[4]
Johannes Gehrke,et al.
Bias Correction in Classification Tree Construction
,
2001,
ICML.
[5]
J. Rice.
Mathematical Statistics and Data Analysis
,
1988
.
[6]
Igor Kononenko,et al.
On Biases in Estimating Multi-Valued Attributes
,
1995,
IJCAI.