论文信息 - Handling of Missing Values in Lexical Acquisition

Handling of Missing Values in Lexical Acquisition

In this work we propose a strategy to reduce the impact of the sparse data problem in the tasks of lexical information acquisition based on the observation of linguistic cues. We propose a way to handle the uncertainty created by missing values, that is, when a zero value could mean either that the cue has not been observed because the word in question does not belong to the class, i.e. negative evidence, or that the word in question has just not been observed in the context sought by chance, i.e. lack of evidence. This uncertainty creates problems to the learner, because zero values for incompatible labelled examples make the cue lose its predictive capacity and even though some samples display the sought context, it is not taken into account. In this paper we present the results of our experiments to try to reduce this uncertainty by, as other authors do (Joanis et al. 2007, for instance), substituting zero values for pre-processed estimates. Here we present a first round of experiments that have been the basis for the estimates of linguistic information motivated by lexical classes. We obtained experimental results that show a clear benefit of the proposed approach.

Núria Bel | Núria Bel

[1] Montserrat Marimon,et al. Automatic Acquisition of Grammatical Types for Nouns , 2007, HLT-NAACL.

[2] Chih-Jen Lin,et al. LIBSVM: A library for support vector machines , 2011, TIST.

[3] G. Āllport. The Psycho-Biology of Language. , 1936 .

[4] G. Zipf,et al. The Psycho-Biology of Language , 1936 .

[5] Ted Briscoe,et al. Automatic Acquisition of Adjectival Subcategorization from Corpora , 2005, ACL.

[6] MerloPaola,et al. Automatic verb classification based on statistical distributions of argument structure , 2001 .

[7] Timothy Baldwin,et al. Learning the Countability of English Nouns from Corpus Data , 2003, ACL.

[8] Brendan S. Gillon,et al. Towards a common semantics for english count and mass nouns , 1992 .

[9] 金田重郎,et al. C4.5: Programs for Machine Learning (書評) , 1995 .

[10] Marc Light,et al. Morphological Cues for Lexical Semantics , 1996, ACL.

[11] A. Ross. Structural Linguistics , 1953, Nature.