On the resolution of ambiguities in the extraction of syntactic categories through chunking

In recent years, several authors have investigated how co-occurrence statistics in natural language can act as a cue that children may use to extract syntactic categories for the language they are learning. While some authors have reported encouraging results, it is difficult to evaluate the quality of the syntactic categories derived. It is argued in this paper that traditional measures of accuracy are inherently flawed. A valid evaluation metric needs to consider the well-formedness of utterances generated through a production end. This paper attempts to evaluate the quality of the categories derived from co-occurrence statistics through the use of MOSAIC, a computational model of syntax acquisition that has already been used to simulate several phenomena in child language. It is shown that derived syntactic categories that may appear to be of high quality quickly give rise to errors that are not typical of child speech. A solution to this problem is suggested in the form of a chunking mechanism that serves to differentiate between alternative grammatical functions of identical word forms. Results are evaluated in terms of the error rates in utterances produced by the system as well as the quantitative fit to the phenomenon of subject omission.

[1]  Fernand Gobet,et al.  Subject Omission in Children’s Language: The Case for Performance Limitations in Learning , 2019, Proceedings of the Twenty-Fourth Annual Conference of the Cognitive Science Society.

[2]  Eytan Ruppin,et al.  Bridging computational, formal and psycholinguistic approaches to language , 2004 .

[3]  J. Pine,et al.  Chunking mechanisms in human learning , 2001, Trends in Cognitive Sciences.

[4]  Julian M. Pine,et al.  A process model of children's early verb use , 2000 .

[5]  S. Gillis,et al.  Root infinitives in Dutch early child language: an effect of input? , 2001, Journal of Child Language.

[6]  Nick Chater,et al.  Distributional Information: A Powerful Cue for Acquiring Syntactic Categories , 1998, Cogn. Sci..

[7]  Toben H. Mintz Frequent frames as a cue for grammatical categories in child directed speech , 2003, Cognition.

[8]  Anna L. Theakston,et al.  The role of performance limitations in the acquisition of verb-argument structure: an alternative account. , 2001, Journal of child language.

[9]  Dietrich Dörner Proceedings of the Fifth International Conference on Cognitive Modeling , 2001 .

[10]  S A Kline The Resolution of Ambiguities , 1983, Canadian journal of psychiatry. Revue canadienne de psychiatrie.

[11]  Fernand Gobet,et al.  Modeling the Development of Children's Use of Optional Infinitives in Dutch and English Using MOSAIC , 2006, Cogn. Sci..

[12]  P. Bloom Subjectlees sentences in child language , 1990 .

[13]  B. MacWhinney The CHILDES project: tools for analyzing talk , 1992 .

[14]  H. Simon,et al.  Perception in chess , 1973 .

[15]  Fernand Gobet,et al.  Modelling children's negation errors using probabilistic learning in MOSAIC. , 2003 .

[16]  Fernand Gobet,et al.  Modelling the Development of Dutch Optional Infinitives in MOSAIC , 2019, Proceedings of the Twenty-Fourth Annual Conference of the Cognitive Science Society.

[17]  L. Gerken,et al.  Grammatical and caregiver cues in early sentence comprehension , 1999, Journal of Child Language.

[18]  Kenneth D. Forbus,et al.  Proceedings of the 26th annual meeting of the Cognitive Science Society , 2004 .