Investigating the Phonetic Organisation of the English Language via Phonological Networks, Percolation and Markov Models

Applying tools from network science and statistical mechanics, this paper represents an interdisciplinary analysis of the phonetic organisation of the English language. By using open datasets, we build phonological networks, where nodes are the phonetic pronunciations of words and edges connect words differing by the addition, deletion, or substitution of exactly one phoneme. We present an investigation of whether the topological features of this phonological network reflect only lower or also higher order correlations in phoneme organisation. We address this question by exploring artificially constructed repertoires of words, constructing phonological networks for these repertoires, and comparing them to the network constructed from the real data. Artificial repertoires of words are built to reflect increasingly higher order statistics of the English corpus. Hence, we start with percolation-type experiments in which phonemes are sampled uniformly at random to construct words, then sample from the real phoneme frequency distribution, and finally we consider repertoires resulting from Markov processes of first, second, and third order. As expected, we find that percolation-type experiments constitute a poor null model for the real data. However, some network features, such as the relatively high assortative mixing by degree and the clustering coefficient of the English PN, can be retrieved by Markov models for word construction. Nevertheless, even Markov processes up to third order cannot fully reproduce other patterns of the empirical network, such as link densities and component sizes. We conjecture that this difference is related to the combinatorial space the real and the artificial phonological networks are embedded into and that the connectivity properties of phonological networks reflect additional patterns in word organisation in the English language which cannot be captured by lower order phoneme correlations.

[1]  Michael S. Vitevitch,et al.  Network Structure Influences Speech Production , 2010, Cogn. Sci..

[2]  Ricard V. Solé,et al.  Ambiguity in language networks , 2014, ArXiv.

[3]  Mark Newman,et al.  Networks: An Introduction , 2010 .

[4]  T. Griffiths,et al.  Google and the Mind , 2007, Psychological science.

[5]  M. Vitevitch What can graph theory tell us about word learning and lexical retrieval? , 2008, Journal of speech, language, and hearing research : JSLHR.

[6]  D. Pisoni,et al.  Recognizing Spoken Words: The Neighborhood Activation Model , 1998, Ear and hearing.

[7]  Gert Storms,et al.  Word associations: Network and semantic properties , 2008, Behavior research methods.

[8]  Christopher T. Kello,et al.  Scale-Free Networks in Phonological and Orthographic Wordform Lexicons , 2007 .

[9]  Clara D. Martin,et al.  Reconciling Phonological Neighborhood Effects in Speech Production through Single Trial Analysis Reconciling Phonological Neighborhood Effects in Speech Production through Single Trial Analysis , 2022 .

[10]  M. Vitevitch The Neighborhood Characteristics of Malapropisms , 1997 .

[11]  David B Pisoni,et al.  The lexical restructuring hypothesis and graph theoretic analyses of networks based on random lexicons. , 2009, Journal of speech, language, and hearing research : JSLHR.

[12]  S. Roodenrys,et al.  Complex network structure influences processing in long-term and short-term memory. , 2012, Journal of memory and language.

[13]  Jinyun Ke,et al.  Complex networks and human language , 2007, ArXiv.

[14]  J. Aitchison Words in the Mind: An Introduction to the Mental Lexicon , 1987 .

[15]  Ramon Ferrer i Cancho,et al.  The small world of human language , 2001, Proceedings of the Royal Society of London. Series B: Biological Sciences.

[16]  W. Kintsch The role of knowledge in discourse comprehension: a construction-integration model. , 1988, Psychological review.

[17]  Cynthia S. Q. Siew,et al.  Community structure in the phonological network , 2013, Front. Psychol..

[18]  G. Grimmett,et al.  Probability and random processes , 2002 .

[19]  Michael S. Vitevitch,et al.  The Structure of Phonological Networks across Multiple Languages , 2009, Int. J. Bifurc. Chaos.

[20]  Markus Brede,et al.  Patterns in the English language: phonological networks, percolation and assembly models , 2014, ArXiv.

[21]  Nick Chater,et al.  Networks in Cognitive Science , 2013, Trends in Cognitive Sciences.

[22]  J. Elman An alternative view of the mental lexicon , 2004, Trends in Cognitive Sciences.

[23]  Michael S. Vitevitch,et al.  Insights into failed lexical retrieval from network science , 2014, Cognitive Psychology.