论文信息 - Learning pronunciation with the Visual ear

Learning pronunciation with the Visual ear

We recently reported the use of Kohonen's feature map as the hidden layer of an RBF network for the recognition of spoken letters [1], and the analysis of sleep EEG [2]. The feature map was shown to act as an aid to visualization during the initial period of unsupervised learning in the hidden layer. In this paper, we again explore the topology preserving properties of Kohonen's feature map, this time for the visual interpretation of speech. It is shown that speech sounds, such as words or phonemes, may be displayed as moving trajectories on a computer screen and enhanced for ease of interpretation. A system known as the Visual Ear is introduced, in which speech from a normal speaker is displayed alongside that of a pupil learning pronunciation, enabling a visual comparison to be made between the two. The application of the Visual Ear to accelerated learning of foreign languages, or as a general speech therapy tool, are then discussed, and the limitations of the present system are highlighted.

Lionel Tarassenko | Jake Reynolds

[1] Stan Davis,et al. Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Se , 1980 .

[2] R. Lippmann,et al. An introduction to computing with neural nets , 1987, IEEE ASSP Magazine.

[3] Lionel Tarassenko,et al. Spoken Letter Recognition with Neural Networks , 1991, Int. J. Neural Syst..

[4] Teuvo Kohonen,et al. Self-Organization and Associative Memory , 1988 .

[5] W. Hardcastle,et al. Visual display of tongue-palate contact: electropalatography in the assessment and remediation of speech disorders. , 1991, The British journal of disorders of communication.

[6] L. Tarassenko,et al. Analysis of the sleep EEG using a multilayer network with spatial organisation , 1992 .

[7] John Moody,et al. Fast Learning in Networks of Locally-Tuned Processing Units , 1989, Neural Computation.

[8] R. Linggard,et al. Neural arrays for speech recognition , 1990 .