论文信息 - An Open-Source Data Driven Spell Checker for Sinhala

An Open-Source Data Driven Spell Checker for Sinhala

In this paper we describe the construction of a spell checker for Sinhala, the language spoken by the majority in Sri Lanka. Due to its morphological richness, the language is difficult to enumerate completely in a lexicon. The approach described is based on n-gram statistics and is relatively inexpensive to construct without deep linguistic knowledge. This approach is particularly useful as there are very few linguistic resources available for Sinhala at present. The proposed algorithm has been shown to be able to detect and correct many of the common spelling errors of the language. Results show a promising performance achieving an average accuracy of 82%. This technique can also be applied to construct spell checkers for other phonetic languages whose linguistic resources are scarce or non-existent. DOI: http://dx.doi.org/10.4038/icter.v3i1.2844 ICTer Vol.3 No.1 2010

[1] W. S. Karunatillake. An Introduction to spoken Sinhala , 1992 .

[2] Bidyut Baran Chaudhuri,et al. Towards Indian language spell-checker design , 2002, Language Engineering Conference, 2002. Proceedings.

[3] Karen Kukich,et al. Techniques for automatically correcting words in text , 1992, CSUR.

[4] V. Dixit,et al. Design and implementation of a morphology-based spellchecker for Marathi, an Indian language , 2005 .

[5] Kumudu Gamage,et al. Sinhala Grapheme-to-Phoneme Conversion and Rules for Schwa Epenthesis , 2006, ACL.

[6] S. B. Nair,et al. Design and implementation of a spell checker for Assamese , 2002, Language Engineering Conference, 2002. Proceedings.