论文信息 - Evaluating HeLI with Non-Linear Mappings

Evaluating HeLI with Non-Linear Mappings

In this paper we describe the non-linear mappings we used with the Helsinki language identification method, HeLI, in the 4th edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial 2017 workshop. Our SUKI team participated in the closed track together with 10 other teams. Our system reached the 7th position in the track. We describe the HeLI method and the non-linear mappings in mathematical notation. The HeLI method uses a probabilistic model with character n-grams and word-based backoff. We also describe our trials using the non-linear mappings instead of relative frequencies and we present statistics about the back-off function of the HeLI method.

Krister Lindén | Tommi Jauhiainen | Heidi Jauhiainen

[1] Marcos Zampieri,et al. Automatic identification of language varieties: The case of Portuguese , 2012, KONVENS.

[2] Preslav Nakov,et al. Overview of the DSL Shared Task 2015 , 2015 .

[3] Jörg Tiedemann,et al. Efficient Discrimination Between Closely Related Languages , 2012, COLING.

[4] Jörg Tiedemann,et al. A Report on the DSL Shared Task 2014 , 2014, VarDial@COLING.

[5] Shervin Malmasi,et al. Arabic Dialect Identification Using a Parallel Multidialectal Corpus , 2015, PACLING.

[6] Chew Yew Choong,et al. Optimizing n‑gram Order of an n‑gram Based Language Identification Algorithm for 68 Written Languages , 2009 .

[7] Marcos Zampieri,et al. N-gram Language Models and POS Distribution for the Identification of Spanish Varieties (Ngrammes et Traits Morphosyntaxiques pour la Identification de Variétés de l’Espagnol) [in French] , 2013, JEP/TALN/RECITAL.

[8] Marcos Zampieri,et al. Using bag-of-words to distinguish similar languages: How efficient are they? , 2013, 2013 IEEE 14th International Symposium on Computational Intelligence and Informatics (CINTI).

[9] Nikola Ljubešić,et al. Discriminating between VERY similar languages among Twitter users , 2014 .

[10] Krister Lindén,et al. Discriminating Similar Languages with Token-Based Backoff , 2015 .

[11] Yves Bestgen,et al. Improving the Character Ngram Model for the DSL Task with BM25 Weighting and Less Frequently Used Feature Sets , 2017, VarDial.