论文信息 - ArchiMob - A Corpus of Spoken Swiss German

ArchiMob - A Corpus of Spoken Swiss German

Swiss dialects of German are, unlike most dialects of well standardised languages, widely used in everyday communication. Despite this fact, automatic processing of Swiss German is still a considerable challenge due to the fact that it is mostly a spoken variety rarely recorded and that it is subject to considerable regional variation. This paper presents a freely available general-purpose corpus of spoken Swiss German suitable for linguistic research, but also for training automatic tools. The corpus is a result of a long design process, intensive manual work and specially adapted computational processing. We first describe how the documents were transcribed, segmented and aligned with the sound source, and how inconsistent transcriptions were unified through an additional normalisation layer. We then present a bootstrapping approach to automatic normalisation using different machine-translation-inspired methods. Furthermore, we evaluate the performance of part-of-speech taggers on our data and show how the same bootstrapping approach improves part-of-speech tagging by 10% over four rounds. Finally, we present the modalities of access of the corpus as well as the data format.

[1] Thomas Schmidt. EXMARaLDA and the FOLK tools - two toolsets for transcribing and annotating spoken language , 2012, LREC.

[2] Tanja Samardzic,et al. Lemmatisation as a Tagging Task , 2012, ACL.

[3] Yves Scherrer,et al. Normalising orthographic and dialectal variants for the automatic processing of Swiss German , 2015 .

[4] Adrian Leemann,et al. Dialäkt Äpp: communicating dialectology to the public – crowdsourcing dialects from the public , 2015 .

[5] Pavel Rychlý,et al. Manatee/Bonito - A Modular Corpus Manager , 2007, RASLAN.

[6] Adam Kilgarriff,et al. The Sketch Engine: ten years on , 2014 .

[7] Florian Schiel,et al. Signal processing via web services: The use case WebMAUS , 2012 .

[8] Amir Zeldes,et al. ANNIS3: A new architecture for generic corpus query and visualization , 2016, Digit. Scholarsh. Humanit..

[9] Yves Scherrer,et al. Natural Language Processing for the Swiss German Dialect Area , 2010, KONVENS.

[10] Jürgen Eichhoff,et al. Sprachatlas der deutschen Schweiz , 2000 .

[11] Nora Hollenstein,et al. Compilation of a Swiss German Dialect Corpus and its Application to PoS Tagging , 2014, VarDial@COLING.

[12] H. Baumgartner,et al. Sprachatlas der deutschen Schweiz , 1964 .