论文信息 - Artificial Error Generation with Fluency Filtering

Artificial Error Generation with Fluency Filtering

The quantity and quality of training data plays a crucial role in grammatical error correction (GEC). However, due to the fact that obtaining human-annotated GEC data is both time-consuming and expensive, several studies have focused on generating artificial error sentences to boost training data for grammatical error correction, and shown significantly better performance. The present study explores how fluency filtering can affect the quality of artificial errors. By comparing artificial data filtered by different levels of fluency, we find that artificial error sentences with low fluency can greatly facilitate error correction, while high fluency errors introduce more noise.

Jungyeul Park | Mengyang Qiu | Jungyeul Park | Mengyang Qiu

[1] Mariano Felice,et al. Artificial error generation for translation-based grammatical error correction , 2016 .

[2] Wei Zhao,et al. Improving Grammatical Error Correction via Pre-Training a Copy-Augmented Architecture with Unlabeled Data , 2019, NAACL.

[3] Ted Briscoe,et al. Automatic Annotation and Evaluation of Error Types for Grammatical Error Correction , 2017, ACL.

[4] Kenneth Heafield,et al. KenLM: Faster and Smaller Language Model Queries , 2011, WMT@EMNLP.

[5] Daniel Jurafsky,et al. Noising and Denoising Natural Language: Diverse Backtranslation for Grammar Correction , 2018, NAACL.

[6] Zheng Yuan,et al. Constrained Grammatical Error Correction using Statistical Machine Translation , 2013, CoNLL Shared Task.

[7] Ming Zhou,et al. Fluency Boost Learning and Inference for Neural Grammatical Error Correction , 2018, ACL.

[8] H. Ng,et al. A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction , 2018, AAAI.

[9] Sebastian Riedel,et al. Wronging a Right: Generating Better Errors to Improve Grammatical Error Detection , 2018, EMNLP.

[10] Quoc V. Le,et al. Sequence to Sequence Learning with Neural Networks , 2014, NIPS.