Developmental corpus (without language corrections) Šolar 2.0 Clear

Solar 2.0 Clear is an adapted version of the Solar 2.0 corpus, cf. http://hdl.handle.net/11356/1214. The Solar 2.0 Clear corpus consists of texts written by students in Slovene primary and secondary schools. School essays form the majority of the corpus while other material includes texts created during lessons, such as text recapitulations or descriptions, examples of formal applications etc. For each text, the information on school (elementary or secondary), subject, level (grade or year), type of text, region and date of production is provided. Unlike the original Solar 2.0 corpus (http://hdl.handle.net/11356/1214), Solar 2.0 Clear includes student texts only: error annotations and other types of feedback from the teachers have been removed. The corpus can thus be used for processing tasks where the inclusion of corrections hinders or complicates the procedures (e.g. for comparative data extraction, training of language models etc).