The Web contains vast amounts of linguistic data. One key issue for linguists and language technologists is how to access it. Commercial search engines give highly compromised access. An alternative is to crawl the Web ourselves, which also allows us to remove duplicates and near-duplicates, navigational material, and a range of other kinds of non-linguistic matter. We can also tokenize, lemmatise and part-of-speech tag the corpus, and load the data into a corpus query tool which supports sophisticated linguistic queries. We have now done this for German and Italian, with corpus sizes of over 1 billion words in each case. We provide Web access to the corpora in our query tool, the Sketch Engine.
[1]
Geoffrey Zweig,et al.
Syntactic Clustering of the Web
,
1997,
Comput. Networks.
[2]
Nicholas Ostler,et al.
Corpus Design Criteria
,
1992
.
[3]
Adam Kilgarriff,et al.
Introduction to the Special Issue on the Web as Corpus
,
2003,
CL.
[4]
William H. Fletcher.
Making the Web More Useful as a Source for Linguistic Corpora
,
2004
.
[5]
Serge Sharo.
Creating General-Purpose Corpora Using Automated Search Engine Queries
,
2006
.
[6]
Soumen Chakrabarti,et al.
Mining the web - discovering knowledge from hypertext data
,
2002
.