论文信息 - Building an Arabic Social Corpus for Dangerous Profile Extraction on Social Networks

Building an Arabic Social Corpus for Dangerous Profile Extraction on Social Networks

Social networks are considered today as revolutionary tools of communication that have a tremendous impact on our lives. However, these tools can be manipulated by vicious users namely terrorists. The process of collecting and analyzing such profiles is a considerably challenging task which has not yet been well established. For this purpose, we propose, in this paper, a new method for data extraction and annotation of suspicious users from social networks threatening the national security. Our method allows constructing a rich Arabic corpus designed for detecting terrorist users spreading on social networks. The amendment of our corpora is ensured following a set of rules defined by a domain expert. All these steps are described in details, and some typical examples are given. Also, some statistics are reported from the data collection and annotation stages as well as the evaluation of the annotated features based on the intra-agreement measurement between different experts.

[1] Abdelmajid Ben Hamadou,et al. An extraction and unification methodology for social networks data: an application to public security , 2017, iiWAS.

[2] Brendan T. O'Connor,et al. Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments , 2010, ACL.

[3] J. Klausen. Tweeting the Jihad: Social Media Networks of Western Foreign Fighters in Syria and Iraq , 2015 .

[4] Robert R. Faulkner,et al. Qualitative Sociology: A Method to the Madness , 1979, Contemporary Sociology: A Journal of Reviews.

[5] Preslav Nakov,et al. Developing a successful SemEval task in sentiment analysis of Twitter and other social media texts , 2016, Language Resources and Evaluation.

[6] J. Sim,et al. The kappa statistic in reliability studies: use, interpretation, and sample size requirements. , 2005, Physical therapy.

[7] Abdelmajid Ben Hamadou,et al. GLIO: A New Method for Grouping Like-Minded Users , 2015, Trans. Comput. Collect. Intell..

[8] Amal Rekik,et al. Deep Learning for Hot Topic Extraction from Social Streams , 2016, HIS.

[9] J. R. Landis,et al. The measurement of observer agreement for categorical data. , 1977, Biometrics.

[10] A. Stefanidis,et al. Harvesting ambient geospatial information from social media feeds , 2011, GeoJournal.