Constructing Linguistic Resources for the Tunisian Dialect Using Textual User-Generated Contents on the Social Web

In Arab countries, the dialect is daily gaining ground in the social interaction on the web and swiftly adapting to globalization. Strengthening the relationship of its practitioners with the outside world and facilitating their social exchanges, the dialect encompasses every day new transcriptions...

Full description

Saved in:

Bibliographic Details
Published in	Current Trends in Web Engineering Vol. 9396; pp. 3 - 14
Main Authors	Younes, Jihen, Achour, Hadhemi, Souissi, Emna
Format	Book Chapter
Language	English
Published	Switzerland Springer International Publishing AG 2015 Springer International Publishing
Series	Lecture Notes in Computer Science
Subjects	Corpus construction Data mining Dictionary construction Information retrieval Language identification Social web textual contents Software Engineering Tunisian dialect
Online Access	Get full text
ISBN	3319247999 9783319247991
ISSN	0302-9743 1611-3349
DOI	10.1007/978-3-319-24800-4_1

Cover

More Information
Summary:	In Arab countries, the dialect is daily gaining ground in the social interaction on the web and swiftly adapting to globalization. Strengthening the relationship of its practitioners with the outside world and facilitating their social exchanges, the dialect encompasses every day new transcriptions that arouse the curiosity of researchers in the NLP community. In this article, we focus specifically on the Tunisian dialect processing. Our goal is to build corpora and dictionaries allowing us to begin our study of this language and to identify its specificities. As a first step, we extract textual user-generated contents on the social Web, we then conduct an automatic content filtering and classification, leaving only the texts containing Tunisian dialect. Finally, we present some of its salient features from the built corpora.
ISBN:	3319247999 9783319247991
ISSN:	0302-9743 1611-3349
DOI:	10.1007/978-3-319-24800-4_1