Title of article
Steps for Creating Two Persian Specialized Corpora
Author/Authors
Alayiaboozar, Elham Iranian Research Institute for Information Science and Technology (IranDoc), Tehran, Iran , Hojjatpanah, Ali Asghar Iranian Research Institute for Information Science and Technology (IranDoc), Tehran, Iran
Pages
13
From page
231
To page
243
Abstract
Currently, most linguistic studies benefit from valid linguistic data available at corpora. Compiling corpora is a common practice in linguistic research. The present study introduces two specialized corpora in Persian; a specialized corpus is used to study a particular type of language or language variety. For building such corpora, first, a set of texts were compiled based on pre-established criteria used in the sampling process (including the mode of the texts, type of the texts, domain of the texts, language/ language varieties of the texts and the date of the texts). The corpora are specialized because they include technical terms in information processing and management, librarianship, linguistics, computational linguistics, thesaurus building, managing, policy-making, natural language processing, information technology, information retrieval, ontology and other related interdisciplinary domains. After compiling data and Metadata, the texts were preprocessed (normalized and tokenized) and annotated (automated POS tagging); finally, the tags were manually checked. Each corpus includes more than four million words. Since not many specialized corpora are built in Persian, such corpora could be considered valuable resources for researchers interested in studying linguistic variations in Persian interdisciplinary texts.
Keywords
Persian Corpus , Specialized Corpus , Building Corpora , Text Preprocessing , Corpus Annotation
Journal title
International Journal of Information Science and Management (IJISM)
Serial Year
2022
Record number
2730086
Link To Document