• Title of article

    Steps for Creating Two Persian Specialized Corpora

  • Author/Authors

    Alayiaboozar, Elham Iranian Research Institute for Information Science and Technology (IranDoc), Tehran, Iran , Hojjatpanah, Ali Asghar Iranian Research Institute for Information Science and Technology (IranDoc), Tehran, Iran

  • Pages
    13
  • From page
    231
  • To page
    243
  • Abstract
    Currently, most linguistic studies benefit from valid linguistic data available at corpora. Compiling corpora is a common practice in linguistic research. The present study introduces two specialized corpora in Persian; a specialized corpus is used to study a particular type of language or language variety. For building such corpora, first, a set of texts were compiled based on pre-established criteria used in the sampling process (including the mode of the texts, type of the texts, domain of the texts, language/ language varieties of the texts and the date of the texts). The corpora are specialized because they include technical terms in information processing and management, librarianship, linguistics, computational linguistics, thesaurus building, managing, policy-making, natural language processing, information technology, information retrieval, ontology and other related interdisciplinary domains. After compiling data and Metadata, the texts were preprocessed (normalized and tokenized) and annotated (automated POS tagging); finally, the tags were manually checked. Each corpus includes more than four million words. Since not many specialized corpora are built in Persian, such corpora could be considered valuable resources for researchers interested in studying linguistic variations in Persian interdisciplinary texts.
  • Keywords
    Persian Corpus , Specialized Corpus , Building Corpora , Text Preprocessing , Corpus Annotation
  • Journal title
    International Journal of Information Science and Management (IJISM)
  • Serial Year
    2022
  • Record number

    2730086