• DocumentCode
    3696043
  • Title

    Using Wikipedia Categories for Discovering the Themes of Text Documents

  • Author

    Abdullah Bawakid

  • Author_Institution
    Fac. of Comput. &
  • Volume
    1
  • fYear
    2015
  • Firstpage
    452
  • Lastpage
    455
  • Abstract
    This paper describes a new unsupervised approach for identifying the main themes of any text document with the aid of Wikipedia. In contrast to others, the proposed algorithm relies on merely two main aspects of Wikipedia, namely its articles titles and categories structure. The inner content of the articles of Wikipedia are not employed in our algorithm. We describe in this paper how to build a Term-Categories vector that defines how strong a term is associated to a Wikipedia concept. We also explain how this vector is employed when processing a text document to discover its main themes. We report the performance of our method by attempting to predict the most representative categories for a subset of Wikipedia articles.
  • Keywords
    "Encyclopedias","Electronic publishing","Internet","Ontologies","Data mining","Prediction algorithms"
  • Publisher
    ieee
  • Conference_Titel
    Intelligent Human-Machine Systems and Cybernetics (IHMSC), 2015 7th International Conference on
  • Print_ISBN
    978-1-4799-8645-3
  • Type

    conf

  • DOI
    10.1109/IHMSC.2015.68
  • Filename
    7334744