• Title of article

    Affiliation disambiguation for constructing semantic digital libraries

  • Author/Authors

    Yong Jiang1، نويسنده , , Hai-Tao Zheng1، نويسنده , , Xinmin Wang1، نويسنده , , Binggan Lu2، نويسنده , , Kaihua Wu2، نويسنده ,

  • Issue Information
    ماهنامه با شماره پیاپی سال 2011
  • Pages
    13
  • From page
    1029
  • To page
    1041
  • Abstract
    With increasing digital information availability, semantic web technologies have been employed to construct semantic digital libraries in order to ease information comprehension. The use of semantic web enables users to search or visualize resources in a semantic fashion. Semantic web generation is a key process in semantic digital library construction, which converts metadata of digital resources into semantic web data. Many text mining technologies, such as keyword extraction and clustering, have been proposed to generate semantic web data. However, one important type of metadata in publications, called affiliation, is hard to convert into semantic web data precisely because different authors, who have the same affiliation, often express the affiliation in different ways. To address this issue, this paper proposes a clustering method based on normalized compression distance for the purpose of affiliation disambiguation. The experimental results show that our method is able to identify different affiliations that denote the same institutes. The clustering results outperform the well-known k-means clustering method in terms of average precision, F-measure, entropy, and purity.
  • Journal title
    Journal of the American Society for Information Science and Technology
  • Serial Year
    2011
  • Journal title
    Journal of the American Society for Information Science and Technology
  • Record number

    994443