• DocumentCode
    1180289
  • Title

    Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA

  • Author

    Zhou, Ke ; Xue Gui-Rong ; Yang, Qiang ; Yu, Yong

  • Author_Institution
    Dept. of Comput. Sci. & Eng., Shanghai Jiao-Tong Univ., Shanghai, China
  • Volume
    22
  • Issue
    1
  • fYear
    2010
  • Firstpage
    46
  • Lastpage
    58
  • Abstract
    It is often difficult and time-consuming to provide a large amount of positive and negative examples for training a classification system in many applications such as information retrieval. Instead, users often find it easier to indicate just a few positive examples of what he or she likes, and thus, these are the only labeled examples available for the learning system. A large amount of unlabeled data are easier to obtain. How to make use of the positive and unlabeled data for learning is a critical problem in machine learning and information retrieval. Several approaches for solving this problem have been proposed in the past, but most of these methods do not work well when only a small amount of labeled positive data are available. In this paper, we propose a novel algorithm called Topic-Sensitive pLSA to solve this problem. This algorithm extends the original probabilistic latent semantic analysis (pLSA), which is a purely unsupervised framework, by injecting a small amount of supervision information from the user. The supervision from users is in the form of indicating which documents fit the users´ interests. The supervision is encoded into a set of constraints. By introducing the penalty terms for these constraints, we propose an objective function that trades off the likelihood of the observed data and the enforcement of the constraints. We develop an iterative algorithm that can obtain the local optimum of the objective function. Experimental evaluation on three data corpora shows that the proposed method can improve the performance especially only with a small amount of labeled positive data.
  • Keywords
    classification; data handling; information retrieval; iterative methods; learning systems; probability; unsupervised learning; classification system; data corpora; information retrieval; iterative algorithm; learning system; machine learning; probabilistic latent semantic analysis; supervision information; topic-sensitive PLSA; topic-sensitive pLSA; unsupervised framework; Semisupervised learning; document classification.; topic-sensitive probabilistic latent semantic analysis;
  • fLanguage
    English
  • Journal_Title
    Knowledge and Data Engineering, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1041-4347
  • Type

    jour

  • DOI
    10.1109/TKDE.2009.56
  • Filename
    4796195