• DocumentCode
    3125598
  • Title

    Understanding Propagation Error and Its Effect on Collective Classification

  • Author

    Xiang, Rongjing ; Neville, Jennifer

  • Author_Institution
    Dept. of Comput. Sci., Purdue Univ., West Lafayette, IN, USA
  • fYear
    2011
  • fDate
    11-14 Dec. 2011
  • Firstpage
    834
  • Lastpage
    843
  • Abstract
    Recent empirical evaluation has shown that the performance of collective classification models can vary based on the amount of class label information available for use during inference. In this paper, we further demonstrate that the relative performance of statistical relational models learned with different estimation methods changes as the availability of test set labels increases. We reason about the cause of this phenomenon from an information-theoretic perspective and this points to a previously unidentified consideration in the development of relational learning algorithms. In particular, we characterize the high propagation error of collective inference models that are estimated with maximum pseudolikelihood estimation (MPLE), and show how this affects performance across the spectrum of label availability when compared to MLE, which has low propagation error. Our formal study leads to a quantitative characterization that can be used to predict the confidence of local propagation for MPLE models. We use this to propose a mixture model that can learn the best trade-off between high and low propagation models. Empirical evaluation on synthetic and real-world data show that our proposed method achieves comparable, or superior, results to both MPLE and low propagation models across the full spectrum of label availability.
  • Keywords
    inference mechanisms; information theory; learning (artificial intelligence); maximum likelihood estimation; pattern classification; class label information; collective classification models; empirical evaluation; inference; information theoretic perspective; label availability; maximum pseudolikelihood estimation; mixture model; propagation error; relational learning algorithms; test set labels; Availability; Computational modeling; Data models; Maximum likelihood estimation; Microscopy; Probabilistic logic; Training; collective classification; probabilistic relational models; statistical relational learning;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Mining (ICDM), 2011 IEEE 11th International Conference on
  • Conference_Location
    Vancouver,BC
  • ISSN
    1550-4786
  • Print_ISBN
    978-1-4577-2075-8
  • Type

    conf

  • DOI
    10.1109/ICDM.2011.151
  • Filename
    6137288