DocumentCode
3125598
Title
Understanding Propagation Error and Its Effect on Collective Classification
Author
Xiang, Rongjing ; Neville, Jennifer
Author_Institution
Dept. of Comput. Sci., Purdue Univ., West Lafayette, IN, USA
fYear
2011
fDate
11-14 Dec. 2011
Firstpage
834
Lastpage
843
Abstract
Recent empirical evaluation has shown that the performance of collective classification models can vary based on the amount of class label information available for use during inference. In this paper, we further demonstrate that the relative performance of statistical relational models learned with different estimation methods changes as the availability of test set labels increases. We reason about the cause of this phenomenon from an information-theoretic perspective and this points to a previously unidentified consideration in the development of relational learning algorithms. In particular, we characterize the high propagation error of collective inference models that are estimated with maximum pseudolikelihood estimation (MPLE), and show how this affects performance across the spectrum of label availability when compared to MLE, which has low propagation error. Our formal study leads to a quantitative characterization that can be used to predict the confidence of local propagation for MPLE models. We use this to propose a mixture model that can learn the best trade-off between high and low propagation models. Empirical evaluation on synthetic and real-world data show that our proposed method achieves comparable, or superior, results to both MPLE and low propagation models across the full spectrum of label availability.
Keywords
inference mechanisms; information theory; learning (artificial intelligence); maximum likelihood estimation; pattern classification; class label information; collective classification models; empirical evaluation; inference; information theoretic perspective; label availability; maximum pseudolikelihood estimation; mixture model; propagation error; relational learning algorithms; test set labels; Availability; Computational modeling; Data models; Maximum likelihood estimation; Microscopy; Probabilistic logic; Training; collective classification; probabilistic relational models; statistical relational learning;
fLanguage
English
Publisher
ieee
Conference_Titel
Data Mining (ICDM), 2011 IEEE 11th International Conference on
Conference_Location
Vancouver,BC
ISSN
1550-4786
Print_ISBN
978-1-4577-2075-8
Type
conf
DOI
10.1109/ICDM.2011.151
Filename
6137288
Link To Document