• Title of article

    Bayesian network models for hierarchical text classification from a thesaurus Original Research Article

  • Author/Authors

    Luis M. de Campos، نويسنده , , Alfonso E. Romero، نويسنده ,

  • Issue Information
    روزنامه با شماره پیاپی سال 2009
  • Pages
    13
  • From page
    932
  • To page
    944
  • Abstract
    We propose a method which, given a document to be classified, automatically generates an ordered set of appropriate descriptors extracted from a thesaurus. The method creates a Bayesian network to model the thesaurus and uses probabilistic inference to select the set of descriptors having high posterior probability of being relevant given the available evidence (the document to be classified). Our model can be used without having preclassified training documents, although it improves its performance as long as more training data become available. We have tested the classification model using a document dataset containing parliamentary resolutions from the regional Parliament of Andalucía at Spain, which were manually indexed from the Eurovoc thesaurus, also carrying out an experimental comparison with other standard text classifiers.
  • Keywords
    thesauri , Bayesian networks , hierarchical classification , Document categorization
  • Journal title
    International Journal of Approximate Reasoning
  • Serial Year
    2009
  • Journal title
    International Journal of Approximate Reasoning
  • Record number

    1182723