• DocumentCode
    1273460
  • Title

    Induction by attribute elimination

  • Author

    Wu, Xindong ; Urpani, David

  • Author_Institution
    Dept. of Math. & Comput. Sci., Colorado Sch. of Mines, Golden, CO, USA
  • Volume
    11
  • Issue
    5
  • fYear
    1999
  • Firstpage
    805
  • Lastpage
    812
  • Abstract
    In most data mining applications where induction is used as the primary tool for knowledge extraction from real world databases, it is difficult to precisely identify a complete set of relevant attributes. The paper introduces a novel rule induction algorithm called Rule Induction Two In One (RITIO), which eliminates attributes in the order of decreasing irrelevancy. Like ID3-like decision tree construction algorithms, RITIO makes use of the entropy measure as a means of constraining the hypothesis search space; but, unlike IDS-like algorithms, the hypotheses language is the rule structure and RITIO generates rules without constructing decision trees. The final concept description produced by RITIO is shown to be largely based on only the most relevant attributes. Experimental results confirm that, even on noisy, industrial databases, RITIO achieves high levels of predictive accuracy
  • Keywords
    data mining; entropy; inference mechanisms; search problems; ID3-like decision tree construction algorithms; RITIO; Rule Induction Two In One; attribute elimination; concept description; data mining applications; entropy measure; hypotheses language; hypothesis search space; industrial databases; knowledge extraction; predictive accuracy; real world databases; relevant attributes; rule induction algorithm; rule structure; Accuracy; Cleaning; Databases; Decision trees; Humans; Induction generators; Information entropy; Noise level; Research and development; Senior members;
  • fLanguage
    English
  • Journal_Title
    Knowledge and Data Engineering, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1041-4347
  • Type

    jour

  • DOI
    10.1109/69.806938
  • Filename
    806938