• DocumentCode
    2429712
  • Title

    A strategy for identification of Web query interfaces using supervised learning

  • Author

    Marin-Castro, Heidy M. ; Sosa-Sosa, Victor J. ; Lopez-Arevalo, Ivan

  • Author_Institution
    Inf. Technol. Lab., Nat. Polytech. Inst., Tamaulipas, Mexico
  • fYear
    2011
  • fDate
    19-21 Oct. 2011
  • Firstpage
    233
  • Lastpage
    237
  • Abstract
    The Deep Web is an enormous source of information constantly growing. It comprises a large amount of databases on the Web that are accessed through Web query interfaces related to different domains. The content of the Deep Web can not be reachable by traditional search engines, what makes almost impossible for common users to get access to this useful information. There are several problems related to the search for content in the Deep Web. One of them is the automatic identification of Web query interfaces, being this a mean to access the information in the Deep Web. The task of classify HTML forms contained inside Web page as Web query interface is challenging due to their enormous heterogeneity. This paper introduce a strategy that automatically identifies Web query interfaces independent of their domain. We make an adequate selection of HTML elements and use them appropriately to build characteristic vectors that are used as input of a supervised classifier to determine if a Web page contains or not a Web query interface. The experimental results show that the proposed strategy is efficient and accurate, achieving better classification results than works previously reported.
  • Keywords
    Internet; learning (artificial intelligence); query processing; HTML; Web query interface; deep Web; supervised classifier; supervised learning; Accuracy; Complexity theory; Databases; Feature extraction; HTML; Support vector machines; Web pages; Databases on the Web; Deep Web; Web query interfaces; supervised learning;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Next Generation Web Services Practices (NWeSP), 2011 7th International Conference on
  • Conference_Location
    Salamanca
  • Print_ISBN
    978-1-4577-1125-1
  • Type

    conf

  • DOI
    10.1109/NWeSP.2011.6088183
  • Filename
    6088183