DocumentCode
2429712
Title
A strategy for identification of Web query interfaces using supervised learning
Author
Marin-Castro, Heidy M. ; Sosa-Sosa, Victor J. ; Lopez-Arevalo, Ivan
Author_Institution
Inf. Technol. Lab., Nat. Polytech. Inst., Tamaulipas, Mexico
fYear
2011
fDate
19-21 Oct. 2011
Firstpage
233
Lastpage
237
Abstract
The Deep Web is an enormous source of information constantly growing. It comprises a large amount of databases on the Web that are accessed through Web query interfaces related to different domains. The content of the Deep Web can not be reachable by traditional search engines, what makes almost impossible for common users to get access to this useful information. There are several problems related to the search for content in the Deep Web. One of them is the automatic identification of Web query interfaces, being this a mean to access the information in the Deep Web. The task of classify HTML forms contained inside Web page as Web query interface is challenging due to their enormous heterogeneity. This paper introduce a strategy that automatically identifies Web query interfaces independent of their domain. We make an adequate selection of HTML elements and use them appropriately to build characteristic vectors that are used as input of a supervised classifier to determine if a Web page contains or not a Web query interface. The experimental results show that the proposed strategy is efficient and accurate, achieving better classification results than works previously reported.
Keywords
Internet; learning (artificial intelligence); query processing; HTML; Web query interface; deep Web; supervised classifier; supervised learning; Accuracy; Complexity theory; Databases; Feature extraction; HTML; Support vector machines; Web pages; Databases on the Web; Deep Web; Web query interfaces; supervised learning;
fLanguage
English
Publisher
ieee
Conference_Titel
Next Generation Web Services Practices (NWeSP), 2011 7th International Conference on
Conference_Location
Salamanca
Print_ISBN
978-1-4577-1125-1
Type
conf
DOI
10.1109/NWeSP.2011.6088183
Filename
6088183
Link To Document