Title :
Extraction of Relevant Snippets from Web Pages Using Hybrid Features
Author :
Zeng, Jun ; Xiong, Qingyu ; Wen, Junhao ; Hirokawa, Sachio
Author_Institution :
Grad. Sch. of Inf. Sci. & Electr. Eng., Kyushu Univ., Fukuoka, Japan
Abstract :
As the amount of web pages increase, identifying and retrieving distinct contents from the web has increasingly become more and more difficult. The traditional approach for extracting data from web page documents is to analyze the DOM (Document Object Model) structure of a HTML page and find a common pattern. However, the number of possible DOM layout patterns is virtually infinite, which means that there is no common pattern that can be used for all kinds of web pages. In this paper, we focus on the pages that are linked to a search engine and aim to analyze the features of relevant and meaningful contents instead of a common pattern. Three features of relevant snippets are introduced. They are: quantity of text, correlation between snippet and query that is inputted into a search engine, and HTML structure. Nine parameters are used to describe the three features. Also, a SVM learning experiment is conducted to verify the effectiveness of the three features. The results show that the HTML structure feature is the most effective feature which can determine whether a snippet is relevant or not.
Keywords :
Internet; document handling; hypermedia markup languages; object-oriented programming; support vector machines; DOM layout patterns; DOM structure; HTML page; HTML structure feature; SVM learning experiment; Web page documents; Web pages; common pattern; document object model; hybrid features; relevant snippets extraction; search engine; Correlation; Diseases; HTML; Search engines; Silicon; Support vector machines; Web pages; SVM learning; content extraction; hybrid feature; relevant snippets;
Conference_Titel :
Advanced Applied Informatics (IIAIAAI), 2012 IIAI International Conference on
Conference_Location :
Fukuoka
Print_ISBN :
978-1-4673-2719-0
DOI :
10.1109/IIAI-AAI.2012.50