DocumentCode
2308274
Title
CUTER: An Efficient Useful Text Extraction Mechanism
Author
Adam, George ; Bouras, Christos ; Poulopoulos, Vassilis
Author_Institution
Res. Acad. Comput. Technol. Inst., Patras
fYear
2009
fDate
26-29 May 2009
Firstpage
703
Lastpage
708
Abstract
In this paper we present CUTER, a system that processes HTML pages in order to extract the useful text from them. The mechanism is focalized on HTML pages that include news articles from major portals and blogs. As useful text we define the body of the article that contains the news report. In order to extract the body of the article we deconstruct the HTML page to its DOM model and we apply a set of algorithms in order to clean and correct the HTML code, locate and characterize each node of the DOM model and finally store the text from the nodes that are characterized as useful text nodes. CUTER is a subsystem of peRSSonal, a Web tool that is used to obtain news articles from all over the world, process them and present them back to the end users in a personalized manner. The role of CUTER is to feed peRSSonal with the body of the. In this paper we present the basic algorithms and experimental results on the efficiency of the CUTER text extractor.
Keywords
Web sites; document handling; hypermedia markup languages; portals; text analysis; HTML page; Web portal; document object model; personalized manner; text extraction mechanism; Application software; Blogs; Computer networks; Data mining; Feeds; HTML; Informatics; Information analysis; Portals; Web pages; DOM analysis; HTML analysis; Text extraction; useful text;
fLanguage
English
Publisher
ieee
Conference_Titel
Advanced Information Networking and Applications Workshops, 2009. WAINA '09. International Conference on
Conference_Location
Bradford
Print_ISBN
978-1-4244-3999-7
Electronic_ISBN
978-0-7695-3639-2
Type
conf
DOI
10.1109/WAINA.2009.60
Filename
5136731
Link To Document