مرکز منطقه ای اطلاع رساني علوم و فناوري - Intelligent focused crawler: Learning which links to crawl

DocumentCode :

3642097

Title :

Intelligent focused crawler: Learning which links to crawl

Author :

Duygu Taylan;Mitat Poyraz;Selim Akyokuş;Murat Can Ganiz

Author_Institution :

Department of Computer Engineering, Doğ

fYear :

2011

fDate :

6/1/2011 12:00:00 AM

Firstpage :

504

Lastpage :

508

Abstract :

A web crawler is defined as an automated program that methodically scans through Internet pages and downloads any page that can be reached via links. With the exponential growth of the Web, fetching information about a special-topic is gaining importance. A focused crawler is a web crawler that attempts to download only web pages that are relevant to a predefined topic or set of topics. In order to determine a web page is about a particular topic, focused crawlers use classification techniques. In this study we focus on the classification of links instead of downloaded web pages to determine relevancy. We combine a Naïve Bayes classifier for classification of URLs with a simple URL scoring optimization to improve the system performance. Our results demonstrate that proposed approach performs better.

Keywords :

"Crawlers","Niobium","Optimization","Classification algorithms","Mathematical model","Web pages","Equations"

Publisher :

ieee

Conference_Titel :

Innovations in Intelligent Systems and Applications (INISTA), 2011 International Symposium on

Print_ISBN :

978-1-61284-919-5

Type :

conf

DOI :

10.1109/INISTA.2011.5946150

Filename :

5946150

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=3642097