• DocumentCode
    2861707
  • Title

    A Weighted Freshness Metric for Maintaining Search Engine Local Repository

  • Author

    Han, Jianchao ; Cercone, Nick ; Hu, Xiaohua

  • Author_Institution
    California State University, Dominguez Hills
  • fYear
    2004
  • fDate
    20-24 Sept. 2004
  • Firstpage
    677
  • Lastpage
    680
  • Abstract
    Current search engines maintain a local repository to improve the search efficiency. A crawler is used to periodically poll the remote web pages to update the contents of the local repository. Due to the resource limitations, some local pages may be stale. To maintain the high freshness of the repository, the crawler is expected to revisit remote web pages in optimized order and frequency. The intuitive metric of freshness of the local repository is defined as the fraction of up-to-date web pages in the repository, which is merely based on the repository content, and does not, unfortunately, reflect the perspective of the search engine users, e.g., how often is a web page queried? We propose a novel weighted metric of the repository freshness with the importance of web pages being the weights. This metric not only takes into account the local web pages themselves but also the perspectives of the search engine users. We study the repository synchronization policy under this new metric, compare this metric with others, analyze its features, and discuss how the web page importance is determined.
  • Keywords
    Computer science; Crawlers; Educational institutions; Frequency synchronization; Information science; Navigation; Search engines; Web pages; Web search; Web server;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Web Intelligence, 2004. WI 2004. Proceedings. IEEE/WIC/ACM International Conference on
  • Print_ISBN
    0-7695-2100-2
  • Type

    conf

  • DOI
    10.1109/WI.2004.10071
  • Filename
    1410895