• DocumentCode
    2727683
  • Title

    Extracting Relevant Snippets from Web Documents through Language Model based Text Segmentation

  • Author

    Li, Qing ; Candan, K. Selçuk ; Qi, Yan

  • fYear
    2007
  • fDate
    2-5 Nov. 2007
  • Firstpage
    287
  • Lastpage
    290
  • Abstract
    Extracting a query-oriented snippet (or passage) and highlighting the relevant information in long document can help reduce the result navigation cost of end users. While the traditional approach of highlighting matching keywords helps when the search is keyword oriented, finding appropriate snippets to represent matches to more complex queries requires novel techniques that can help characterize the relevance of various parts of a document to the given query, succinctly. In this paper, we present a languagemodel based method for accurately detecting the most relevant passages of a given document. Unlike previous works in passage retrieval which focus on searching relevance nodes for filtering of preoccupied passages, we focus on query-informed segmentation for snippet extraction. The algorithms presented in this paper are currently being deployed in OASIS, a system to help reduce the navigational load of blind users in accessing Web-based digital libraries.
  • Keywords
    Costs; Data mining; Filtering; History; Indexing; Navigation; Search engines; Software libraries; Visualization; Web pages;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Web Intelligence, IEEE/WIC/ACM International Conference on
  • Conference_Location
    Fremont, CA
  • Print_ISBN
    978-0-7695-3026-0
  • Type

    conf

  • DOI
    10.1109/WI.2007.115
  • Filename
    4427103