DocumentCode
2360019
Title
Syntactic similarity of Web documents
Author
Pereira, Álvaro R., Jr. ; Ziviani, Nivio
Author_Institution
Dept. of Comput. Sci., Fed. Univ. of Minas Gerais, Belo Horizonte, Brazil
fYear
2003
fDate
10-12 Nov. 2003
Firstpage
194
Lastpage
200
Abstract
We present and compare two methods for evaluating the syntactic similarity between documents. The first method uses the Patricia tree, constructed from the original document, and the similarity is computed searching the text of each candidate document in the tree. The second method uses shingles concept to obtain the similarity measure for every document pairs, and each shingle from the original document is inserted in a hash table, where shingles of each candidate document are searched. Given an original document and some candidates, two methods find documents that have some similarity relationship with the original document. Experimental results were obtained by using a plagiarized documents generator system, from 900 documents collected from the Web. Considering the arithmetic average of the absolute differences between the expected and obtained similarity, the algorithm that uses shingles obtained a performance of 4.13% and the algorithm that uses Patricia tree a performance of 7.50%.
Keywords
Internet; computational complexity; document handling; search engines; tree data structures; tree searching; Patricia tree; Web document syntactic similarity evaluation; candidate document; hash table; plagiarized documents generator system; shingles concept; tree searching; Arithmetic; Computer science; Fingerprint recognition; Indexing; Information retrieval; Metasearch; Plagiarism; Search engines; World Wide Web;
fLanguage
English
Publisher
ieee
Conference_Titel
Web Congress, 2003. Proceedings. First Latin American
Print_ISBN
0-7695-2058-8
Type
conf
DOI
10.1109/LAWEB.2003.1250297
Filename
1250297
Link To Document