DocumentCode
1940826
Title
PhishDef: URL names say it all
Author
Le, Anh ; Markopoulou, Athina ; Faloutsos, Michalis
Author_Institution
Univ. of California, Irvine, CA, USA
fYear
2011
fDate
10-15 April 2011
Firstpage
191
Lastpage
195
Abstract
Phishing is an increasingly sophisticated method to steal personal user information using sites that pretend to be legitimate. In this paper, we take the following steps to identify phishing URLs. First, we carefully select lexical features of the URLs that are resistant to obfuscation techniques used by attackers. Second, we evaluate the classification accuracy when using only lexical features, both automatically and hand-selected, vs. when using additional features. We show that lexical features are sufficient for all practical purposes. Third, we thoroughly compare several classification algorithms, and we propose to use an online method (AROW) that is able to overcome noisy training data. Based on the insights gained from our analysis, we propose PhishDef, a phishing detection system that uses only URL names and combines the above three elements. PhishDef is a highly accurate method (when compared to state-of-the-art approaches over real datasets), lightweight (thus appropriate for online and client-side deployment), proactive (based on online classification rather than blacklists), and resilient to training data inaccuracies (thus enabling the use of large noisy training data).
Keywords
Web sites; computer crime; AROW; PhishDef; URL names; classification accuracy; lexical features; obfuscation technique resilience; online classification; online method; personal user information stealing; phishing detection system; Accuracy; Classification algorithms; Error analysis; Feature extraction; Noise; Prediction algorithms; Servers;
fLanguage
English
Publisher
ieee
Conference_Titel
INFOCOM, 2011 Proceedings IEEE
Conference_Location
Shanghai
ISSN
0743-166X
Print_ISBN
978-1-4244-9919-9
Type
conf
DOI
10.1109/INFCOM.2011.5934995
Filename
5934995
Link To Document