Title :
A mining-based category evolution approach to managing online document categories
Author :
Wei, Chih-Ping ; Dong, Yuan-Xin
Author_Institution :
Dept. of Inf. Manage., Nat. Sun Yat-Sen Univ., Kaohsiung, Taiwan
Abstract :
With rapid expansion of the numbers and sizes of text repositories and improvements in global connectivity, the quantity of information available online as free-format text is growing exponentially. Many large organizations create and maintain huge volumes of textual information online, and there is a pressing need for support of efficient and effective information retrieval, filtering, and management. Text categorization, or the assignment of textual documents to one or more pre-defined categories based on their content, is an essential component of efficient management and retrieval of documents. Previously, research has focused predominantly on developing or adopting statistical classification or inductive learning methods for automatically discovering text categorization patterns for a pre-defined set of categories. However, as documents accumulate, such categories may not capture a document´s characteristics correctly. In this study, we propose a mining-based category evolution (MiCE) technique to adjust document categories based on existing categories and their associated documents. Empirical evaluation results indicate that the proposed technique, MiCE, was more effective than the category discovery approach and was insensitive to the quality of original categories.
Keywords :
classification; data mining; document handling; information retrieval; learning by example; MiCE; free-format text; global connectivity; inductive learning methods; information retrieval; mining-based category evolution approach; online document categories management; statistical classification; text repositories; Content based retrieval; Content management; Information filtering; Information filters; Information management; Information retrieval; Learning systems; Mice; Pressing; Text categorization;
Conference_Titel :
System Sciences, 2001. Proceedings of the 34th Annual Hawaii International Conference on
Conference_Location :
Maui, HI, USA
Print_ISBN :
0-7695-0981-9
DOI :
10.1109/HICSS.2001.927093