DocumentCode :
20693
Title :
On the Use of Genetic Programming for Mining Comprehensible Rules in Subgroup Discovery
Author :
Luna, Jose Marcio ; Romero, Jose Raul ; Romero, C. ; Ventura, Sebastian
Author_Institution :
Dept. of Comput. Sci. & Numerical Anal., Univ. of Cordoba, Cordoba, Spain
Volume :
44
Issue :
12
fYear :
2014
fDate :
Dec. 2014
Firstpage :
2329
Lastpage :
2341
Abstract :
This paper proposes a novel grammar-guided genetic programming algorithm for subgroup discovery. This algorithm, called comprehensible grammar-based algorithm for subgroup discovery (CGBA-SD), combines the requirements of discovering comprehensible rules with the ability to mine expressive and flexible solutions owing to the use of a context-free grammar. Each rule is represented as a derivation tree that shows a solution described using the language denoted by the grammar. The algorithm includes mechanisms to adapt the diversity of the population by self-adapting the probabilities of recombination and mutation. We compare the approach with existing evolutionary and classic subgroup discovery algorithms. CGBA-SD appears to be a very promising algorithm that discovers comprehensible subgroups and behaves better than other algorithms as measures by complexity, interest, and precision indicate. The results obtained were validated by means of a series of nonparametric tests.
Keywords :
context-free grammars; data mining; genetic algorithms; probability; statistical testing; comprehensible grammar-based algorithm for subgroup discovery; comprehensible rule mining; comprehensible subgroup discovery algorithms; context-free grammar; grammar-guided genetic programming algorithm; mutation probabilities; nonparametric tests; recombination probabilities; Complexity theory; Fuzzy systems; Genetic programming; Grammar; Production; Sociology; Statistics; Data mining (DM); genetic programming (GP); grammar-guided genetic programming (G3P); subgroup discovery (SD);
fLanguage :
English
Journal_Title :
Cybernetics, IEEE Transactions on
Publisher :
ieee
ISSN :
2168-2267
Type :
jour
DOI :
10.1109/TCYB.2014.2306819
Filename :
6756991
Link To Document :
بازگشت