000 03856nam a2200265 a 4500
001 vtls000072570
003 KUKTEM
005 20251114204537.0
008 130621t2012 my a f m 000 0 eng d
020 _aTHE0002015(Local)
039 9 _a201905131604
_byusri
_c201710121547
_daishah
_c201306210939
_dFida
_y201306210937
_zFida
040 _aUMP
090 _aQA278 .Q56 2012 rs Thesis
100 0 _aQin Hongwu
245 1 0 _aThe new efficient and accurate attribute-oriented clustering algorithms for categorical data /
_cQin Hongwu
260 _aKuantan, Pahang :
_bUMP,
_c2012
300 _axix, 164 p. :
_bill. (some col.) ;
_c30 cm. +
_e1 CD-ROM
502 _aThesis (Doctor of Philosophy in Computer Science) -- Universiti Malaysia Pahang - 2012
504 _aBibliography: p. 133-138
520 3 _aCategorical data clustering has attracted much attention recently due to the fact that much of the data contained in today’s databases is categorical in nature. Many algorithms for clustering categorical data have been proposed, in which attribute-oriented hierarchical divisive clustering algorithm Min-Min Roughness (MMR) has the highest efficiency among these algorithms with low clustering accuracy, conversely, genetic clustering algorithm Genetic-Average Normalized Mutual Information (G-ANMI) has the highest clustering accuracy among these algorithms with low clustering efficiency. This work firstly reveals the significance of attributes in categorical data clustering, and then investigates the limitations of algorithms MMR and G-ANMI respectively, and correspondingly proposes a new attribute-oriented hierarchical divisive clustering algorithm termed Mean Gain Ratio (MGR) and an improved genetic clustering algorithm termed Improved G-ANMI (IG-ANMI) for categorical data. MGR includes two steps: selecting clustering attribute and selecting equivalence class on the clustering attribute. Information theory based concepts of mean gain ratio and entropy of clusters are used to implement these two steps, respectively. MGR can be run with or without specifying the number of clusters while few existing clustering algorithms for categorical data can be run without specifying the number of clusters. IG-ANMI algorithm improves G-ANMI by developing a new attribute-oriented initialization method in which part of initial chromosomes is generated by using the attributes partitions. Four real-life data sets obtained from University of California Irvine (UCI) machine learning repository and ten synthetically generated data sets are used to evaluate MGR and IG-ANMI algorithms, and other four algorithms are used to compare with these two algorithms. The experimental results show that MGR overcomes the limitations of MMR and the average clustering accuracy is improved by 19% (from 0.696 to 0.83), at the same time maintains the highest efficiency. IG-ANMI greatly improves the efficiency of G-ANMI (improved by 31% on the Zoo data set, 74% on the Votes data set, 59% on the Breast Cancer data set, and 3428% on the Mushroom data set) as well as the clustering accuracy of G-ANMI (the average clustering accuracy on four UCI data sets is improved by 10.6%, from 0.815 to 0.901), at the same time maintains the highest clustering accuracy. IG-ANMI has obvious advantage against G-ANMI on large data sets in terms of clustering efficiency as well as clustering accuracy. In addition, both of MGR and IG-ANMI have good scalability. The running time of MGR and IG-ANMI algorithms tend to vary linearly with the increase of the number of objects as well as the number of clusters.
650 0 _aCluster analysis
650 0 _aCluster analysis
_xData processing
856 4 0 _uhttp://ecollib.ump.edu.my/24671/
_zLibrary access only
999 _aVIRTUA40
_c3866
_d3872
999 _aVTLSSORT0080*0200*0400*0900*1000*2450*2600*3000*5020*5040*5200*6500*6501*8560*9992