<?xml version="1.0" encoding="UTF-8"?>
<record
    xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
    xsi:schemaLocation="http://www.loc.gov/MARC21/slim http://www.loc.gov/standards/marcxml/schema/MARC21slim.xsd"
    xmlns="http://www.loc.gov/MARC21/slim">

  <leader>03856nam a2200265 a 4500</leader>
  <controlfield tag="001">vtls000072570</controlfield>
  <controlfield tag="003">KUKTEM</controlfield>
  <controlfield tag="005">20251114204537.0</controlfield>
  <controlfield tag="008">130621t2012    my a   f m    000 0 eng d</controlfield>
  <datafield tag="020" ind1=" " ind2=" ">
    <subfield code="a">THE0002015(Local)</subfield>
  </datafield>
  <datafield tag="039" ind1=" " ind2="9">
    <subfield code="a">201905131604</subfield>
    <subfield code="b">yusri</subfield>
    <subfield code="c">201710121547</subfield>
    <subfield code="d">aishah</subfield>
    <subfield code="c">201306210939</subfield>
    <subfield code="d">Fida</subfield>
    <subfield code="y">201306210937</subfield>
    <subfield code="z">Fida</subfield>
  </datafield>
  <datafield tag="040" ind1=" " ind2=" ">
    <subfield code="a">UMP</subfield>
  </datafield>
  <datafield tag="090" ind1=" " ind2=" ">
    <subfield code="a">QA278 .Q56 2012 rs Thesis</subfield>
  </datafield>
  <datafield tag="100" ind1="0" ind2=" ">
    <subfield code="a">Qin Hongwu</subfield>
  </datafield>
  <datafield tag="245" ind1="1" ind2="0">
    <subfield code="a">The new efficient and accurate attribute-oriented clustering algorithms for categorical data /</subfield>
    <subfield code="c">Qin Hongwu</subfield>
  </datafield>
  <datafield tag="260" ind1=" " ind2=" ">
    <subfield code="a">Kuantan, Pahang :</subfield>
    <subfield code="b">UMP,</subfield>
    <subfield code="c">2012</subfield>
  </datafield>
  <datafield tag="300" ind1=" " ind2=" ">
    <subfield code="a">xix, 164 p. :</subfield>
    <subfield code="b">ill. (some col.) ;</subfield>
    <subfield code="c">30 cm. +</subfield>
    <subfield code="e">1 CD-ROM</subfield>
  </datafield>
  <datafield tag="502" ind1=" " ind2=" ">
    <subfield code="a">Thesis (Doctor of Philosophy in Computer Science)  -- Universiti Malaysia Pahang - 2012</subfield>
  </datafield>
  <datafield tag="504" ind1=" " ind2=" ">
    <subfield code="a">Bibliography: p. 133-138</subfield>
  </datafield>
  <datafield tag="520" ind1="3" ind2=" ">
    <subfield code="a">Categorical data clustering has attracted much attention recently due to the fact that much of the data contained in today&#x2019;s databases is categorical in nature. Many algorithms for clustering categorical data have been proposed, in which attribute-oriented hierarchical divisive clustering algorithm Min-Min Roughness (MMR) has the highest efficiency among these algorithms with low clustering accuracy, conversely, genetic clustering algorithm Genetic-Average Normalized Mutual Information (G-ANMI) has the highest clustering accuracy among these algorithms with low clustering efficiency. This work firstly reveals the significance of attributes in categorical data clustering, and then investigates the limitations of algorithms MMR and G-ANMI respectively, and correspondingly proposes a new attribute-oriented hierarchical divisive clustering algorithm termed Mean Gain Ratio (MGR) and an improved genetic clustering algorithm termed Improved G-ANMI (IG-ANMI) for categorical data. MGR includes two steps: selecting clustering attribute and selecting equivalence class on the clustering attribute. Information theory based concepts of mean gain ratio and entropy of clusters are used to implement these two steps, respectively. MGR can be run with or without specifying the number of clusters while few existing clustering algorithms for categorical data can be run without specifying the number of clusters. IG-ANMI algorithm improves G-ANMI by developing a new attribute-oriented initialization method in which part of initial chromosomes is generated by using the attributes partitions. Four real-life data sets obtained from University of California Irvine (UCI) machine learning repository and ten synthetically generated data sets are used to evaluate MGR and IG-ANMI algorithms, and other four algorithms are used to compare with these two algorithms. The experimental results show that MGR overcomes the limitations of MMR and the average clustering accuracy is improved by 19% (from 0.696 to 0.83), at the same time maintains the highest efficiency. IG-ANMI greatly improves the efficiency of G-ANMI (improved by 31% on the Zoo data set, 74% on the Votes data set, 59% on the Breast Cancer data set, and 3428% on the Mushroom data set) as well as the clustering accuracy of G-ANMI (the average clustering accuracy on four UCI data sets is improved by 10.6%, from 0.815 to 0.901), at the same time maintains the highest clustering accuracy. IG-ANMI has obvious advantage against G-ANMI on large data sets in terms of clustering efficiency as well as clustering accuracy. In addition, both of MGR and IG-ANMI have good scalability. The running time of MGR and IG-ANMI algorithms tend to vary linearly with the increase of the number of objects as well as the number of clusters.</subfield>
  </datafield>
  <datafield tag="650" ind1=" " ind2="0">
    <subfield code="a">Cluster analysis</subfield>
  </datafield>
  <datafield tag="650" ind1=" " ind2="0">
    <subfield code="a">Cluster analysis</subfield>
    <subfield code="x">Data processing</subfield>
  </datafield>
  <datafield tag="856" ind1="4" ind2="0">
    <subfield code="u">http://ecollib.ump.edu.my/24671/</subfield>
    <subfield code="z">Library access only</subfield>
  </datafield>
  <datafield tag="952" ind1=" " ind2=" ">
    <subfield code="0">0</subfield>
    <subfield code="1">0</subfield>
    <subfield code="2">lcc</subfield>
    <subfield code="4">0</subfield>
    <subfield code="7">1</subfield>
    <subfield code="a">10000</subfield>
    <subfield code="b">10000</subfield>
    <subfield code="d">2019-09-04</subfield>
    <subfield code="l">0</subfield>
    <subfield code="o">QA278 .Q56 2012 rs Thesis</subfield>
    <subfield code="p">0000067936</subfield>
    <subfield code="r">2019-09-04 00:00:00</subfield>
    <subfield code="t">1</subfield>
    <subfield code="w">2019-09-04</subfield>
    <subfield code="y">THESIS</subfield>
  </datafield>
  <datafield tag="952" ind1=" " ind2=" ">
    <subfield code="0">0</subfield>
    <subfield code="1">0</subfield>
    <subfield code="2">lcc</subfield>
    <subfield code="4">0</subfield>
    <subfield code="7">1</subfield>
    <subfield code="a">10000</subfield>
    <subfield code="b">10000</subfield>
    <subfield code="d">2019-09-04</subfield>
    <subfield code="l">0</subfield>
    <subfield code="o">CD 6310 | QA278 .Q56 2012 rs Thesis</subfield>
    <subfield code="p">0000067937</subfield>
    <subfield code="r">2019-09-04 00:00:00</subfield>
    <subfield code="t">1</subfield>
    <subfield code="w">2019-09-04</subfield>
    <subfield code="y">THESIS</subfield>
  </datafield>
  <datafield tag="999" ind1=" " ind2=" ">
    <subfield code="a">VIRTUA40</subfield>
    <subfield code="c">3866</subfield>
    <subfield code="d">3872</subfield>
  </datafield>
  <datafield tag="999" ind1=" " ind2=" ">
    <subfield code="a">VTLSSORT0080*0200*0400*0900*1000*2450*2600*3000*5020*5040*5200*6500*6501*8560*9992</subfield>
  </datafield>
</record>
