Web page classification based on recursive Convolutional Neural Network (CNN) and word cloud image (Record no. 103739)

MARC details
000 -LEADER
fixed length control field 03224ntm a2200313 i 4500
003 - CONTROL NUMBER IDENTIFIER
control field MY-KuUP
005 - DATE AND TIME OF LATEST TRANSACTION
control field 20251216111123.0
006 - FIXED-LENGTH DATA ELEMENTS--ADDITIONAL MATERIAL CHARACTERISTICS
fixed length control field t||||fr|||| 000 0
007 - PHYSICAL DESCRIPTION FIXED FIELD--GENERAL INFORMATION
fixed length control field ta
008 - FIXED-LENGTH DATA ELEMENTS--GENERAL INFORMATION
fixed length control field 251216t20252025my a|||fr|||| 000 0 eng d
020 ## - INTERNATIONAL STANDARD BOOK NUMBER
International Standard Book Number THE0010157 (Local)
Qualifying information Hardback
040 ## - CATALOGING SOURCE
Original cataloging agency UMPSA
Language of cataloging eng
Transcribing agency UMPSA
Description conventions rda
090 ## - LOCALLY ASSIGNED LC-TYPE CALL NUMBER (OCLC); LOCAL CALL NUMBER (RLIN)
Classification number (OCLC) (R) ; Classification number, CALL (RLIN) (NR) FKOM .H39 2025 r Thesis
100 1# - MAIN ENTRY--PERSONAL NAME
Personal name Siti Hawa Binti Apandi,
Relator term author.
245 10 - TITLE STATEMENT
Title Web page classification based on recursive Convolutional Neural Network (CNN) and word cloud image
Statement of responsibility, etc. Siti Hawa Binti Apandi
264 ## - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Place of production, publication, distribution, manufacture Kuantan, Pahang :
Name of producer, publisher, distributor, manufacturer UMPSA,
Date of production, publication, distribution, manufacture, or copyright notice 2025
264 ## - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Date of production, publication, distribution, manufacture, or copyright notice © 2025
300 ## - PHYSICAL DESCRIPTION
Extent xi, 131 pages :
Other physical details illustrations ;
Dimensions 30 cm. +
Accompanying material CD-ROM
336 ## - CONTENT TYPE
Source rdacontent
Content type term text
337 ## - MEDIA TYPE
Source rdamedia
Media type term unmediated
338 ## - CARRIER TYPE
Source rdacarrier
Carrier type term volume
347 ## - DIGITAL FILE CHARACTERISTICS
Source rda
File type text file
Encoding format PDF
500 ## - GENERAL NOTE
General note Faculty of Computing
502 ## - DISSERTATION NOTE
Dissertation note Thesis (Doctor of Philosophy) -- Universiti Malaysia Pahang – 2025
504 ## - BIBLIOGRAPHY, ETC. NOTE
Bibliography, etc. note The increasing quantity and diversity of web content pose tremendous challenges to web page categorization, particularly in handling unstructured data and adapting to the dynamic nature of web content. Traditional classification methods relying on Convolutional Neural Networks (CNN) are confronted with hierarchical feature learning and long-range dependencies in text data. To overcome these constraints, this research proposes a novel web page classification model by integrating Recursive CNN with word cloud images as an alternative feature representation approach. The study aims to improve the accuracy of classification, interpretability, and computational expense by applying CNN's image processing capability to text data. The research process entails three phases: (1) Identification of effective text representation methods, (2) Developing a CNN-based classification model, and (3) Performance evaluation of the model against existing classification techniques. Web page text is extracted from HTML source codes, vectorized using the bag-of-words method, and depicted as word cloud images with high-frequency words presented in bold. The Recursive CNN model is trained using a training dataset of 391-word cloud images (274 Gaming and 117 Online Video Streaming) and tuned through grid search hyperparameter tuning. The maximum validation accuracy of the model is 89.50% for Adam optimizer, learning rate of 0.001, batch size 32, and 20 epochs, with a training time of 1 minute and 34 seconds. The overall accuracy is 89.50%, confirming the effectiveness of Recursive CNN in detecting deeper word relationships and reducing misclassification errors. Comparative analysis with existing CNN-based models demonstrates the improved adaptability and efficiency of the proposed method, confirming that word cloud images can be a good input representation for deep learning-based classification tasks. The findings provide a new benchmark for classifying web pages, reaffirming the feasibility of recursive feature extraction approaches as an instrument for future document classification, sentiment, and automated content filtering applications.
610 20 - SUBJECT ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element Faculty of Computing
General subdivision Dissertations
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Universities and colleges
General subdivision Dissertations
942 ## - ADDED ENTRY ELEMENTS (KOHA)
Source of classification or shelving scheme Library of Congress Classification
Koha item type Thesis
Holdings
Withdrawn status Lost status Source of classification or shelving scheme Damaged status Not for loan Permanent Location Current Location Date acquired Total Checkouts Full call number Barcode Date last seen Copy number Price effective from Koha item type
  Not lost Library of Congress Classification   Not for loan UMPLIB PEKAN UMPLIB PEKAN 16/12/2025   FKOM .H39 2025 r Thesis T000003847 16/12/2025 1 16/12/2025 Thesis
  Not lost Library of Congress Classification   Final Processing UMPLIB PEKAN UMPLIB PEKAN 16/12/2025   CD13794 T000003848 16/12/2025   16/12/2025 Thesis

Perpustakaan Universiti Malaysia Pahang Al-Sultan Abdullah
26600 Pekan, Pahang Darul Makmur
Phone: +609 431 5063 (Gambang) / +609 431 5035 (Pekan)
Email: umplibrary@umpsa.edu.my

Connect With Us