Siti Hawa Binti Apandi,

Web page classification based on recursive Convolutional Neural Network (CNN) and word cloud image Siti Hawa Binti Apandi - xi, 131 pages : illustrations ; 30 cm. + CD-ROM

Faculty of Computing

Thesis (Doctor of Philosophy) -- Universiti Malaysia Pahang – 2025

The increasing quantity and diversity of web content pose tremendous challenges to web page categorization, particularly in handling unstructured data and adapting to the dynamic nature of web content. Traditional classification methods relying on Convolutional Neural Networks (CNN) are confronted with hierarchical feature learning and long-range dependencies in text data. To overcome these constraints, this research proposes a novel web page classification model by integrating Recursive CNN with word cloud images as an alternative feature representation approach. The study aims to improve the accuracy of classification, interpretability, and computational expense by applying CNN's image processing capability to text data. The research process entails three phases: (1) Identification of effective text representation methods, (2) Developing a CNN-based classification model, and (3) Performance evaluation of the model against existing classification techniques. Web page text is extracted from HTML source codes, vectorized using the bag-of-words method, and depicted as word cloud images with high-frequency words presented in bold. The Recursive CNN model is trained using a training dataset of 391-word cloud images (274 Gaming and 117 Online Video Streaming) and tuned through grid search hyperparameter tuning. The maximum validation accuracy of the model is 89.50% for Adam optimizer, learning rate of 0.001, batch size 32, and 20 epochs, with a training time of 1 minute and 34 seconds. The overall accuracy is 89.50%, confirming the effectiveness of Recursive CNN in detecting deeper word relationships and reducing misclassification errors. Comparative analysis with existing CNN-based models demonstrates the improved adaptability and efficiency of the proposed method, confirming that word cloud images can be a good input representation for deep learning-based classification tasks. The findings provide a new benchmark for classifying web pages, reaffirming the feasibility of recursive feature extraction approaches as an instrument for future document classification, sentiment, and automated content filtering applications.

THE0010157 (Local)


Faculty of Computing--Dissertations


Universities and colleges--Dissertations