Modified word representation vector based scalar weight for contextual text classification / (Record no. 101215)

MARC details
000 -LEADER
fixed length control field 04409ntm a2200337 i 4500
003 - CONTROL NUMBER IDENTIFIER
control field MY-KuUP
005 - DATE AND TIME OF LATEST TRANSACTION
control field 20251125110922.0
006 - FIXED-LENGTH DATA ELEMENTS--ADDITIONAL MATERIAL CHARACTERISTICS
fixed length control field t||||fr|||| 000 0
007 - PHYSICAL DESCRIPTION FIXED FIELD--GENERAL INFORMATION
fixed length control field ta
008 - FIXED-LENGTH DATA ELEMENTS--GENERAL INFORMATION
fixed length control field 240902s20242024my a|||fr|||| 00| 0 eng d
020 ## - INTERNATIONAL STANDARD BOOK NUMBER
International Standard Book Number THE0009940 (Local)
Qualifying information Hardback
040 ## - CATALOGING SOURCE
Original cataloging agency UMPSA
Language of cataloging eng
Transcribing agency UMP
Description conventions rda
090 ## - LOCALLY ASSIGNED LC-TYPE CALL NUMBER (OCLC); LOCAL CALL NUMBER (RLIN)
Classification number (OCLC) (R) ; Classification number, CALL (RLIN) (NR) FKOM .A23 2024 r Thesis
100 0# - MAIN ENTRY--PERSONAL NAME
Personal name Abbas Saliimi Lokman,
Relator term author.
245 10 - TITLE STATEMENT
Title Modified word representation vector based scalar weight for contextual text classification /
Statement of responsibility, etc. Abbas Saliimi Lokman
264 #1 - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Place of production, publication, distribution, manufacture Kuantan, Pahang :
Name of producer, publisher, distributor, manufacturer UMPSA,
Date of production, publication, distribution, manufacture, or copyright notice 2024
264 #4 - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Date of production, publication, distribution, manufacture, or copyright notice © 2024
300 ## - PHYSICAL DESCRIPTION
Extent xii, 120 pages :
Other physical details illustration ;
Materials specified 0 cm. +30 cm. +
Accompanying material 1 CD-ROM
336 ## - CONTENT TYPE
Source rdacontent
Content type term text
337 ## - MEDIA TYPE
Source rdamedia
Media type term unmediated
338 ## - CARRIER TYPE
Source rdacarrier
Carrier type term volume
347 ## - DIGITAL FILE CHARACTERISTICS
Source rda
File type text file
Encoding format PDF
500 ## - GENERAL NOTE
General note Faculty of Computing
502 ## - DISSERTATION NOTE
Dissertation note Thesis (Master of Science) -- Universiti Malaysia Pahang – 2022
504 ## - BIBLIOGRAPHY, ETC. NOTE
-- ncludes bibliographical referencesIncludes bibliographical references
520 3# - SUMMARY, ETC.
Summary, etc. This thesis investigates contextual text classification, which is the process of categorising textual data into different classes or categories based on its meaning within a given context. Central to this process is the representation of words through vectors for computational interpretation. Current practices employ Large Language Models (LLMs) to generate contextualised word representation vectors, achieved through pre-training the LLM on vast corpora that enables it to grasp intricate language patterns and context. For contextual text classification, the pre-trained LLM is further train on classificationspecific labeled data in a process called fine-tuning. Although this approach is currently considered the most optimal in the field, it poses a notable challenge due to the substantial demand for computing resources stemming from the vast number of trainable parameters in LLMs. Furthermore, although pre-trained LLMs can generate contextualised word representation vectors, they lack the flexibility to modify the semantic significance of these vectors outside of the LLM, necessitating fine-tuning for the modification of word vectors. To bridge this gap, a five-phase research methodology is structured to propose and evaluate an algorithm enabling the external modification of LLM-generated word vectors using scalar values as the focus weightage. To validate this algorithm, the modified word vectors are compared with original LLM-generated word vectors to evaluate their reflection of the intended context. In addition, a contextual text classification experiment is conducted using benchmarked datasets to assess the performance of the modified word vectors in the targeted classification task. For this experiment, the modified word vectors serve as input to train a Machine Learning (ML) model for the text classification process, aiming for the developed ML model to have a significantly smaller parameter count. This experiment aims to determine the effectiveness of the modified word vectors in contextual text classification tasks, utilizing a more computationally efficient approach. Based on the acquired results, the experiments reveal that the modified word vectors algorithm can effectively alter original LLM-generated word vectors to reflect intended contexts and can outperform baseline scores in contextual text classification tasks. Evaluation metrics including Accuracy, Precision, Recall, and F1 score are employed in the evaluation process, with Accuracy and F1 score serving as primary metrics. The evaluation showcases significant improvements, with the test ML model achieving a best accuracy score of 0.571, a 46% increase from the baseline, and a best F1 score of 0.727, a 30% increment from the baseline. Overall, this thesis presents five contributions: the proposed modified word vectors algorithm, the new contextual classification dataset named QCoC, the efficient question-type classifier based on the feed-forward neural network algorithm, the potential transferability of the presented work to other domains, and the practical implications of the presented work towards cases where computational resources are limited or costly.
610 20 - SUBJECT ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element Faculty of Computing
General subdivision Dissertations
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Universities and colleges
General subdivision Dissertations
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Theses
General subdivision Dissertations
942 ## - ADDED ENTRY ELEMENTS (KOHA)
Source of classification or shelving scheme Library of Congress Classification
Koha item type Thesis
Holdings
Withdrawn status Lost status Source of classification or shelving scheme Damaged status Not for loan Home library Current library Date acquired Total checkouts Full call number Barcode Date last seen Price effective from Koha item type Copy number
  Not lost Library of Congress Classification   Not for loan UMPLIB PEKAN UMPLIB PEKAN 02/09/2024   FKOM .A23 2024 r Thesis T000003327 02/09/2024 02/09/2024 Thesis  
  Not lost Library of Congress Classification   Final Processing UMPLIB PEKAN UMPLIB PEKAN 02/09/2024   CD13668 T000003328 02/09/2024 02/09/2024 Thesis 1

Perpustakaan Universiti Malaysia Pahang Al-Sultan Abdullah
26600 Pekan, Pahang Darul Makmur
Phone: +609 431 5063 (Gambang) / +609 431 5035 (Pekan)
Email: umplibrary@umpsa.edu.my

Connect With Us