Enhanced ontological query expansion model using bigram combinations and combined similarity measure for improving information retrieval / (Record no. 94669)

MARC details
000 -LEADER
fixed length control field 04538nam a2200349 i 4500
003 - CONTROL NUMBER IDENTIFIER
control field MY-KuUP
005 - DATE AND TIME OF LATEST TRANSACTION
control field 20251125105757.0
006 - FIXED-LENGTH DATA ELEMENTS--ADDITIONAL MATERIAL CHARACTERISTICS
fixed length control field a||||fr|||| 000 0
007 - PHYSICAL DESCRIPTION FIXED FIELD--GENERAL INFORMATION
fixed length control field ta
008 - FIXED-LENGTH DATA ELEMENTS--GENERAL INFORMATION
fixed length control field 201105t20202020my a|||frma|| 000 0 eng d
020 ## - INTERNATIONAL STANDARD BOOK NUMBER
International Standard Book Number THE0008568(Local)
040 ## - CATALOGING SOURCE
Original cataloging agency UMP
Language of cataloging eng
Transcribing agency UMP
Description conventions rda
090 ## - LOCALLY ASSIGNED LC-TYPE CALL NUMBER (OCLC); LOCAL CALL NUMBER (RLIN)
Classification number (OCLC) (R) ; Classification number, CALL (RLIN) (NR) FKOM .A37 2020 r Thesis
100 1# - MAIN ENTRY--PERSONAL NAME
Personal name Raza, Muhammad Ahsan,
Relator term author.
245 10 - TITLE STATEMENT
Title Enhanced ontological query expansion model using bigram combinations and combined similarity measure for improving information retrieval /
Statement of responsibility, etc. Muhammad Ahsan Raza
264 #1 - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Place of production, publication, distribution, manufacture Kuantan, Pahang :
Name of producer, publisher, distributor, manufacturer UMP ,
Date of production, publication, distribution, manufacture, or copyright notice 2020
264 #4 - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Place of production, publication, distribution, manufacture ©2020
300 ## - PHYSICAL DESCRIPTION
Extent xii, 155 pages :
Other physical details illustrations (some color) ;
Dimensions 30 cm. +
Accompanying material 1 CD ROM
336 ## - CONTENT TYPE
Content type term text
Source rdacontent
336 ## - CONTENT TYPE
Content type term text
Source rdacontent
337 ## - MEDIA TYPE
Media type term unmediated
Source rdamedia
337 ## - MEDIA TYPE
Media type term computer
Source rdamedia
338 ## - CARRIER TYPE
Carrier type term volume
Source rdacarrier
338 ## - CARRIER TYPE
Carrier type term computer disc
Source rdacarrier
347 ## - DIGITAL FILE CHARACTERISTICS
File type text file
Encoding format PDF
Source rda
500 ## - GENERAL NOTE
General note Faculty of Computing
502 ## - DISSERTATION NOTE
Dissertation note Thesis (Doctor of Philosophy) -- Universiti Malaysia Pahang – 2020
504 ## - BIBLIOGRAPHY, ETC. NOTE
Bibliography, etc. note The information explosion over the Web has been increasing and changing rapidly over the time, thus the effective retrieval of information is increasingly gaining in importance. Most Information Retrieval (IR) systems typically rely on query and document keyword matching, in order to search over huge amounts of Web data, examples being famous search engines such as Google, Bing or Yahoo. Problem arising with these simple keyword matching IR systems is vocabulary mismatch issue: the searcher’s query terms may not be matched with those of the corpus. IR systems cannot always expect a user to type the exact keyword in query as present in corpus in order to obtain relevant documents. To deal with this issue, several efforts have been made such as query expansion, whereby the search query is expanded with additional relevant terms using original query keywords. In recent years, ontology based query expansion (OQE) emerges as an advance query expansion model to expand search query semantically using ontology knowledgebase. However, common problems with existing OQE model include (i) the inherent ambiguity of natural language search query, (ii) term-based expansion to support unstructured search query rather than considering multiple query terms together (iii) The expansion of search query with irrelevant terms. The main objective of this research is to improve existing OQE model and propose an enhanced ontological query expansion model (EOQE) for effective information retrieval. The EOQE model attempts to semantically expand unstructured natural language search queries in order to retrieve relevant documents for computer science discipline. The model overcomes the limitations of existing OQE model by following three main steps. First, the query refinement step performs linguistic processing of search query and generates valid search term. Second, the enhanced ontology based expansion step disambiguates the search query and generates the additional expansion concepts on the basis of bigram combinations technique. Third, the query formulation step filters irrelevant terms from expansion concepts set using combined similarity measure technique. The performance evaluation EOQE model was based on comparing the retrieval results of queries expanded with EOQE model and the original queries (called as baseline model). On Vector Space Model (VSM) standard IR system, the EOQE model showed 32% improvements in terms of mean average precision against baseline model, and achieved above 90% recall values for most of search queries. The EOQE model also attained a 17% and 15% increase in P@20 and P@40 values, respectively, than baseline model over famous Google search system. Furthermore, the EOQE model demonstrated competitive performance in terms of precision, recall, average precision, and mean average precision values against EOQE model variants based on single similarity measure. The main contributions of this research are to introduce a model to semantically expand unstructured and ambiguous natural language query using bigram combinations and strong combined similarity measure techniques. These contributions enable exploiting multiple query terms in procedure of OQE rather than using individual query terms, and formulating expanded queries with more relevant semantic concepts.
610 20 - SUBJECT ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element Faculty of Computing
General subdivision Dissertations
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Universities and colleges
General subdivision Dissertations
942 ## - ADDED ENTRY ELEMENTS (KOHA)
Source of classification or shelving scheme Library of Congress Classification
Koha item type Thesis
Holdings
Withdrawn status Lost status Source of classification or shelving scheme Damaged status Not for loan Collection Home library Current library Shelving location Date acquired Total checkouts Full call number Barcode Date last seen Price effective from Koha item type
  Not lost Library of Congress Classification     Reference UMPLIB PEKAN UMPLIB PEKAN Reference 05/11/2020   FKOM .A37 2020 r Thesis T000001028 31/12/2020 05/11/2020 Thesis
  Not lost Library of Congress Classification   Not for loan Reference UMPLIB PEKAN UMPLIB PEKAN   05/11/2020   CD12742 T000001029 19/02/2021 05/11/2020 Thesis

Perpustakaan Universiti Malaysia Pahang Al-Sultan Abdullah
26600 Pekan, Pahang Darul Makmur
Phone: +609 431 5063 (Gambang) / +609 431 5035 (Pekan)
Email: umplibrary@umpsa.edu.my

Connect With Us