A tree based keyphrase extraction technique for academic literature / (Record no. 91238)

MARC details
000 -LEADER
fixed length control field 03224nam a22003737a 4500
003 - CONTROL NUMBER IDENTIFIER
control field MY-KuUP
005 - DATE AND TIME OF LATEST TRANSACTION
control field 20251125105432.0
006 - FIXED-LENGTH DATA ELEMENTS--ADDITIONAL MATERIAL CHARACTERISTICS
fixed length control field a|||||r|||| 000 0
007 - PHYSICAL DESCRIPTION FIXED FIELD--GENERAL INFORMATION
fixed length control field ta
008 - FIXED-LENGTH DATA ELEMENTS--GENERAL INFORMATION
fixed length control field 191119t20192019my a|||| |m|| 000 0 eng d
020 ## - INTERNATIONAL STANDARD BOOK NUMBER
International Standard Book Number THE0008292(Local)
040 ## - CATALOGING SOURCE
Original cataloging agency UMP
Language of cataloging eng
Transcribing agency UMP
Description conventions rda
090 ## - LOCALLY ASSIGNED LC-TYPE CALL NUMBER (OCLC); LOCAL CALL NUMBER (RLIN)
Classification number (OCLC) (R) ; Classification number, CALL (RLIN) (NR) FSKKP .G65 2019 r Thesis
100 1# - MAIN ENTRY--PERSONAL NAME
Personal name Rabby, Gollam,
Relator term author.
245 12 - TITLE STATEMENT
Title A tree based keyphrase extraction technique for academic literature /
Statement of responsibility, etc. Gollam Rabby
264 #1 - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Place of production, publication, distribution, manufacture Kuantan, Pahang :
Name of producer, publisher, distributor, manufacturer UMP,
Date of production, publication, distribution, manufacture, or copyright notice 2019
264 #4 - PRODUCTION, PUBLICATION, DISTRIBUTION, MANUFACTURE, AND COPYRIGHT NOTICE
Place of production, publication, distribution, manufacture © 2019
300 ## - PHYSICAL DESCRIPTION
Extent xii, 102 pages :
Other physical details illustrations (some color) ;
Dimensions 30 cm. +
Accompanying material 1 CD-ROM
336 ## - CONTENT TYPE
Source rdacontent
Content type term text
336 ## - CONTENT TYPE
Source rdacontent
Content type term text
337 ## - MEDIA TYPE
Source rdamedia
Media type term unmediated
337 ## - MEDIA TYPE
Source rdamedia
Media type term computer
338 ## - CARRIER TYPE
Source rdacarrier
Carrier type term volume
338 ## - CARRIER TYPE
Source rdacarrier
Carrier type term computer disc
347 ## - DIGITAL FILE CHARACTERISTICS
Source rda
File type text file
Encoding format PDF
500 ## - GENERAL NOTE
General note Faculty of Computer Systems and Software Engineering
502 ## - DISSERTATION NOTE
Dissertation note Thesis (Master of Science) -- Universiti Malaysia Pahang – 2019
504 ## - BIBLIOGRAPHY, ETC. NOTE
Bibliography, etc. note Includes bibliographical references
520 3# - SUMMARY, ETC.
Summary, etc. Automatic keyphrase extraction techniques aim to extract quality keyphrases to summarize a document at a higher level. Among the existing techniques some of them are domain-specific and require application domain knowledge, some of them are based on higher-order statistical methods and are computationally expensive, and some of them require large train data which are rare for many applications. Overcoming these issues, this thesis proposes a new unsupervised automatic keyphrase extraction technique, named TeKET or Tree-based Keyphrase Extraction Technique, which is domain-independent, employs limited statistical knowledge, and requires no train data. The proposed technique also introduces a new variant of the binary tree, called KeyPhrase Extraction (KePhEx) tree to extract final keyphrases from candidate keyphrases. Depending on the candidate keyphrases the KePhEx tree structure is either expanded or shrunk or maintained. In addition, a measure, called Cohesiveness Index or CI, is derived that denotes the degree of cohesiveness of a given node with respect to the root which is used in extracting final keyphrases from a resultant tree in a flexible manner and is utilized in ranking keyphrases alongside Term Frequency. The effectiveness of the proposed technique is evaluated using an experimental evaluation on a benchmark corpus, called SemEval-2010 with total 244 train and test articles, and compared with other relevant unsupervised techniques by taking the representatives from both statistical (such as Term Frequency-Inverse Document Frequency and YAKE) and graph-based techniques (PositionRank, CollabRank (SingleRank), TopicRank, and MultipartiteRank) into account. Three evaluation metrics, namely precision, recall and F1 score are taken into consideration during the experiments. The obtained results demonstrate the improved performance of the proposed technique over other similar techniques in terms of precision, recall, and F1 scores.
610 20 - SUBJECT ADDED ENTRY--CORPORATE NAME
Corporate name or jurisdiction name as entry element Faculty of Computer Systems and Software Engineering
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Disertations
General subdivision Universities and colleges
650 #0 - SUBJECT ADDED ENTRY--TOPICAL TERM
Topical term or geographic name entry element Theses
942 ## - ADDED ENTRY ELEMENTS (KOHA)
Source of classification or shelving scheme Library of Congress Classification
Koha item type Thesis
Holdings
Withdrawn status Lost status Source of classification or shelving scheme Damaged status Not for loan Collection Home library Current library Shelving location Date acquired Total checkouts Full call number Barcode Date last seen Price effective from Koha item type
  Not lost Library of Congress Classification   Not for loan Reference UMPLIB PEKAN UMPLIB PEKAN Reference 19/11/2019   FSKKP .G65 2019 r Thesis T000000421 11/11/2021 19/11/2019 Thesis
  Not lost Library of Congress Classification   Not for loan Reference UMPLIB PEKAN UMPLIB PEKAN Reference 19/11/2019   CD 12380 T000000422 10/01/2022 19/11/2019 Thesis

Perpustakaan Universiti Malaysia Pahang Al-Sultan Abdullah
26600 Pekan, Pahang Darul Makmur
Phone: +609 431 5063 (Gambang) / +609 431 5035 (Pekan)
Email: umplibrary@umpsa.edu.my

Connect With Us