Transfer learning-based deep learning models for intelligent human frontal facial emotion recognition / Anbananthan Pillai A/l Munanday

By: Material type: TextTextPublisher: Kuantan, Pahang : UMPSA, 2024Copyright date: © 2024Description: xvi, 85 pages : illustrations (some color) ; 30 cm. + 1 CD-ROMContent type:
  • text
Media type:
  • unmediated
Carrier type:
  • volume
ISBN:
  • THE0009971 (Local)
Subject(s): Dissertation note: Thesis (Master of Science) -- Universiti Malaysia Pahang – 2024 Abstract: Achieving precise facial emotion recognition (FER) is a critical endeavor in artificial intelligence to enhance interactions between humans and computers. However, current Convolutional Neural Network (CNN) architectures face substantial obstacles, including structural limitations and reliance on controlled datasets, leading to diminished accuracy and processing speed. These issues hinder the ability to accurately capture and interpret nuanced human emotions, especially in non-frontal or profile images. This study addresses the recognizing and interpreting human emotions from facial expressions by using an improved Convolutional Neural Network (CNN) architecture. This architecture is an adaptation of the well-known VGG-16 model. Numerous research efforts have applied to CNNs, incorporating a restricted layer structure, to solve challenges in Facial Emotion Recognition (FER). Nonetheless, conventional shallow CNNs, with their basic learning strategies, struggle to adequately extract features necessary for discerning emotional content in high-definition images. A common limitation among many existing approaches is their focus solely on frontal images, disregarding the significance of profile images captured from various angles, which are crucial for the functionality of a comprehensive FER system. To forge a highly precise FER system, this research introduces an advanced Deep CNN (DCNN) approach via the Transfer Learning (TL) method. This method involves the use of a preexisting DCNN framework, specifically VGG-16, by altering its upper dense layer(s) to suit FER requirements, and further refining the model with facial emotion datasets. The efficacy of this refined CNN model is evaluated against five different pre-trained DCNN architectures (ResNet50, VGG-16, MobileNetV2, EfficientNetB0, Xception) using three benchmark datasets FER2013, CK+, own dataset and KDEF with 80:20 training and testing split. The study focused on optimizing different parameters of the CNN, including the number of layers, activation functions, learning rates, and dropout rates. These adjustments were aimed at enhancing the performance of facial emotion recognition (FER) across several datasets. The goal was to accurately identify the seven universally recognized facial expressions: happiness, sadness, anger, fear, disgust, surprise, and neutrality with both frontal and side views. The suggested approach attained significant precision on these datasets using pre-trained models. Through a 10- fold cross-validation method, the highest facial emotion recognition (FER) accuracies obtained with a refined CNN model on the test sets of FER-2013, CK+, own dataset and KDEF were 93.02%, 97.62%, 93.95% and 86.04%, respectively. Additionally, the performance on the KDEF dataset, specifically with profile views, was encouraging, showcasing the necessary effectiveness for practical real-world application. The major findings indicated that the optimized CNN model achieved significant improvements in test accuracy, reaching 95.47% on FER2013, 93.51% on CK+, 91.11% on own dataset and 93.64% on KDEF, outperforming the traditional VGG-16 model and several other benchmarks from the literature. Significantly, the refined model demonstrated the ability to identify emotions across a broad range of facial expressions, encompassing profile views, among different demographic categories, and even when facial features were partially obscured. This research not only highlights the improvements in accuracy and generalizability of the refined CNN model but also sets the stage for future studies. Future research will focus on integrating more diverse and complex datasets, exploring real-time FER applications, and further optimizing CNN architectures to enhance their robustness and applicability in real-world scenarios.
Tags from this library: No tags from this library for this title. Log in to add tags.
Star ratings
    Average rating: 0.0 (0 votes)
Holdings
Item type Current library Call number Status Date due Barcode
Thesis Thesis UMPLIB PEKAN FTKPM .A53 2024 r Thesis (Browse shelf(Opens below)) Not for loan T000003311
Thesis Thesis UMPLIB PEKAN CD13660 (Browse shelf(Opens below)) Not for loan T000003312

Faculty of Manufacturing and Mechatronic Engineering Technology

Thesis (Master of Science) -- Universiti Malaysia Pahang – 2024

Includes bibliographical references

Achieving precise facial emotion recognition (FER) is a critical endeavor in artificial intelligence to enhance interactions between humans and computers. However, current Convolutional Neural Network (CNN) architectures face substantial obstacles, including structural limitations and reliance on controlled datasets, leading to diminished accuracy and processing speed. These issues hinder the ability to accurately capture and interpret nuanced human emotions, especially in non-frontal or profile images. This study addresses the recognizing and interpreting human emotions from facial expressions by using an improved Convolutional Neural Network (CNN) architecture. This architecture is an adaptation of the well-known VGG-16 model. Numerous research efforts have applied to CNNs, incorporating a restricted layer structure, to solve challenges in Facial Emotion Recognition (FER). Nonetheless, conventional shallow CNNs, with their basic learning strategies, struggle to adequately extract features necessary for discerning emotional content in high-definition images. A common limitation among many existing approaches is their focus solely on frontal images, disregarding the significance of profile images captured from various angles, which are crucial for the functionality of a comprehensive FER system. To forge a highly precise FER system, this research introduces an advanced Deep CNN (DCNN) approach via the Transfer Learning (TL) method. This method involves the use of a preexisting DCNN framework, specifically VGG-16, by altering its upper dense layer(s) to suit FER requirements, and further refining the model with facial emotion datasets. The efficacy of this refined CNN model is evaluated against five different pre-trained DCNN architectures (ResNet50, VGG-16, MobileNetV2, EfficientNetB0, Xception) using three benchmark datasets FER2013, CK+, own dataset and KDEF with 80:20 training and testing split. The study focused on optimizing different parameters of the CNN, including the number of layers, activation functions, learning rates, and dropout rates. These adjustments were aimed at enhancing the performance of facial emotion recognition (FER) across several datasets. The goal was to accurately identify the seven universally recognized facial expressions: happiness, sadness, anger, fear, disgust, surprise, and neutrality with both frontal and side views. The suggested approach attained significant precision on these datasets using pre-trained models. Through a 10- fold cross-validation method, the highest facial emotion recognition (FER) accuracies obtained with a refined CNN model on the test sets of FER-2013, CK+, own dataset and KDEF were 93.02%, 97.62%, 93.95% and 86.04%, respectively. Additionally, the performance on the KDEF dataset, specifically with profile views, was encouraging, showcasing the necessary effectiveness for practical real-world application. The major findings indicated that the optimized CNN model achieved significant improvements in test accuracy, reaching 95.47% on FER2013, 93.51% on CK+, 91.11% on own dataset and 93.64% on KDEF, outperforming the traditional VGG-16 model and several other benchmarks from the literature. Significantly, the refined model demonstrated the ability to identify emotions across a broad range of facial expressions, encompassing profile views, among different demographic categories, and even when facial features were partially obscured. This research not only highlights the improvements in accuracy and generalizability of the refined CNN model but also sets the stage for future studies. Future research will focus on integrating more diverse and complex datasets, exploring real-time FER applications, and further optimizing CNN architectures to enhance their robustness and applicability in real-world scenarios.

Perpustakaan Universiti Malaysia Pahang Al-Sultan Abdullah
26600 Pekan, Pahang Darul Makmur
Phone: +609 431 5063 (Gambang) / +609 431 5035 (Pekan)
Email: umplibrary@umpsa.edu.my

Connect With Us