Enhanced Hybrid Vision Transformer with Multi-Scale Feature Integration and Patch Dropping for Facial Expression Recognition

被引：1

作者：

Li, Nianfeng ^{[1
]}

Huang, Yongyuan ^{[1
]}

Wang, Zhenyan ^{[1
]}

Fan, Ziyao ^{[1
]}

Li, Xinyuan ^{[1
]}

Xiao, Zhiguo ^{[1
,2
]}

机构：

[1] Changchun Univ, Coll Food Sci & Engn, 6543 Satellite Rd, Changchun 130022, Peoples R China

[2] Beijing Inst Technol, Sch Comp Sci Technol, Beijing 100811, Peoples R China

来源：

SENSORS | 2024年 / 24卷 / 13期

关键词：

facial expression recognition; lightweight network; attention module; transformer; CONVOLUTIONAL NEURAL-NETWORK; ATTENTION;

D O I：

10.3390/s24134153

中图分类号：

O65 [分析化学];

学科分类号：

070302 ; 081704 ;

摘要：

Convolutional neural networks (CNNs) have made significant progress in the field of facial expression recognition (FER). However, due to challenges such as occlusion, lighting variations, and changes in head pose, facial expression recognition in real-world environments remains highly challenging. At the same time, methods solely based on CNN heavily rely on local spatial features, lack global information, and struggle to balance the relationship between computational complexity and recognition accuracy. Consequently, the CNN-based models still fall short in their ability to address FER adequately. To address these issues, we propose a lightweight facial expression recognition method based on a hybrid vision transformer. This method captures multi-scale facial features through an improved attention module, achieving richer feature integration, enhancing the network's perception of key facial expression regions, and improving feature extraction capabilities. Additionally, to further enhance the model's performance, we have designed the patch dropping (PD) module. This module aims to emulate the attention allocation mechanism of the human visual system for local features, guiding the network to focus on the most discriminative features, reducing the influence of irrelevant features, and intuitively lowering computational costs. Extensive experiments demonstrate that our approach significantly outperforms other methods, achieving an accuracy of 86.51% on RAF-DB and nearly 70% on FER2013, with a model size of only 3.64 MB. These results demonstrate that our method provides a new perspective for the field of facial expression recognition.

引用

页数：18

共 50 条

[1] Feature fusion of multi-granularity and multi-scale for facial expression recognition
Xia, Haiying
Lu, Lidan
Song, Shuxiang
VISUAL COMPUTER, 2024, 40 (03): : 2035 - 2047
[2] Feature fusion of multi-granularity and multi-scale for facial expression recognition
Haiying Xia
Lidan Lu
Shuxiang Song
The Visual Computer, 2024, 40 : 2035 - 2047
[3] Patch attention convolutional vision transformer for facial expression recognition with occlusion
Liu, Chang
Hirota, Kaoru
Dai, Yaping
INFORMATION SCIENCES, 2023, 619 : 781 - 794
[4] A multi-scale feature fusion convolutional neural network for facial expression recognition
Zhang, Xiufeng
Fu, Xingkui
Qi, Guobin
Zhang, Ning
EXPERT SYSTEMS, 2024, 41 (04)
[5] Progressive Multi-Scale Vision Transformer for Facial Action Unit Detection
Wang, Chongwen
Wang, Zicheng
FRONTIERS IN NEUROROBOTICS, 2022, 15 (15):
[6] Facial Expression Recognition Based on Multi-scale CNNs
Zhou, Shuai
Liang, Yanyan
Wan, Jun
Li, Stan Z.
BIOMETRIC RECOGNITION, 2016, 9967 : 503 - 510
[7] Facial Expression Recognition Based on Vision Transformer with Hybrid Local Attention
Tian, Yuan
Zhu, Jingxuan
Yao, Huang
Chen, Di
APPLIED SCIENCES-BASEL, 2024, 14 (15):
[8] Facial expression-based emotion recognition across diverse age groups: a multi-scale vision transformer with contrastive learning approach
Balachandran, G.
Ranjith, S.
Chenthil, T. R.
Jagan, G. C.
JOURNAL OF COMBINATORIAL OPTIMIZATION, 2025, 49 (01)
[9] Multi-scale Feature Relation Modeling for Facial Expression Restoration
Liu, Zhilei
Wu, Yunpeng
Zhang, Cuicui
2021 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2021,
[10] Multi-Scale Feature For Recognition
Lei, Songze
Hao, Chongyang
Qi, Min
ICECT: 2009 INTERNATIONAL CONFERENCE ON ELECTRONIC COMPUTER TECHNOLOGY, PROCEEDINGS, 2009, : 277 - 280

← 1 2 3 4 5 →