Hybrid-attention and frame difference enhanced network for micro-video venue recognition

被引:2
|
作者
Wang, Bing [1 ]
Huang, Xianglin [1 ]
Cao, Gang [1 ]
Yang, Lifang [1 ]
Wei, Xiaolong [1 ]
Tao, Zhulin [1 ]
机构
[1] Commun Univ China, State Key Lab Media Convergence & Commun, Dingfuzhuang 1, Beijing 100024, Peoples R China
基金
中国国家自然科学基金;
关键词
Micro-video venue recognition; robust visual features; hybrid attention module; difference enhanced module; SCENE;
D O I
10.3233/JIFS-213191
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Many micro-video related applications, such as personalized location recommendation and micro-video verification, can be benefited greatly from the venue information. Most existing works focus on integrating the information from multi-modal for exact venue category recognition. It is important to make full use of the information from different modalities. However, the performance may be limited by the lacked acoustic modality or textual descriptions in uploaded micro-videos. Therefore, in this paper visual modality is explored as the only modality according to its rich and indispensable semantic information. To this end, a hybrid-attention and frame difference enhanced network (HAFDN) is proposed to generate the comprehensive venue representation. Such network mainly contains two parallel branches: content and motion branches. Specifically, in the content branch, a domain-adaptive CNN model combined with temporal shift module (TSM) is employed to extract discriminative visual features. Then, a novel hybrid attention module (HAM) is introduced to enhance extracted features via three attention mechanisms. In HAM, channel attention, local and global spatial attention mechanisms are used to capture salient visual information from different views. In addition, convolutional Long Short-Term Memory (convLSTM) is enforced after HAM to better encode the long spatial-temporal dependency. A difference-enhanced module parallel with HAM is devised to learn the content variations among adjacent frames, which is usually ignored in prior works. Moreover, in the motion branch, 3D-CNNs and LSTM are used to capture movement variation as a supplement of content branch in a different form. Finally, the features from two branches are fused to generate robust video-level representations for predicting venue categories. Extensive experimental results on public datasets verify the effectiveness of the proposed micro-video venue recognition scheme. The source code is available at https://github.com/hs8945/HAFDN.
引用
收藏
页码:3337 / 3353
页数:17
相关论文
共 50 条
  • [21] Hybrid-Attention Network for RGB-D Salient Object Detection
    Chen, Yuzhen
    Zhou, Wujie
    [J]. APPLIED SCIENCES-BASEL, 2020, 10 (17):
  • [22] Intention-convolution and hybrid-attention network for vehicle trajectory prediction
    Li, Chao
    Liu, Zhanwen
    Lin, Shan
    Wang, Yang
    Zhao, Xiangmo
    [J]. Expert Systems with Applications, 2024, 236
  • [23] Heterogeneous Hierarchical Feature Aggregation Network for Personalized Micro-Video Recommendation
    Cai, Desheng
    Qian, Shengsheng
    Fang, Quan
    Xu, Changsheng
    [J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 : 805 - 818
  • [24] Heterogeneous Graph Contrastive Learning Network for Personalized Micro-Video Recommendation
    Cai, Desheng
    Qian, Shengsheng
    Fang, Quan
    Hu, Jun
    Ding, Wenkui
    Xu, Changsheng
    [J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2023, 25 : 2761 - 2773
  • [25] Spatiotemporal information deep fusion network with frame attention mechanism for video action recognition
    Ou, Hongshi
    Sun, Jifeng
    [J]. JOURNAL OF ELECTRONIC IMAGING, 2019, 28 (02)
  • [26] An adaptive frame selection network with enhanced dilated convolution for video smoke recognition
    Tao, Huanjie
    Duan, Qianyue
    [J]. EXPERT SYSTEMS WITH APPLICATIONS, 2023, 215
  • [27] Concept-Aware Denoising Graph Neural Network for Micro-Video Recommendation
    Liu, Yiyu
    Liu, Qian
    Tian, Yu
    Wang, Changping
    Niu, Yanan
    Song, Yang
    Li, Chenliang
    [J]. PROCEEDINGS OF THE 30TH ACM INTERNATIONAL CONFERENCE ON INFORMATION & KNOWLEDGE MANAGEMENT, CIKM 2021, 2021, : 1099 - 1108
  • [28] A Vlogger-augmented Graph Neural Network Model for Micro-video Recommendation
    Lai, Weijiang
    Jin, Beihong
    Li, Beibei
    Zheng, Yiyuan
    Zhao, Rui
    [J]. MACHINE LEARNING AND KNOWLEDGE DISCOVERY IN DATABASES: APPLIED DATA SCIENCE AND DEMO TRACK, ECML PKDD 2023, PT VI, 2023, 14174 : 684 - 699
  • [29] Multimodal Progressive Modulation Network for Micro-Video Multi-Label Classification
    Jing, Peiguang
    Zhao, Xuan
    Fan, Fugui
    Yang, Fan
    Li, Yun
    Su, Yuting
    [J]. IEEE Transactions on Multimedia, 2024, 26 : 10134 - 10144
  • [30] Hybrid-attention guided network with multiple resolution features for person re-identification
    Zhang, Guoqing
    Yang, Junchuan
    Zheng, Yuhui
    Wang, Ye
    Wu, Yi
    Chen, Shengyong
    [J]. INFORMATION SCIENCES, 2021, 578 : 525 - 538