Enhancing Classification with Hierarchical Scalable Query on Fusion Transformer

被引:2
|
作者
Sahoo, Sudeep Kumar [1 ]
Chalasani, Sathish [1 ]
Joshi, Abhishek [1 ]
Iyer, Kiran Nanjunda [1 ]
机构
[1] Samsung R&D Inst, Bangalore, India
关键词
D O I
10.1109/ICCE56470.2023.10043496
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Real-world vision based applications require finegrained classification for various applications of interest like e-commerce, mobile applications, warehouse management, etc. where reducing the severity of mistakes and improving the classification accuracy is of utmost importance. This paper proposes a method to boost fine-grained classification through a hierarchical approach via learnable independent query embeddings. This is achieved through a classification network that uses coarse class predictions to improve the fine class accuracy in a stage-wise sequential manner. We exploit the idea of hierarchy to learn query embeddings that are scalable across all levels, thus making this a relevant approach even for extreme classification where we have a large number of classes. The query is initialized with a weighted Eigen image calculated from training samples to best represent and capture the variance of the object. We introduce transformer blocks to fuse intermediate layers at which query attention happens to enhance the spatial representation of feature maps at different scales. This multi-scale fusion helps improve the accuracy of small-size objects. We propose a twofold approach for the unique representation of learnable queries. First, at each hierarchical level, we leverage cluster based loss that ensures maximum separation between inter-class query embeddings and helps learn a better (query) representation in higher dimensional spaces. Second, we fuse coarse level queries with finer level queries weighted by a learned scale factor. We additionally introduce a novel block called Cross Attention on Multi-level queries with Prior (CAMP) Block that helps reduce error propagation from coarse level to finer level, which is a common problem in all hierarchical classifiers. Our method is able to outperform the existing methods with an improvement of about 11% at the fine-grained classification.
引用
收藏
页数:6
相关论文
共 50 条
  • [31] Deep Hierarchical Vision Transformer for Hyperspectral and LiDAR Data Classification
    Xue, Zhixiang
    Tan, Xiong
    Yu, Xuchu
    Liu, Bing
    Yu, Anzhu
    Zhang, Pengqiang
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2022, 31 : 3095 - 3110
  • [32] Hierarchical Pretrained Backbone Vision Transformer for Image Classification in Histopathology
    Zedda, Luca
    Loddo, Andrea
    Di Ruberto, Cecilia
    IMAGE ANALYSIS AND PROCESSING, ICIAP 2023, PT II, 2023, 14234 : 223 - 234
  • [33] Enhancing time series forecasting: A hierarchical transformer with probabilistic decomposition representation
    Tong, Junlong
    Xie, Liping
    Yang, Wankou
    Zhang, Kanjian
    Zhao, Junsheng
    INFORMATION SCIENCES, 2023, 647
  • [34] Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement
    Hou, Xiuquan
    Liu, Meiqin
    Zhang, Senlin
    Wei, Ping
    Chen, Badong
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 17574 - 17583
  • [35] Enhancing Natural Language Query to SQL Query Generation Through Classification-Based Table Selection
    Chopra, Ankush
    Azam, Rauful
    ENGINEERING APPLICATIONS OF NEURAL NETWORKS, EANN 2024, 2024, 2141 : 152 - 165
  • [36] DRCCT: Enhancing Diabetic Retinopathy Classification with a Compact Convolutional Transformer
    Touati, Mohamed
    Touati, Rabeb
    Nana, Laurent
    Benzarti, Faouzi
    Ben Yahia, Sadok
    BIG DATA AND COGNITIVE COMPUTING, 2025, 9 (01)
  • [37] Enhancing dynamic ECG heartbeat classification with lightweight transformer model
    Meng, Lingxiao
    Tan, Wenjun
    Ma, Jiangang
    Wang, Ruofei
    Yin, Xiaoxia
    Zhang, Yanchun
    ARTIFICIAL INTELLIGENCE IN MEDICINE, 2022, 124
  • [38] Enhancing Auditory Brainstem Response Classification Based On Vision Transformer
    Ahmed, Hunar Abubakir
    Majidpour, Jafar
    Ahmed, Mohammed Hussein
    Jameel, Samer Kais
    Majidpour, Amir
    COMPUTER JOURNAL, 2023, 67 (05): : 1872 - 1878
  • [39] Enhancing Few-Shot Image Classification With Cosine Transformer
    Nguyen, Quang-Huy
    Nguyen, Cuong Q.
    Le, Dung D. D.
    Pham, Hieu H.
    IEEE ACCESS, 2023, 11 : 79659 - 79672
  • [40] Hierarchical Spectral-Spatial Transformer for Hyperspectral and Multispectral Image Fusion
    Zhu, Tianxing
    Liu, Qin
    Zhang, Lixiang
    REMOTE SENSING, 2024, 16 (22)