Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

被引:1
|
作者
Li, Jiaming [1 ]
Zhang, Jiacheng [1 ]
Li, Jichang [1 ,2 ]
Li, Ge [3 ]
Liu, Si [4 ]
Lin, Liang [1 ]
Li, Guanbin [1 ,5 ,6 ]
机构
[1] Sun Yat Sen Univ, Sch Comp Sci & Engn, Guangzhou, Peoples R China
[2] Univ Hong Kong, Dept Comp Sci, Hong Kong, Peoples R China
[3] Peking Univ, Shenzhen Grad Sch, SECE, Shenzhen, Peoples R China
[4] Beihang Univ, Inst Artificial Intelligence, Beijing, Peoples R China
[5] GuangDong Prov Key Lab Informat Secur Technol, Shenzhen, Guangdong, Peoples R China
[6] Sun Yat Sen Univ, Res Inst, Shenzhen, Guangdong, Peoples R China
基金
中国国家自然科学基金;
关键词
D O I
10.1109/CVPR52733.2024.01578
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of object detection, significantly generalizing the powerful capabilities of the detector to identify more unknown object categories. However, these methods face significant challenges in background interpretation and model overfitting and thus often result in the loss of crucial back-ground knowledge, giving rise to sub-optimal inference performance of the detector. To mitigate these issues, we present a novel OVD framework termed LBP to propose learning background prompts to harness explored implicit background knowledge, thus enhancing the detection performance w.r.t. base and novel categories. Specifically, we devise three modules: Background Category-specific Prompt, Background Object Discovery, and Inference Probability Rectification, to empower the detector to discover, represent, and leverage implicit object knowledge explored from background proposals. Evaluation on two benchmark datasets, OV-COCO and OV-LVIS, demonstrates the superiority of our proposed method over existing state-of-the-art approaches in handling the OVD tasks.
引用
收藏
页码:16678 / 16687
页数:10
相关论文
共 50 条
  • [41] Open-Vocabulary Object Detection by Novel-Class Feature Perception Enhancement
    Hui, Kanghua
    Cai, Xianqiao
    Zhang, Zhi
    Huang, Rui
    Liu, Qing
    ADVANCED INTELLIGENT COMPUTING TECHNOLOGY AND APPLICATIONS, PT IV, ICIC 2024, 2024, 14865 : 220 - 231
  • [42] Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers
    Kim, Dahun
    Angelova, Anelia
    Kuo, Weicheng
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 11144 - 11154
  • [43] YOLO-World: Real-Time Open-Vocabulary Object Detection
    Cheng, Tianheng
    Sone, Lin
    Ge, Yixiao
    Liu, Wenyu
    Wang, Xinggang
    Shan, Yong
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 16901 - 16911
  • [44] MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
    Wang, Kuo
    Cheng, Lechao
    Chen, Weikai
    Zhang, Pingping
    Lin, Liang
    Zhou, Fan
    Li, Guanbin
    COMPUTER VISION - ECCV 2024, PT XVII, 2025, 15075 : 106 - 122
  • [45] Learning What to Remember: Vocabulary Knowledge and Children's Memory for Object Names and Features
    Perry, Lynn K.
    Axelsson, Emma L.
    Horst, Jessica S.
    INFANT AND CHILD DEVELOPMENT, 2016, 25 (04): : 247 - 258
  • [46] Open-World Human-Object Interaction Detection via Multi-modal Prompts
    Yang, Jie
    Li, Bingliang
    Zeng, Ailing
    Zhang, Lei
    Zhang, Ruimao
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 16954 - 16964
  • [47] FOREGROUND DETECTION: COMBINING BACKGROUND SUBSPACE LEARNING WITH OBJECT SMOOTHING MODEL
    Xue, Gengjian
    Song, Li
    Sun, Jun
    Zhou, Jun
    2013 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO (ICME 2013), 2013,
  • [48] Feature Fusion Based Background Model Learning for Video Object Detection
    Padhi, Aditya Narayan
    Acharya, Subhabrata
    Nanda, Pradipta Kumar
    2020 IEEE REGION 10 SYMPOSIUM (TENSYMP) - TECHNOLOGY FOR IMPACTFUL SUSTAINABLE DEVELOPMENT, 2020, : 126 - 129
  • [49] Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection
    Rasheed, Hanoona
    Maaz, Muhammad
    Khattak, Muhammad Uzair
    Khan, Salman
    Khan, Fahad Shahbaz
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 35 (NEURIPS 2022), 2022,
  • [50] Open Source Assessment of Deep Learning Visual Object Detection
    Paniego, Sergio
    Sharma, Vinay
    Maria Canas, Jose
    SENSORS, 2022, 22 (12)