Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

被引:1
|
作者
Li, Jiaming [1 ]
Zhang, Jiacheng [1 ]
Li, Jichang [1 ,2 ]
Li, Ge [3 ]
Liu, Si [4 ]
Lin, Liang [1 ]
Li, Guanbin [1 ,5 ,6 ]
机构
[1] Sun Yat Sen Univ, Sch Comp Sci & Engn, Guangzhou, Peoples R China
[2] Univ Hong Kong, Dept Comp Sci, Hong Kong, Peoples R China
[3] Peking Univ, Shenzhen Grad Sch, SECE, Shenzhen, Peoples R China
[4] Beihang Univ, Inst Artificial Intelligence, Beijing, Peoples R China
[5] GuangDong Prov Key Lab Informat Secur Technol, Shenzhen, Guangdong, Peoples R China
[6] Sun Yat Sen Univ, Res Inst, Shenzhen, Guangdong, Peoples R China
基金
中国国家自然科学基金;
关键词
D O I
10.1109/CVPR52733.2024.01578
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of object detection, significantly generalizing the powerful capabilities of the detector to identify more unknown object categories. However, these methods face significant challenges in background interpretation and model overfitting and thus often result in the loss of crucial back-ground knowledge, giving rise to sub-optimal inference performance of the detector. To mitigate these issues, we present a novel OVD framework termed LBP to propose learning background prompts to harness explored implicit background knowledge, thus enhancing the detection performance w.r.t. base and novel categories. Specifically, we devise three modules: Background Category-specific Prompt, Background Object Discovery, and Inference Probability Rectification, to empower the detector to discover, represent, and leverage implicit object knowledge explored from background proposals. Evaluation on two benchmark datasets, OV-COCO and OV-LVIS, demonstrates the superiority of our proposed method over existing state-of-the-art approaches in handling the OVD tasks.
引用
收藏
页码:16678 / 16687
页数:10
相关论文
共 50 条
  • [31] Moving Object Detection Based on Background Compensation and Deep Learning
    Zhu, Juncai
    Wang, Zhizhong
    Wang, Songwei
    Chen, Shuli
    SYMMETRY-BASEL, 2020, 12 (12): : 1 - 17
  • [32] Implicit-Explicit Motion Learning for Video Camouflaged Object Detection
    Hui, Wenjun
    Zhu, Zhenfeng
    Gu, Guanghua
    Liu, Meiqin
    Zhao, Yao
    IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 7188 - 7196
  • [33] Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation
    Hinami, Ryota
    Satoh, Shin'ichi
    2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP 2018), 2018, : 2605 - 2615
  • [34] Teaching for breadth and depth of vocabulary knowledge: Learning from explicit and implicit instruction and the storybook texts
    Dickinson, David K.
    Nesbitt, Kimberly T.
    Collins, Molly F.
    Hadley, Elizabeth B.
    Newman, Katherine
    Rivera, Bretta L.
    Ilgez, Hande
    Nicolopoulou, Ageliki
    Golinkoff, Roberta Michnick
    Hirsh-Pasek, Kathy
    EARLY CHILDHOOD RESEARCH QUARTERLY, 2019, 47 : 341 - 356
  • [35] Learning Efficient Object Detection Models with Knowledge Distillation
    Chen, Guobin
    Choi, Wongun
    Yu, Xiang
    Han, Tony
    Chandraker, Manmohan
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 30 (NIPS 2017), 2017, 30
  • [36] Simple Image-Level Classification Improves Open-Vocabulary Object Detection
    Fang, Ruohuan
    Pang, Guansong
    Bai, Xiao
    THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 2, 2024, : 1716 - 1725
  • [37] Relation Knowledge Distillation by Auxiliary Learning for Object Detection
    Wang, Hao
    Jia, Tong
    Wang, Qilong
    Zuo, Wangmeng
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2024, 33 : 4796 - 4810
  • [38] Open-vocabulary object detection via debiased curriculum self-training
    Zhang, Hanlue
    Guan, Dayan
    Ke, Xiangrui
    El Saddik, Abdulmotaleb
    Lu, Shijian
    EXPERT SYSTEMS WITH APPLICATIONS, 2024, 255
  • [39] DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection
    Yao, Lewei
    Pi, Renjie
    Hang, Jianhua
    Liang, Xiaodan
    Xu, Hang
    Zhang, Wei
    Li, Zhenguo
    Xu, Dan
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 27381 - 27391
  • [40] OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection
    Zhang, Hu
    Ku, Jianhua
    Tang, Tao
    Sun, Haiyang
    Huang, Xin
    Huang, Zi
    Yu, Kaicheng
    COMPUTER VISION - ECCV 2024, PT LXXXIV, 2025, 15142 : 1 - 19