MonoATT: Online Monocular 3D Object Detection with Adaptive Token Transformer

被引：11

作者：

Zhou, Yunsong ^{[1
]}

Zhu, Hongzi ^{[1
]}

Liu, Quan ^{[1
]}

Chang, Shan ^{[2
]}

Guo, Minyi ^{[1
]}

机构：

[1] Shanghai Jiao Tong Univ, Shanghai, Peoples R China

[2] Donghua Univ, Shanghai, Peoples R China

来源：

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2023年

基金：

上海市自然科学基金; 中国国家自然科学基金;

关键词：

D O I：

10.1109/CVPR52729.2023.01678

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Mobile monocular 3D object detection (Mono3D) (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Existing transformer-based offline Mono3D models adopt grid-based vision tokens, which is suboptimal when using coarse tokens due to the limited available computational power. In this paper, we propose an online Mono3D framework, called MonoATT, which leverages a novel vision transformer with heterogeneous tokens of varying shapes and sizes to facilitate mobile Mono3D. The core idea of MonoATT is to adaptively assign finer tokens to areas of more significance before utilizing a transformer to enhance Mono3D. To this end, we first use prior knowledge to design a scoring network for selecting the most important areas of the image, and then propose a token clustering and merging network with an attention mechanism to gradually merge tokens around the selected areas in multiple stages. Finally, a pixel-level feature map is reconstructed from heterogeneous tokens before employing a SOTA Mono3D detector as the underlying detection core. Experiment results on the real-world KITTI dataset demonstrate that MonoATT can effectively improve the Mono3D accuracy for both near and far objects and guarantee low latency. MonoATT yields the best performance compared with the state-of-the-art methods by a large margin and is ranked number one on the KITTI 3D benchmark.

引用

页码：17493 / 17503

页数：11

共 50 条

[31] A New Monocular 3D Object Detection with Neural Network
Hong, Weijie
Liu, Yiguang
Zheng, Yunan
Wang, Ying
Shi, Xuelei
PATTERN RECOGNITION AND COMPUTER VISION (PRCV 2018), PT IV, 2018, 11259 : 174 - 185
[32] 3D Visual Object Detection from Monocular Images
Wang, Qiaosong
Rasmussen, Christopher
ADVANCES IN VISUAL COMPUTING, ISVC 2019, PT I, 2020, 11844 : 168 - 180
[33] Competition for roadside camera monocular 3D object detection
Jia, Jinrang
Shi, Yifeng
Qu, Yuli
Wang, Rui
Xu, Xing
Zhang, Hai
NATIONAL SCIENCE REVIEW, 2023, 10 (06)
[34] Monocular 3D Object Detection with Depth from Motion
Wang, Tai
Pang, Jiangmiao
Lin, Dahua
COMPUTER VISION, ECCV 2022, PT IX, 2022, 13669 : 386 - 403
[35] Objects are Different: Flexible Monocular 3D Object Detection
Zhang, Yunpeng
Lu, Jiwen
Zhou, Jie
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 3288 - 3297
[36] Monocular 3D object detection for construction scene analysis
Shen, Jie
Jiao, Lang
Zhang, Cong
Peng, Keran
COMPUTER-AIDED CIVIL AND INFRASTRUCTURE ENGINEERING, 2024, 39 (09) : 1370 - 1389
[37] Delving into Localization Errors for Monocular 3D Object Detection
Ma, Xinzhu
Zhang, Yinmin
Xu, Dan
Zhou, Dongzhan
Yi, Shuai
Li, Haojie
Ouyang, Wanli
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 4719 - 4728
[38] Shape-Aware Monocular 3D Object Detection
Chen, Wei
Zhao, Jie
Zhao, Wan-Lei
Wu, Song-Yuan
IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, 2023, 24 (06) : 6416 - 6424
[39] Competition for roadside camera monocular 3D object detection
Jinrang Jia
Yifeng Shi
Yuli Qu
Rui Wang
Xing Xu
Hai Zhang
NationalScienceReview, 2023, 10 (06) : 34 - 37
[40] MonoGRNet: A General Framework for Monocular 3D Object Detection
Qin, Zengyi
Wang, Jinglu
Lu, Yan
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (09) : 5170 - 5184

← 1 2 3 4 5 →